Using Built-in Functions for Shadow Denoising in Ray Tracing Applications

Visibility sampling and built-in wave function detection of penumbra areas through the dispatchable unit of the parallel processor, generating a penumbra mask, solving the problem of low shadow denoising efficiency in ray tracing rendering, and achieving a more efficient rendering process.

CN114764841BActive Publication Date: 2025-07-18NVIDIA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210033280.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-14
Filing Date
2022-01-12
Publication Date
2025-07-18
Estimated Expiration
2042-01-12

AI Technical Summary

Technical Problem

The existing ray tracing rendering technology requires a large amount of shadow ray sampling when generating shadows, resulting in high computing resources consumption and long rendering time. At the same time, traditional denoising technology requires expensive post-processing transfer and global memory access, which affects efficiency.

Method used

The threads of the scheduling unit of the parallel processor perform visibility sampling, use the built-in function to detect the penumbra area, and generate a penumbra mask through statistical values to avoid the application of denoising filters on the external penumbra area, and dynamically determine the denoising filter parameters to reduce processing time.

Benefits of technology

It effectively reduces rendering time and computing resource consumption, improves shadow denoising efficiency, avoids expensive post-processing transfer and global memory access, and realizes a faster rendering process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114764841B_ABST
    Figure CN114764841B_ABST
Patent Text Reader

Abstract

Disclosed is the use of built-in functions for shadow denoising in ray tracing applications. In an example, threads of a dispatch unit of a parallel processor (e.g., a warp or wavefront) can be used to sample the visibility of pixels relative to one or more light sources. The threads can receive results of samples executed by other threads in the dispatch unit to compute a value indicating whether a region corresponds to a penumbra (e.g., using wave built-in functions). Each thread can correspond to respective pixels, and the region can correspond to pixels of the dispatch unit. A frame can be divided into regions, where each region corresponds to a respective dispatch unit. When denoising ray traced shadow information, the values of the regions can be used to avoid applying a denoising filter to pixels of regions outside the penumbra while applying the denoising filter to pixels of regions within the penumbra.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] Ray tracing is a method for rendering images by tracing the path of light in a virtual environment and simulating the interaction of light with virtual objects. Ray tracing techniques can be used to simulate various optical effects - such as shadows, reflections and refractions, scattering phenomena, and dispersion phenomena (such as chromatic aberration). When using ray tracing to render soft shadows, conventional methods of shadow tracing can project any number of shadow rays from a location in the virtual environment to sample the illumination conditions of a pixel relative to a light source. Samples of the ray tracing can be combined and applied to the pixel. In the penumbra (the region where light is partially blocked in the shadow), some shadow rays may be visible to the light source while other shadow rays may be blocked. A large number of shadow rays may be required to converge the combined illumination conditions to an accurate result. To conserve computational resources and reduce rendering time, the shadow rays can be sparsely sampled, resulting in noisy shadow data. The noisy shadow data can be filtered using denoising techniques to reduce the noise and produce a final rendering that more closely approximates the rendering of a fully sampled scene.

[0002] Computational resources for denoising shadow data can be reduced by focusing the denoising on pixels within the penumbra. For example, fully illuminated or fully shadowed pixels outside the penumbra do not require denoising because the corresponding ray tracing samples reflect those pixels as being in shadow. A penumbra mask can be generated and used to indicate which pixels are within the penumbra during denoising. Generating the penumbra mask typically involves a post-processing pass performed on the shadow data and can be computationally expensive due to accessing global memory. SUMMARY OF THE INVENTION

[0003] Embodiments of the present disclosure relate to using wave intrinsic functions to detect penumbra regions for shadow denoising. Specifically, the present disclosure relates in part to leveraging threads of a dispatch unit of a parallel processor, the dispatch unit being used to sample visibility in ray tracing in order to identify penumbra regions for denoising the shadows of ray tracing.

[0004] Compared with conventional methods, the disclosed method can be used to determine which pixels of a frame are within the penumbra while avoiding a post - processing pass. According to aspects of the present disclosure, threads in a schedulable unit (e.g., a warp or wavefront) of a parallel processor can be used to sample the visibility of pixels relative to one or more light sources. At least one of the threads can receive the results of sampling performed by other threads (e.g., each other thread) in the schedulable unit to calculate a value (e.g., using wave intrinsics of the parallel processor) indicating whether a region corresponds to the penumbra. In at least one embodiment, each thread can correspond to a respective pixel, and the region can correspond to the pixels of the schedulable unit. Further, a frame can be divided into pixel regions, where each region corresponds to a respective schedulable unit. When applying a denoising pass to ray - traced shadow information, the value of the region can be used to avoid applying a denoising filter to pixels in regions outside the penumbra while applying the denoising filter to pixels in regions within the penumbra. For example, the value can be used to generate a penumbra mask, and the penumbra mask can be used to denoise a shadow mask.

[0005] The present disclosure further provides a method for determining parameters of a denoising filter. According to aspects of the present disclosure, threads of a schedulable unit can be used to sample one or more aspects of a scene (e.g., visibility, global illumination, ambient occlusion, etc.). At least one of the threads can receive the results of sampling performed by other threads (e.g., each other thread) in the schedulable unit to calculate a value (e.g., using wave intrinsics of the parallel processor) indicating one or more properties of a region of the scene. When applying a denoising pass to rendering data, the value of the region can be used to determine one or more parameters of the denoising filter applied to the rendering data. For example, the value can be used to determine a filter radius and / or a range of values included in the filtering. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The following describes in detail the present system and method for detecting a penumbra region for shadow denoising using wave intrinsics with reference to the accompanying drawings, where:

[0007] Figure 1 is a data - flow diagram showing an example process for generating an output image using an image rendering system according to some embodiments of the present disclosure;

[0008] Figure 2 is a diagram showing an example of how rendered values can correspond to mask values according to some embodiments of the present disclosure;

[0009] Figure 3 is a diagram showing an example of capturing ray - traced samples of a virtual environment according to some embodiments of the present disclosure;

[0010] Figure 4is a flowchart showing an example of a method for using dispatchable units to determine visibility values and values indicating that positions in a scene correspond to penumbra values, according to some embodiments of the present disclosure;

[0011] Figure 5 is a flowchart showing an example of a method for using a thread group of one or more dispatchable units to determine ray tracing samples for visibility and values indicating whether a pixel corresponds to a penumbra, according to some embodiments of the present disclosure;

[0012] Figure 6 is a flowchart showing an example of a method for using dispatchable units to determine ray tracing samples and one or more values of one or more parameters for determining a denoising filter, according to some embodiments of the present disclosure;

[0013] Figure 7 is a block diagram of an example computing environment suitable for implementing some embodiments of the present disclosure; and

[0014] Figure 8 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0015] The present disclosure relates to using wave built-in functions to detect penumbra regions for shadow denoising. Specifically, the present disclosure provides in part the use of threads of dispatchable units of a parallel processor, the threads of the dispatchable units being used to sample one or more aspects of a virtual environment relative to a pixel (e.g., by executing shader code). In embodiments for determining which pixels are in the penumbra, the conditions may include the visibility of the pixel relative to one or more light sources.

[0016] The disclosed method can be used to determine statistical values for informing the denoising of rendered data without the need for a dedicated post-processing pass. For example, the statistical values can be used to determine which pixels of a frame are within the penumbra during the denoising of rendered data and / or to guide the filtering of rendered data. The rendered data can include spatial and / or temporal ray tracing samples.

[0017] According to aspects of the present disclosure, threads of dispatchable units (e.g., warps or wavefronts) of a parallel processor can be used to sample one or more aspects of a virtual environment relative to a pixel (e.g., by executing shader code). In embodiments for determining which pixels are in the penumbra, the conditions may include the visibility of the pixel relative to one or more light sources.

[0018] Threads can be arranged into thread groups, where a thread group can refer to each thread of a scheduling unit or a subset of threads of a schedulable unit. At least one thread from the threads in the group can receive the results of sampling performed by other threads within the group. One or more threads can compute statistical values regarding samples of ray tracing. For example, for visibility, each thread can compute a value indicating whether a region of a frame corresponds to penumbra. In at least one embodiment, wave built-in functions of a parallel processor can be used to retrieve values corresponding to ray tracing samples from other threads. For example, a wave active sum function can return the sum of values (statistical value) to a thread. The statistical values computed by the threads can be used to inform filtering of rendering data. For example, the statistical value can be used as a mask value or can be used by a thread to compute a mask value. The mask value can be stored in a mask, which can be accessed during denoising. In at least one embodiment, the mask can be a penumbra mask indicating which pixels correspond to penumbra.

[0019] In at least one embodiment, each thread can correspond to respective pixels, and the region of the frame for which the statistical value is computed can correspond to the pixels of a thread group. Further, a frame can be divided into pixel regions, where each region corresponds to respective thread groups and / or schedulable units. Using the disclosed method, it may not be necessary to have a post-processing pass to determine the statistical values for informing denoising of rendering data, thereby reducing the processing time for denoising rendering data. For example, threads of a schedulable unit can determine samples of a virtual environment and statistical values from the samples (e.g., as part of executing a ray generation shader). The statistical values can be computed from registers of the threads, which can have a significantly lower access time than the memory used for post-processing.

[0020] In at least one embodiment, when applying a denoising pass to rendering data (e.g., ray tracing samples), the statistical values of regions can be used to avoid applying a denoising filter to one or more pixels of the region. For example, in a case where the mask value of a penumbra mask indicates that a region is outside the penumbra, the denoising filter may not be applied to the pixels within the region. The present disclosure further provides a method for determining one or more parameters of a denoising filter. For example, in addition to or instead of using statistical values to determine which pixels to skip when applying a denoising filter, statistical values can be used to determine one or more parameters of the denoising filter for pixels. Examples of parameters include parameters defining filter radius, filter weights, and / or ranges of values to include in filtering.

[0021] See Figure 1 , Figure 1is a data flow diagram showing an example process 140 for generating an output image 120 using an image rendering system 100, in accordance with some embodiments of the present disclosure. Such and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or in place of those shown, and some elements may be omitted altogether. Further, many of the elements described herein may be implemented as discrete or distributed components or as functional entities combined with other components and implemented in any suitable combination, arrangement, or location. The various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software. For example, the various functions may be performed by a processor executing instructions stored in a memory.

[0022] In at least one embodiment, the image rendering system 100 may be implemented at least partially in Figure 8 the data center 800. As a different example, the image rendering system 100 may be included in or include one or more of a system for performing simulation operations, a system for performing simulation operations to test or validate autonomous machine applications, a system for performing deep learning operations, a system implemented using edge devices, a system incorporating one or more virtual machines (VMs), a system implemented at least partially in a data center, or a system implemented at least partially using cloud computing resources.

[0023] The image rendering system 100 may be configured to render images of a virtual environment, such as Figure 3 the virtual environment 300. To render an image of a virtual environment (which may also be referred to as a scene), the image rendering system 100 may employ a ray tracer 102, an image filter 104, an image combiner 106, and a three-dimensional (3D) renderer 108.

[0024] The ray tracer 102 may be configured to trace rays passing through the virtual environment using any of a variety of potential ray tracing techniques in order to generate ray tracing samples of one or more aspects of the virtual environment (e.g., lighting conditions such as visibility) with respect to positions in the virtual environment. One or more parallel processors may be used, such as Figure 7One or more dispatch units of the GPU 708) to determine ray tracing samples. Further, the ray tracing samples can be captured or otherwise used to generate rendering data 122 (e.g., by the dispatch units). The ray tracer 102 can also be configured (e.g., by the dispatch units) to compute values from the ray tracing samples, such as statistical values (e.g., the sum of visibility values in the dispatch units). The values can be determined from the rendering data 122 using the dispatch units and can indicate whether the corresponding location and / or pixel corresponds to the penumbra of a shadow. In an embodiment, these values can be captured or otherwise used to generate mask data 124.

[0025] The image filter 104 can be configured to filter the rendering data 122 (or other rendering data) from the ray tracer 102 based at least on the values computed from the ray tracing samples (e.g., using the mask data 124). For example, in the case where the value indicates that the location or pixel corresponds to the penumbra (e.g., Figure 1 a white pixel in the mask data 124), the denoiser can apply a denoising filter to that location or pixel. In the case where the value indicates that the location or pixel corresponds to a fully illuminated or fully shadowed region (e.g., Figure 1 a black pixel in the mask data 124), the denoiser can skip applying the denoising filter to that location or pixel (also referred to as "early exit").

[0026] In some embodiments, the image combiner 106 can combine data corresponding to the filtered rendering data 122 (e.g., filtered shadow / visibility data) with data representing a 3D rendering of the virtual environment (e.g., shadow data without ray tracing) to generate an output image 120. The 3D renderer 108 can be configured to generate the 3D rendering using any suitable method that may or may not include ray tracing. In an embodiment, the 3D rendering can include pixel color information for a frame of the virtual environment.

[0027] As described herein, the ray tracer 102 can be configured to trace rays through a virtual environment using any of a variety of potential ray tracing techniques to generate ray tracing samples of one or more aspects of the virtual environment relative to positions in the virtual environment. The ray tracer 102 can also be configured to compute values from the ray tracing samples, such as statistical values (e.g., sum of visibility values in a dispatch unit), which can be utilized by other components of the rendering pipeline, such as the image filter 104. In various embodiments, the ray tracer 102 can utilize dispatch units of one or more processors for parallel processing to generate ray tracing samples and values derived from the ray tracing samples. In doing so, values (e.g., as reflected in the mask data 124) can be derived without the need for a post-processing pass. Among other potential advantages, values can be computed faster than using conventional methods because, as opposed to shared or global memory, data (e.g., rendering data 122) for computing the values can be accessed from the registers of the threads.

[0028] In various embodiments, a dispatch unit can refer to a group of hardware dispatchable threads that can be used for parallel processing. A thread can refer to a GPU thread or a CPU thread. In different examples, threads can be implemented using at least in part a single instruction, multiple threads (SIMT) execution model. A thread can also be referred to as a work item, a basic element of data to be processed, an individual lane, or a sequence of single instruction, multiple data (SIMD) lane operations.

[0029] Examples of dispatch units include warps associated with NVIDIA terminology (e.g., CUDA-based technologies) or wavefronts associated with AMD terminology (e.g., OpenCL-based technologies). For CUDA-based technologies, by way of example and not limitation, a dispatch unit can include 32 threads. For OpenCL-based technologies, by way of example and not limitation, a dispatch unit can include 64 threads. In one or more embodiments, a dispatch unit can refer to threads of SIMD instructions. In one or more embodiments, a dispatch unit can include a set of operations that execute in lockstep, run the same instruction, and follow the same control flow path. In some embodiments, individuals or groups of lanes or threads of a dispatch unit can be masked from execution.

[0030] In various embodiments, the ray tracer 102 can operate one or more shaders or programs that are executed by one or more dispatch units for parallel processing to generate ray tracing samples and values derived from the ray tracing samples. For example, the ray tracing samples can be generated by the same shader as the values derived from the ray tracing samples. A shader can be, for example, a ray generation shader, where the code of the ray generation shader can be executed by one or more thread groups and / or dispatch units (e.g., in parallel).

[0031] Now refer to Figure 2 , Figure 2 which is a diagram showing an example of how the values of rendering 200 can correspond to the values of mask 202 according to some embodiments of the present disclosure. Rendering 200 may correspond to Figure 1 the rendering data 122 and mask 202 may correspond to Figure 1 the mask data 124. In at least one embodiment, the ray tracer 102 may divide the rendering or frame into regions, where each region may correspond to one or more pixels and / or positions of the virtual environment 300. For example, rendering 200 may be divided into regions 210A, 210B, 210C, 210D, 210E to 210N (collectively also referred to herein as "regions 210"). In at least one embodiment, the ray tracer 102 (e.g., shader code executed by a thread) may configure regions 210 such that they do not overlap and enclose the entire frame or rendering. For example, in Figure 2 , each region 210 corresponds to a rectangular region of the pixels of the frame.

[0032] In the example shown, each region 210 corresponds to a dispatch unit, and each thread within the dispatch unit corresponds to a respective pixel or unit of region 210. In particular, the example shown relates to a warp, where each region 210 may correspond to 32 pixels and threads. Where the rendering 200 is H render pixels x V render pixels and each region is H region pixels x V region pixels, there may be H render / H region x V render / V region regions in the rendering 200. Figure 2 Each region 210 in

[0033] As described herein, the ray tracer 102 may use threads of dispatchable units to determine values corresponding to ray tracing samples. In at least one embodiment, pixels or cells of region 210 may store values of ray tracing samples and / or values derived from ray tracing samples. For example, each thread may store the value of the ray tracing sample for the pixel or cell generated by that thread in a register.

[0034] Regarding Figure 3 Examples of ray tracing samples are described. Now refer to Figure 3 , Figure 3 FIG. is a diagram illustrating an example of capturing ray tracing samples of a virtual environment 300 according to some embodiments of the present disclosure. The image rendering system 100 may be configured to use the ray tracer 102 to render an image using any number of ray tracing passes in order to sample the conditions of the virtual environment.

[0035] Figure 3 An example of is a sample regarding visibility and, more specifically, a sample regarding the visibility of one or more pixels relative to a light source in the virtual environment 300. In such an example, Figure 2 the rendering 200 of may correspond to a shadow mask for a frame. However, the disclosed method may be implemented with other types of ray tracing samples, which may include those that form binary signals (e.g., having values of 0 or 1) or non-binary signals. In some embodiments, the ray tracing samples may represent, indicate, or otherwise correspond to the ambient occlusion, global illumination, or other properties of one or more pixels and / or locations relative to the virtual environment 300. When sampling different aspects of the virtual environment, the ray tracing techniques may be adapted to suit the effects being simulated. Further, in the present example, when a ray interacts with a location in the virtual environment 300 (e.g., at the light source 320 or an occluder 322), no additional rays may be projected from that location. However, for other ray tracing effects or techniques, one or more additional rays may be projected from it.

[0036] In at least one embodiment, the ray tracer 102 may project or trace rays using a ray generation shader. Figure 3Various examples of rays through virtual environment 300 that ray tracer 102 can trace (e.g., using one ray per pixel) are shown relative to ray tracing pass 314. For example, rays 340, 342, and 344 are individually labeled among the nine rays shown for ray tracing pass 314. Ray tracer 102 can use the rays to sample one or more aspects of virtual environment 300 jointly with respect to positions in virtual environment 300. Examples of thirty-two positions in region 210D are shown, where positions 330, 332, and 334 are individually labeled. However, each region 210 can be sampled similarly in ray tracing pass 314.

[0037] In at least one embodiment, each ray is associated with (e.g., projected from) one of these positions and is used to generate a ray tracing sample for that position. For example, ray 340 is associated with position 332, ray 342 is associated with position 330, and ray 344 is associated with position 334. In some embodiments, each position from which ray tracer 102 projects a ray corresponds to a respective pixel of region 210, as shown. For example, positions (such as positions 330, 332, and 334) can be determined by transforming the virtual screen of the pixel (e.g., from a z-buffer) into world space. The virtual screen can represent the view of a camera in virtual environment 300, and in some embodiments, the positions can be referred to as (e.g., of rendering 200) pixels or world space pixels. In other examples, the positions may not have such a one-to-one correspondence with pixels. Further, in other examples, the positions can be determined as the respective points and / or regions where the respective eye rays (e.g., projected from the camera through the virtual screen including the pixel) interact with virtual environment 300.

[0038] In various embodiments, the accuracy of a sample at a location may be limited because each ray may provide only partial information about that location. As such, sampling the virtual environment 300 using a limited number of rays may result in noise in the image (particularly for certain locations in the virtual environment 300). To illustrate the foregoing, the rays used in the example shown are shadow rays for sampling one or more aspects of the lighting conditions at a location relative to the light source 320 in the virtual environment 300. The image rendering system 100 may use this information, for example, to render shadows in the image based on the lighting conditions at each location. In some embodiments, rays are projected from each location to sample random or pseudo-random locations at the light source 320. The ray tracer 102 may use any suitable ray tracing method, such as stochastic ray tracing. Examples of stochastic ray tracing techniques that may be used include those employing Monte Carlo or quasi-Monte Carlo sampling strategies. In the example shown, the ray tracer 102 (e.g., each thread) projects one ray per location and / or pixel in the ray tracing pass 314 for sampling. In other embodiments, a different number of rays may be projected per location or pixel, no rays may be projected for certain locations or pixels, and / or a different number of rays may be projected for different locations or pixels (e.g., by each thread). In the case where multiple rays are projected for a pixel or location, the value of the pixel or cell in the rendering 200 may correspond to the sum (e.g., average) of the ray tracing sample values for the pixel or location.

[0039] Although only the light source 320 is shown, the lighting conditions at each location may similarly be sampled relative to other light sources and / or objects in the virtual environment 300, which may be combined with the ray tracing samples obtained with respect to the light source 320, or may be used to generate additional renderings 200 (and masks 202) that may be filtered by the image filter 104 and provided to the image combiner 106. For example, the lighting conditions of different light sources may be determined and filtered separately (e.g., using the filtering techniques described with respect to Figure 1 ), and combined by the image combiner 106 (e.g., as another input to the image combiner 106).

[0040] As shown, some rays (such as ray 344) may interact with the light source 320, resulting in a ray tracing sample indicating that the light from the light source 320 can illuminate the corresponding location. In some embodiments, rays falling into this category may be assigned a visibility value of 1 to indicate that they are visible relative to the light source 320 (in Figure 2 and Figure 3(indicated by no shading in). Other rays (such as ray 340 and ray 342) can interact with the object, causing the ray tracing samples to indicate that the light from light source 320 is at least partially blocked and / or prevented from reaching these locations. An example of such an object is the light shield 322, which can block light from reaching the light source 320. In some embodiments, rays falling into this category can be assigned a visibility value of 0 to indicate that they are invisible relative to the light source 320 (in Figure 2 and Figure 3 (indicated by shading in). Since the visibility value can assume one of two potential values, it can correspond to a binary signal.

[0041] In at least one embodiment, a thread can determine the visibility of a corresponding location and can store the corresponding value for one or more pixels of the region 210 corresponding to the thread (e.g., according to the shader code). For example, each thread can determine the visibility value (e.g., 1 or 0) of a location / pixel and store it in a register. In at least one embodiment, each thread can have a dedicated register for storing one or more values. Figure 1 The rendering data 122 of

[0042] In Figure 3 example, the ray tracer 102 can determine that in region 210D, locations 330 and 332 are invisible to the light source 320 and all other locations are visible. Location 330 is an example of a location that can be within the penumbra of the shadow cast by the light shield 322, and the lighting conditions can be calculated more accurately by combining ray tracing samples from multiple rays. For example, the ray tracing sample of location 330 generated using only ray 342 can indicate that location 330 is completely blocked from receiving light from the light source 320. However, if a ray tracing sample of location 330 is generated using another ray, it can indicate that location 330 is at least partially illuminated by the light source 320, such that location 330 is within the penumbra.

[0043] Accordingly, restricting the number of rays used to generate samples for a location can result in noise, which in turn can result in visual artifacts in the data rendered by the image rendering system 100. The image filter 104 can be used to implement a denoising technique to reduce the noise. In different examples, the denoising technique can include: the image filter 104 filters the illumination condition data or other rendering data corresponding to the ray tracing samples from the ray tracer 102 spatially and / or temporally. For example, the image filter 104 can apply one or more spatial filter passes and / or temporal filter passes to the rendering data 122 from the ray tracer 102. According to the present disclosure, the image filter 104 can use the mask data 124 or otherwise use data corresponding to values generated by the thread from the values of the ray tracing samples as input to inform the denoising.

[0044] Return Figure 2 , a thread can store one or more values generated by the thread from the values of the ray tracing samples of at least one other thread in an area of the mask 202. The area of the mask 202 can include one or more pixels or cells of the mask 202. In this example, each area of the mask 202 is a single pixel or cell, but in other cases, different areas can include different numbers of pixels or cells and / or each area can include more than one pixel or cell. The mask 202 includes areas 212A, 212B, 212C, 212D, and 212E through 212N (collectively referred to as areas 212). In some examples, each area 210 of the rendering 200 is mapped to a single area 212 (e.g., via shader code executed by the thread). Accordingly, the mask 202 can include 64,800 pixels or cells. In other examples, an area 210 can be mapped to multiple areas 212 and / or an area 212 can correspond to multiple areas 210 (e.g., values from multiple areas 210 can be blended or otherwise aggregated by the thread to form a value in one or more areas 210).

[0045] For example, at least one thread of a dispatch unit corresponding to area 210A can store a value generated by the thread in area 212A of the mask 202, where the value is generated from the values of the threads in the dispatch unit. Additionally, a thread of a dispatch unit corresponding to area 210B can store a value generated by the threads in the dispatch unit in area 212B of the mask 202, where the value is generated from the values of the threads in the dispatch unit. Similarly, area 210C can correspond to area 212C, area 210D can correspond to area 212D, area 210E can correspond to area 212E, and area 210N can correspond to area 212N.

[0046] As described herein, a thread can compute a value for mask 202 (which may also be referred to as a mask value) based at least on ray tracing samples for each thread within a thread group and / or a dispatch unit. Typically, threads of a dispatch unit may only be able to access values of ray tracing samples generated by other threads within the dispatch unit. Thus, each thread group can be within the same dispatch unit. In the example shown, each dispatch unit includes a single thread group, and the thread group includes all the threads of the dispatch unit. At least one thread (e.g., each thread) can aggregate ray tracing samples of the dispatch unit, and at least one thread (e.g., one of the threads) can store the result in an area of mask 202. Thereby, mask 202 can have a lower resolution than the frame being rendered, which can reduce processing and storage requirements.

[0047] In other examples, a dispatch unit can be divided into multiple thread groups. In the case where the dispatch unit includes multiple thread groups, an area 212 of mask 202 can be provided for each thread group or the groups can share area 212. For example, area 210N of render 200 can include a group of 16 threads corresponding to a left hand side 4X4 pixel group and a group of 16 threads corresponding to a right hand side 4X4 pixel group. In this example, area 212N of mask 202 can alternatively include two adjacent areas - one area for each subgroup of area 212D. As an example, the same threads of a dispatch unit can compute and store values for two thread groups, or different threads can compute and store mask values for each thread group. For example, for each thread group, each thread within the group can compute a mask value, and at least one of those threads can store the value in mask 202 (e.g., in a buffer). Dividing the dispatch unit into multiple thread groups can be used to increase the resolution of mask 202.

[0048] In at least one embodiment, a thread can receive ray tracing samples and compute a value for mask 202 using one or more wave built-in functions. Wave built-in functions can refer to built-in functions available for use in code executed by one or more threads of a dispatch unit. Wave built-in functions can allow a thread to access values from another thread within the dispatch unit. Different wave built-in functions can be employed, which can depend on the format and / or the information desired to be captured by one or more values being computed for mask 202. As an example, a thread can perform wave activities and functions. Wave activities and functions can receive values of ray tracing samples (e.g., visibility values) from each thread of the dispatch unit (from registers), compute the sum of those values, and return the computed sum as a result.

[0049] In the example shown, the calculated value may indicate whether one or more pixels and / or positions of the virtual environment are within the penumbra. For example, the value of a ray tracing sample may be a visibility value, which is 0 or 1. For region 210 of rendering 200, the sum of the visibility values may be between 0 and 32. A value of 0 may indicate that the position corresponding to the region is completely within the shadow. Region 210E is an example of a region that may be indicated as being completely within the shadow. A value of 32 may indicate that the position corresponding to region 210 is fully illuminated. Regions 210A, 210B, and 210C are examples of regions that may be indicated as being fully illuminated. A value between 0 and 32 (the total number of threads in the group) may indicate that the respective positions corresponding to the region are in the penumbra. Regions 210D and 210N are examples of regions that may be indicated as being in the penumbra. Although examples are provided for the entire region 210, they may be similarly applied to subgroups or regions of region 210.

[0050] As described herein, mask 202 may be generated based at least on the value calculated from the ray tracing sample. Mask 202 may be referred to as a penumbra mask when the mask value of mask 202 indicates whether a pixel and / or a position in virtual environment 300 corresponds to the penumbra. Although these values (e.g., returned by wave built-in functions) may be used as the mask value of mask 202, in at least one embodiment, a thread uses the value to calculate a mask value and stores the mask value in mask 202. As an example, the mask value of mask 202 may be a binary value, and each binary value may indicate whether region 212 corresponds to the penumbra. In at least one embodiment, a value of 1 may indicate that region 212 corresponds to the penumbra, and a value of 0 may indicate that region 212 is outside the penumbra or does not correspond to the penumbra (being indicated as fully illuminated or fully within the shadow). Thus, using the example above, when the value calculated from the ray tracing sample is 0 or 32, the thread may calculate and store a mask value of 1. When the value calculated from the ray tracing sample is greater than 0 and less than 32, the thread may calculate and store a mask value of 0. Accordingly, regions 212D and 212N have a mask value of 0, and the other regions 212 have a value of 1. Although the example using 0 and 32 as thresholds is used, different thresholds may be used to determine the mask value and / or only one or the other of the thresholds may be employed.

[0051] Although in this example the one or more values calculated by a thread from ray tracing samples of at least one other thread are sums, other types of statistical values can be calculated. For example, in at least one embodiment, the calculated value is the variance of the values of the thread group. Further, more than one statistical value can be calculated and can be used to determine a mask value, or can be used for another purpose. For example, the sum can be used by the image filter 104 to determine whether to skip applying a denoising filter to a pixel, and the variance can be used to determine the radius of the denoising filter applied to the pixel and / or the range of values to be included in the denoising performed by the denoising filter. Statistical information about ray tracing samples can be useful for denoising any of a variety of different types of ray tracing samples, and the disclosed embodiments are not limited to shadow or visibility-related samples. For example, as described herein, ray tracing samples can have any condition, light condition, or other aspect of a virtual environment (e.g., hit distance, depth, etc.).

[0052] As described herein, the image filter 104 can use the mask data 124 or otherwise data corresponding to the values generated by a thread from ray tracing samples (e.g., not necessarily a mask) as input to inform denoising. The image filter 104 can use any of a variety of possible filtering techniques to filter the data. In some examples, the image filter 104 uses a cross (or joint) bilateral filter to perform the filtering. The cross bilateral filter can replace each pixel with a weighted average of nearby pixels using weights of a Gaussian distribution (which takes into account the distance, variance, and / or other differences between pixels) to guide the image. In at least one embodiment, this can involve the mask data 124 or otherwise data corresponding to the values generated by a thread from ray tracing samples analyzed by the image filter 104 to determine the filter weights. An edge stopping function can be used to identify common surfaces using G-buffer attributes to improve the robustness of the cross bilateral filter under input noise.

[0053] In at least one embodiment, the image filter 104 uses the mask data 124 to apply the denoising filter to a pixel early or skip it based on the mask value associated with the pixel. For example, the image filter 104 can skip applying the denoising filter to a pixel based at least on a mask value indicating that the pixel does not correspond to a penumbra. In at least one embodiment, when evaluating a pixel for denoising, the image filter 104 maps one or more values of each region 212 in the mask data to each pixel corresponding to the region 210 (or more generally, a thread group). For example, the value 0 of the region 212A from the mask 202 can be used for each pixel corresponding to the region 210A in the rendering 200. Based on the image filter 104 determining that a pixel is mapped to the value 1 or otherwise associated with the value 1, the image filter can apply the denoising filter to the pixel or can otherwise skip the pixel.

[0054] As described herein, the filter and / or one or more parameters passed by the filter applied by the image filter 104 can be determined based at least on the mask value and / or statistical value calculated by the thread. In at least one embodiment, the filtering of one or more pixels can be guided based at least in part on the value (e.g., variance) calculated by the thread of one or more pixels. For example, one or more parameters define the range of filter values, the filter weights of the pixels, and / or the filter radius of the filter and / or the filter pass. The image filter 104 can filter the spatial and / or temporal samples of the rendering data using one or more parameters. In different instances, the range of filter values can define a set of filter values for one or more pixels and can be based on the variance of one or more pixels. For example, when applying the filter and / or the filter pass to a pixel, the image filter 104 can exclude the set of filter values from the filtering based at least on the set of filter values being outside the range. In an embodiment, the range and / or the filter radius can increase and decrease with the variance.

[0055] Any of the various filters and / or filter passes described herein can be applied using a filter kernel. The filter can also have one or more filter directions. The filter kernel of the filter can refer to a matrix (e.g., a rectangular array) defining one or more convolutions for processing image data (and / or light condition data or rendering data) of an image (e.g., the data values of pixels) to change one or more characteristics of the image, such as the shading and / or color of the pixels of the image. In some examples, the filter kernel can be applied as a separable filter. When applying the filter as a separable filter, multiple sub-matrices or filters can be used to represent the matrix, and the sub-matrices or filters can be applied to the image data separately in multiple passes. When determining or calculating the filter core of the separable filter, the present disclosure anticipates that the sub-matrix can be directly calculated or can be derived from another matrix.

[0056] Each element of the matrix of the filter kernel may correspond to a respective pixel position. One of the pixel positions of the matrix may represent the initial pixel position corresponding to the pixel to which the filter is applied and may be located at the center of the matrix (e.g., for determining the position of the filter). For example, when applying a filter to a pixel corresponding to Figure 3 the pixel at position 332, the pixel may define the initial pixel position. When applying some filters, data values (e.g., visibility values) of other pixels may be used at image positions determined relative to the pixel to determine the data values of the pixels within the footprint of the filter kernel. The filter orientation may define the alignment of the matrix with respect to the image and / or the pixel, and the filter is applied along the filter width to the image and / or the pixel. Thus, when applying a filter to a pixel, the filter orientation and the filter kernel may be used relative to the initial pixel position to determine other pixels at other pixel positions of the matrix for the filter kernel.

[0057] Each element of the matrix of the filter kernel may include a filter weight for the pixel position. The matrix may be applied to an image using convolution, where the data value of each pixel of the image corresponding to the pixel position of the matrix is added to or otherwise combined with the data values of the pixels corresponding to the local neighbors in the matrix, weighted by the filter values (also referred to as filter weights). For one or more of the filters described herein, the filter values may be configured to blur the pixels, for example, by fitting a distribution to the filter kernel (e.g., fitting to the width and height).

[0058] The data values to which the filter is applied may correspond to the illumination condition data (e.g., visibility data) of the pixel. Thus, applying the matrix of the filter kernel to the pixel may cause the illumination condition data to be at least partially shared among the pixels corresponding to the pixel positions of the filter kernel. The sharing of the illumination condition data may reduce noise resulting from sparse sampling of the illumination conditions in ray tracing. In at least one embodiment, the image filter 104 performs spatial filtering of the rendering data. In some cases, the image filter 104 may also perform temporal filtering. Temporal filtering may utilize ray tracing samples, which may be generated similarly to those described with respect to Figure 3 but generated for previous states and / or previous output frames of the virtual environment 300. Thus, temporal filtering may increase the effective sample count of the ray tracing samples used to determine the filtered illumination condition data of the pixels and / or may increase the temporal stability of the filtered illumination condition data of the pixels.

[0059] Since the temporal ray tracing samples can correspond to different states of the virtual environment 300, some samples may be irrelevant or relevant to the current state of the virtual environment 300 (e.g., an object or camera may move, a light source may change), thus presenting a risk of visual artifacts when the visual artifacts are used for filtering. Some embodiments may use the variance of the values corresponding to the temporal ray tracing samples (e.g., computed by threads of a dispatchable unit according to Figure 2 ) to guide temporal filtering in order to reduce or eliminate these potential artifacts. According to a further aspect of the present disclosure, spatial filtering of pixels of a frame may be skipped based at least in part on determining that a mean, a first variance moment, and / or a variance computed by threads of a dispatchable unit from temporal ray tracing samples associated with a pixel is greater than or equal to a first threshold, and / or less than or equal to a second threshold, and a count of the values exceeds a third threshold.

[0060] The method may be used for any suitable ray tracing effect or technique, such as for global illumination, ambient occlusion, shadows, reflections, refractions, scattering phenomena, and dispersion phenomena. Thus, for example, although in some examples the ray tracing samples may correspond to visibility samples, in other examples the ray tracing samples may correspond to color luminance. Further, the method may be implemented in a rendering pipeline different from that shown in Figure 1 , which may or may not use an image combiner 106 to combine the output from the 3D renderer 108.

[0061] Now referring to Figures 4 - 6 , each block of methods 400, 500, and 600 and other methods described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in a memory. The methods may also be embodied as computer-usable instructions stored on a computer storage medium. These methods may be provided by a stand-alone application, a service, or a hosted service (independently or in combination with another hosted service) or a plug-in of another product, to name a few. Additionally, by way of example, these methods are described with respect to the image rendering system 100 ( Figure 1 ). However, these methods may alternatively or additionally be performed by any one system or any combination of systems, including but not limited to those described herein.

[0062] Figure 4It is a flowchart showing an example of method 400 for determining a visibility value and indicating a value corresponding to penumbra at a position in a scene using schedulable units according to some embodiments of the present disclosure. At block B402, method 400 includes: determining a first value corresponding to visibility using a schedulable unit. For example, ray tracer 102 may use threads of the schedulable unit corresponding to region 210D to determine a first value corresponding to the visibility of virtual environment 300 relative to at least light source 320, at least based on casting rays (such as rays 340, 342, and 344) in virtual environment 300. Figure 1 The rendering data 122 may represent the first value, or the first value may be used in other ways to derive the rendering data 122.

[0063] At block B404, method 400 includes: receiving, using at least one thread of the schedulable unit, a second value calculated from the first value, the second value indicating one or more positions corresponding to penumbra. For example, ray tracer 102 may use at least one thread of the schedulable unit to receive a second value calculated from the first value. For example, the thread may execute wave built-in functions to receive the first value from registers associated with one or more other threads of the schedulable unit and calculate the second value from one or more first values. The second value may indicate that one or more positions in virtual environment 300 correspond to penumbra.

[0064] At block B406, method 400 includes: applying a denoising filter to the first value using the second value based on determining that one or more positions correspond to penumbra using the second value. For example, Figure 1 The mask data 124 may represent the second value, or the second value may be used in other ways to derive the mask data 124. For example, the thread may determine a mask value (e.g., a binary value) from the second value and store the mask value as the mask data 124.

[0065] At block B408, method 400 includes: applying a denoising filter to the rendering data corresponding to the first value at least based on determining that one or more positions correspond to penumbra using the second value. For example, image filter 104 may apply a denoising filter to the rendering data 122 at least based on determining that one or more positions correspond to penumbra using the second value. Although method 400 is described with respect to schedulable units, method 400 may be executed using any number of parallel-operating schedulable units. Additionally, in some embodiments, all threads of the schedulable unit may be used for method 400 or a subset of the thread group or threads. Further, it is not necessary to receive the first value from all threads or thread groups of the schedulable unit.

[0066] Figure 5FIG. 500 is a flowchart illustrating an example of a method for using a thread group of one or more dispatch units to determine ray tracing samples for visibility and a value indicating whether a pixel corresponds to a penumbra, according to some embodiments of the present disclosure. At block B502, method 500 includes: determining ray tracing samples for visibility using one or more dispatch units. For example, ray tracer 102 may use a thread group of one or more dispatch units of one or more parallel processors to determine ray tracing samples for the visibility of pixels assigned to the group relative to at least light source 320 in virtual environment 300. In at least one embodiment, a thread group may refer to each thread of a dispatch unit or a subset of threads of a dispatch unit.

[0067] At block B504, method 500 includes: determining respective values for a thread group of one or more dispatch units, wherein at least one thread in the group calculates one of these values from the ray tracing samples of the group, and the value indicates whether the pixels of the group correspond to a penumbra. For example, ray tracer 102 may determine values for thread groups, such as each dispatch unit corresponding to region 210 and / or its sub-regions. For each of these thread groups, at least one thread in the group (e.g., each thread) may calculate one of these values from the ray tracing samples of the group. The value may indicate whether at least one pixel (e.g., each pixel) of the pixels assigned to the group corresponds to a penumbra.

[0068] At block B506, method 500 includes: denoising rendering data based at least on the value. For example, image filter 104 may denoise rendering data 122 corresponding to the ray tracing samples of the multiple thread groups based at least on the values of the groups. As an example, each thread in the group may determine a mask value (e.g., a binary value) from the value, and one or more threads of the threads in the group may store the mask value as mask data 124. Image filter 104 may use mask data 124 to denoise rendering data 122.

[0069] Figure 6 FIG. 600 is a flowchart illustrating an example of a method for using dispatch units to determine ray tracing samples and one or more values for determining one or more parameters of a denoising filter, according to some embodiments of the present disclosure. At block B602, method 600 includes: determining ray tracing samples of a scene using dispatch units. For example, ray tracer 102 may use threads of dispatch units of one or more parallel processors to determine ray tracing samples of virtual environment 300.

[0070] In block B604, the method includes: receiving, using at least one thread of a dispatchable unit, one or more values calculated from ray tracing samples. For example, ray tracer 102 may receive, using at least one thread of a dispatchable unit, one or more values calculated from ray tracing samples. For example, at least one thread may execute wave built-in functions to receive one or more values of ray tracing samples from registers associated with one or more other threads of the dispatchable unit and may calculate one or more values from the ray tracing samples. In some embodiments, the one or more values may indicate that one or more locations in virtual environment 300 correspond to penumbras. However, in other examples, the one or more values may indicate other information related to denoising.

[0071] In block B606, method 600 includes: determining, at least based on the one or more values, one or more parameters of a denoising filter. For example, Figure 1 The mask data 124 of... may represent the one or more values, or the one or more values may otherwise be used to derive the mask data 124. For example, a thread may determine a mask value (e.g., a binary value) from the one or more values and store the mask value as the mask data 124. In other examples, the one or more values may be used as one or more mask values. Image filter 104 may determine, at least based on the one or more values, one or more parameters of a denoising filter, e.g., by utilizing the mask data 124.

[0072] In block B608, method 600 includes: rendering a frame of a scene by applying a denoising filter to rendering data corresponding to ray tracing samples, at least based on using the one or more parameters. For example, image rendering system 100 may generate output image 120 by applying a denoising filter to rendering data 122 corresponding to ray tracing samples, at least based on using the one or more parameters. Although rendering data 122 is provided as an example, the denoising filter may be used to denoise other rendering data (e.g., rendering data corresponding to ray tracing samples other than visibility samples).

[0073] Example computing device

[0074] Figure 7 is a block diagram of an example computing device 700 suitable for implementing some embodiments of the present disclosure. Computing device 700 may include an interconnect system 702 that directly or indirectly couples the following devices: a memory 704, one or more central processing units (CPUs) 706, one or more graphics processing units (GPUs) 708, a communication interface 710, input / output (I / O) ports 712, input / output components 714, a power supply 716, one or more presentation components 718 (e.g., one or more displays), and one or more logic units 720.

[0075] Although Figure 7 each of the blocks of Figure 7 is shown as being connected via an interconnect system 702 having circuitry, this is not intended to be limiting and is for clarity only. For example, in some embodiments, a rendering component 718 such as a display device may be considered an I / O component 714 (e.g., if the display is a touchscreen). As another example, the CPU 706 and / or GPU 708 may include memory (e.g., memory 704 may represent a storage device other than the memory of the GPU 708, CPU 706, and / or other components). In other words, Figure 7 the computing device of Figure 7 is merely illustrative. No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types, as all of these are considered within the scope of the computing device of Figure 7 Figure 7 .

[0076] The interconnect system 702 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 702 may include one or more types of buses or links, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 706 may be directly connected to the memory 704. Further, the CPU 706 may be directly connected to the GPU 708. In cases where there are direct or point-to-point connections between components, the interconnect system 702 may include a PCIe link to effect the connection. In these examples, a PCI bus need not be included in the computing device 700.

[0077] The memory 704 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 700. Computer-readable media can include volatile and nonvolatile media as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.

[0078] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 704 can store computer-readable instructions (e.g., which represent programs and / or program elements such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage devices, magnetic tape cartridges, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by computing device 700. As used herein, computer storage media does not include signals per se.

[0079] Computer storage media can include computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information conveyance medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, computer storage media can include wired media such as a wired network or direct wired connection, and wireless media such as sound, RF, infrared, and other wireless media. Any of the foregoing combinations should also be included within the scope of computer-readable media.

[0080] CPU 706 can be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 700 to perform one or more of the methods and / or processes described herein. Each of the CPUs 706 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of simultaneously processing a large number of software threads. CPU 706 can include any type of processor and can include different types of processors depending on the type of computing device 700 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 700, the processor can be an advanced RISC machine (ARM) processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as a math coprocessor, computing device 700 can also include one or more CPUs 706.

[0081] In addition to or instead of the CPU 706, the GPU 708 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 700 to perform one or more of the methods and / or processes described herein. One or more of the GPUs 708 may be an integrated GPU (e.g., integrated with one or more of the CPUs 706 and / or one or more of the GPUs 708 may be a discrete GPU). In an embodiment, one or more of the GPUs 708 may be a coprocessor of one or more of the CPUs 706. The GPU 708 may be used by the computing device 700 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 708 may be used for general-purpose computing on the GPU (GPGPU). The GPU 708 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 708 may generate pixel data for an output image in response to a rendering command (e.g., a rendering command received from the CPU 706 via a host interface). The GPU 708 may include graphics memory, such as display memory, for storing pixel data or any other suitable data (such as GPGPU data). The display memory may be included as part of the memory 704. The GPU 708 may include two or more GPUs operating in parallel (e.g., via a link). The link may directly connect the GPUs (e.g., using NVLINK) or may connect the GPUs through a switch (e.g., using NVSwitch). When combined, each GPU 708 may generate pixel data or GPGPU data for a different part of the output or for a different output (e.g., the first GPU for a first image and the second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.

[0082] In addition to or instead of CPU 706 and / or GPU 708, one or more logic units 720 may be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 700 to perform one or more of the methods and / or processes described herein. In an embodiment, one or more CPUs 706, one or more GPUs 708, and / or one or more logic units 720 may perform any combination of methods, processes, and / or portions thereof discretely or jointly. One or more of the logic units 720 may be one or more of the CPUs 706 and / or GPUs 708 and / or integrated in one or more of the CPUs 706 and / or GPUs 708, and / or one or more of the logic units 720 may be discrete components or otherwise external to the CPUs 706 and / or GPUs 708. In an embodiment, one or more of the logic units 720 may be a coprocessor of one or more CPUs in the CPU 706 and / or one or more GPUs in the GPU 708.

[0083] Examples of logic units 720 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel vision cores (PVC), vision processing units (VPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic logic units (ALU), application specific integrated circuits (ASIC), floating point units (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.

[0084] Communication interface 710 may include one or more receivers, transmitters, and / or transceivers that enable computing device 700 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. Communication interface 710 may include components and functionality that enable communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., over Ethernet or InfiniBand communication), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, communication interface 710 may also include one or more processing cores and / or their components (such as data processing units (DPU)) to directly access data and store the data in the local memory of other processing units (such as GPUs) of computing device 700.

[0085] The I / O port 712 can enable the computing device 700 to be logically coupled to other devices including I / O components 714, presentation components 718, and / or other components, some of which may be built into (e.g., integrated into) the computing device 700. Illustrative I / O components 714 include microphones, mice, keyboards, joysticks, game pads, game controllers, dish satellite antennas, scanners, printers, wireless devices, and so on. The I / O components 714 can provide a natural user interface (NUI) for processing user-generated air gestures, voice, or other physiological inputs. In some examples, the input can be transmitted to appropriate network elements for further processing. The NUI can implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of the computing device 700 (described in more detail below). The computing device 700 can include depth cameras such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technologies, and combinations thereof for gesture detection and recognition. Additionally, the computing device 700 can include an accelerometer or gyroscope enabling motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 700 to render immersive augmented reality or virtual reality.

[0086] The power supply 716 can include hardwired power, battery power, or a combination thereof. The power supply 716 can power the computing device 700 to enable the components of the computing device 700 to operate.

[0087] The presentation component 718 can include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 718 can receive data from other components (e.g., GPU 708, CPU 706, etc.) and output the data (e.g., as images, videos, sounds, etc.).

[0088] Example network environment

[0089] A network environment suitable for use in implementing embodiments of the present disclosure can include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) can be implemented on one or more instances of Figure 7 one or more computing devices 700 - for example, each device can include similar components, features, and / or functions of one or more computing devices 700.

[0090] Components of a network environment can communicate with each other via one or more networks, which can be wired, wireless, or both. The network can include multiple networks or one of multiple networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or the public switched telephone network (PSTN) and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0091] A compatible network environment can include one or more peer-to-peer network environments (in which case, servers may not be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functions described herein with respect to servers can be implemented on any number of client devices.

[0092] In at least one embodiment, the network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment can include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework that supports software layers and / or one or more applications of an application layer. The software or application can respectively include network-based service software or applications. In an embodiment, one or more client devices can use network-based service software or applications (e.g., by accessing service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, but is not limited to, the type of free and open-source software web application framework that can perform large-scale data processing (e.g., "big data") using a distributed file system.

[0093] A cloud-based network environment can provide any combination of cloud computing and / or cloud storage for performing the computing and / or data storage functions (or one or more parts thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers located in a state, region, country, the globe, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, the core server can assign at least a portion of the function to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0094] The client device may include at least some of the components, features, and functions of one or more of the example computing devices 700 described herein. By way of example and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, camera, surveillance device or system, vehicle, boat, spacecraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronic device, workstation, edge device, any combination of these depicted devices, or any other suitable device. Figure 7

[0095] Example data center

[0096] Figure 8 An example data center 800 that may be used in at least one embodiment is shown. In at least one embodiment, the data center 800 may include a data center infrastructure layer 810, a framework layer 820, a software layer 830, and / or an application layer 840.

[0097] In at least one embodiment, as Figure 8 shown, the data center infrastructure layer 810 may include a resource coordinator 812, grouped computing resources 814, and node computing resources (“node C.R.”) 816(1)-816(N), where “N” represents any whole positive integer. In at least one embodiment, the node C.R. 816(1)-816(N) may include, but is not limited to, any number of central processing units (CPUs) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drives or disk drives), network input / output (NW I / O) devices, network switches, virtual machines (VMs), power modules, and / or cooling modules, etc. In at least one embodiment, one or more of the node C.R. 816(1)-816(N) may be a server having one or more of the above computing resources.

[0098] In at least one embodiment, the grouped computing resources 814 can include separate groupings (not shown) of node C.R.s housed within one or more racks, or numerous racks (also not shown) housed within data centers at various geographical locations. Separate groupings of node C.R.s within the grouped computing resources 814 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R.s that include CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches, in any combination.

[0099] In at least one embodiment, the resource coordinator 822 can configure or otherwise control one or more node C.R.s 816(1)-816(N) and / or the grouped computing resources 814. In at least one embodiment, the resource coordinator 822 can include a software design infrastructure (“SDI”) management entity for the data center 800. In at least one embodiment, the resource coordinator can include hardware, software, or some combination thereof.

[0100] In at least one embodiment, as Figure 8As shown, the framework layer 820 may include a job scheduler 844, a configuration manager 834, a resource manager 836, and a distributed file system 838. In at least one embodiment, the framework layer 820 may include a framework for software 832 that supports the software layer 830 and / or one or more applications 842 of the application layer 940. In at least one embodiment, the software 832 or the application 842 may respectively include network-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 820 may be, but is not limited to, a free and open-source software web application framework, such as Apache SparkTM (hereinafter referred to as "Spark") that can utilize the distributed file system 838 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 844 may include a Spark driver to facilitate scheduling of workloads supported by the various layers of the data center 800. In at least one embodiment, the configuration manager 834 may be able to configure different layers, such as the software layer 830 and the framework layer 820 including Spark and the distributed file system 838 for supporting large-scale data processing. In at least one embodiment, the resource manager 836 is capable of managing cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 838 and the job scheduler 844. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 814 on the data center infrastructure layer 810. In at least one embodiment, the resource manager 836 may coordinate with the resource coordinator 812 to manage these mapped or allocated computing resources.

[0101] In at least one embodiment, the software 832 included in the software layer 830 may include software used by at least a portion of the nodes C.R. 816(1)-816(N), the grouped computing resources 814, and / or the distributed file system 838 of the framework layer 820. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.

[0102] In at least one embodiment, one or more applications 842 included in the application layer 840 may include one or more types of applications used by at least a portion of nodes C.R. 816(1)-816(N), grouped computing resources 814, and / or the distributed file system 838 of the framework layer 820. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0103] In at least one embodiment, any one of the configuration manager 834, the resource manager 836, and the resource coordinator 812 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions may relieve the data center operator of the data center 800 from making potentially bad configuration decisions and may avoid underutilization and / or poorly performing parts of the data center.

[0104] In at least one embodiment, the data center 800 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by computing weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 800. In at least one embodiment, by using the weight parameters calculated by one or more training techniques, the resources described above in connection with the data center 800 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks.

[0105] The present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be practiced in various system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network.

[0106] As used herein, the recitation of "and / or" with respect to two or more elements shall be construed to mean only one element or a combination of elements. For example, "element A, element B, and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. Further, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of at least one of element A and element B. Still further, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of at least one of element A and element B.

[0107] The subject matter of the present disclosure is specifically described herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Rather, the inventors have contemplated that the claimed subject matter may be embodied in other ways, in combination with other current or future technologies, to include different steps or combinations of steps similar to the steps described herein. Additionally, although the terms "step" and / or "block" may be used herein to imply different elements of a method employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and except when the order of individual steps is explicitly described.

Claims

1. A computer-implemented method, comprising: Determining a first value indicative of visibility of a region of a frame of the scene for at least one light source in the scene, based at least on casting one or more rays from at least two positions in the scene, the at least two positions corresponding to the region of the frame; Receiving a second value corresponding to a lower-resolution version of the region of the frame, the second value being calculated for a set of positions among the at least two positions based on data representative of the first value, wherein the second value indicates that the region of the frame corresponds to penumbra; And Applying a denoising filter to rendering data, based at least on determining that the region of the frame corresponds to the penumbra using the lower-resolution version of the region.

2. The method according to claim 1, wherein the one or more rays include one or more first rays for sampling the visibility of one or more first pixels among at least two pixels corresponding to the region, and one or more second rays corresponding to at least a second pixel of the at least two pixels.

3. The method according to claim 1, wherein at least one thread for determining the first value and receiving the second value corresponds to a respective position among the set of positions, and each position among the set of positions corresponds to a respective pixel in the pixel region.

4. The method according to claim 1, wherein the receiving of the second value is received from an output of a wave built-in function using one or more parallel processors, the wave built-in function being called by a thread among one or more threads and accessing data from one or more registers among the one or more threads to generate the second value.

5. The method according to claim 1, further comprising: Generating a penumbra mask having a lower resolution than a frame of the scene using the second value, wherein the penumbra mask includes the lower-resolution version of the region of the frame.

6. The method according to claim 1, wherein the lower-resolution version of the region is represented using fewer pixel values than the region of the frame.

7. The method according to claim 1, wherein determining the first value, receiving the second value, and determining that the region corresponds to the penumbra are performed in one or more ray-tracing passes, and applying the denoising filter is performed in a denoising pass operating on image data produced by the one or more ray-tracing passes.

8. The method according to claim 1, wherein determining the first value and receiving the second value are performed by a ray generation shader executed by one or more threads using one or more parallel processors.

9. The method according to claim 1, wherein determining that the region corresponds to the penumbra is based at least on: comparing the second value with a threshold.

10. The method according to claim 1, wherein the second value includes statistical data regarding the visibility of the region.

11. A computer-implemented method, comprising: Ray tracing samples for determining the visibility of a region of the frame for at least one light source in the scene are determined based at least on sampling at least two positions in the scene depicted by the frame, the at least two positions corresponding to the region of the frame; Based on data of the ray tracing samples representing the region, a value corresponding to a lower resolution version of the region of the frame is determined, the value indicating whether the region of the frame corresponds to a penumbra; And Based at least on determining that the region corresponds to the penumbra using the lower resolution version of the region, denoising is performed on rendering data corresponding to the ray tracing samples.

12. The method according to claim 11, wherein a first set of threads is used to determine that the ray tracing samples belong to a first schedulable unit of one or more parallel processors, and a second set of threads belongs to a second schedulable unit of the one or more parallel processors.

13. The method according to claim 11, wherein said denoising comprises: Apply a denoising filter to pixels of the region using the lower resolution version of the region, the pixels being assigned to a first set of threads of one or more parallel processors.

14. The method according to claim 11, wherein the denoising comprises: Apply a denoising pass to the ray tracing samples using the lower resolution version of the region, wherein the denoising pass skips pixels of the region assigned to a second set of threads of one or more parallel processors based at least on the value indicating that the region is outside the penumbra.

15. The method according to claim 11, wherein a plurality of threads of one or more parallel processors determine the value based on the ray tracing samples.

16. The method according to claim 11, wherein the lower resolution version of the region is included in a penumbra mask of a frame of the scene, and the denoising is based at least on analyzing the penumbra mask.

17. A processor, comprising: One or more circuits for: determining ray tracing samples corresponding to a region of a frame of the scene based at least on projecting one or more rays from at least two positions in the scene, the at least two positions corresponding to the region of the frame; receiving one or more values corresponding to a lower resolution version of the region of the frame, wherein a value among the one or more values is calculated for a set of the positions based on data representing the ray tracing samples; determining one or more parameters of a denoising filter based at least on determining that the region corresponds to a penumbra using the lower resolution version of the region of the frame; And generating the frame of the scene based at least on applying the denoising filter to rendering data corresponding to the ray tracing samples using the one or more parameters.

18. The processor according to claim 17, wherein the one or more parameters define a filter radius of the denoising filter.

19. The processor according to claim 17, wherein the one or more parameters define a range, and values outside the range are excluded from filtering using the denoising filter based on the values being outside the range.

20. The processor according to claim 17, wherein the processor is included in at least one of the following: A system for performing analog operations; A system for performing analog operations to test or verify autonomous machine applications; A system for performing deep learning operations; A system implemented using edge devices; A system incorporating one or more virtual machines (VMs); A system implemented at least partially in a data center; or A system implemented at least partially using cloud computing resources.

Citation Information

Patent Citations

  • Subsurface illumination analysis, a hybrid wave equation-ray-tracing method

    CN1682234A

  • Synthetic image and video generation from ground truth data

    US20080175507A1

  • Processor and method for accelerating ray casting

    US20160350960A1

  • Shadow denoising in ray-tracing applications

    US20190287291A1