Light Importance Cache Using Spatial Hashing in Real-Time Ray Tracing Applications
By determining and caching importance information for scene regions, the method optimizes ray tracing to improve image quality and efficiency in rendering complex scenes with many light sources, addressing resource constraints and maintaining high frame rates.
Patent Information
- Application Number
- CN202110901023.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-08
- Filing Date
- 2021-08-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-08-06
AI Technical Summary
The prior art renders in complex scenes, especially in high frame rate applications, insufficient resources lead to the inability to accurately represent all lighting, affecting image quality.
By determining the importance of light sources in the scene and performing sampling optimization, using the spatial hashed light importance cache method, only important areas are sampled, combined with path spatial filtering and photon mapping technology, the rendering efficiency and quality are improved.
Significantly improves the bandwidth and cache efficiency of image generation, while maintaining high image quality, reducing resource requirements, and is suitable for real-time and offline rendering systems.
Smart Images

Figure CN114663572B_ABST
Abstract
Description
Background
[0001] As the quality of display devices and user expectations continue to increase, it is necessary to continuously improve the quality of the displayed content. This can include tasks such as realistically lighting the objects in the scene to be rendered. In complex scenes with many light sources (including reflective or refractive surfaces), there may not be sufficient resource capacity to accurately represent all the lighting in the scene, especially for applications that require high frame rates.
[0002] Brief Description of the Drawings
[0003] Described with reference to the accompanying drawings, according to various embodiments of the present disclosure, wherein:
[0004] Figure 1A 、 Figure 1B and Figure 1C show ray tracing of an object in a scene according to at least one embodiment;
[0005] Figure 2 show ray tracing for both view and shadow rays according to at least one embodiment;
[0006] Figure 3A and Figure 3B show ray tracing with octahedral voxels according to at least one embodiment;
[0007] Figure 4 show an example process for rendering a frame of a scene according to at least one embodiment;
[0008] Figure 5 show a process for performing ray tracing during rendering according to at least one embodiment;
[0009] Figure 6 show an example image generation system including a light importance cache according to at least one embodiment;
[0010] Figure 7 show an example data center system according to at least one embodiment;
[0011] Figure 8 show a computer system according to at least one embodiment;
[0012] Figure 9 show a computer system according to at least one embodiment;
[0013] Figure 10 show at least a portion of a graphics processor according to one or more embodiments; and
[0014] Figure 11Shows at least a portion of a graphics processor in accordance with one or more embodiments. Detailed description
[0015] Methods according to various embodiments can overcome various deficiencies in existing image generation methods. In particular, various embodiments can provide for the determination and caching of important information for one or more light sources in a scene. The importance information can be determined for various regions of the scene, which may be associated with respective voxels and can include directional data. The importance information can be used to determine the amount of sampling to be performed for these different regions, at least in part based on factors such as the number of important lights and the importance determined for those lights. This approach can significantly improve bandwidth and cache efficiency, while providing high image quality, as compared to sampling all a large number of light sources or approximating such sampling for all such light sources.
[0016] In various real-time and offline rendering systems, a series of images or video frames can be generated using a virtual camera view of a three-dimensional scene. The scene can include various types of static or dynamic objects, as well as one or more light sources for illuminating these objects. The objects can include any type of foreground or background object that can be visually represented in an image or other digital view. The light sources can also be any suitable light source or illumination source that can affect one or more of these scene objects, where the light can have different colors or intensities. At least some of the objects in these scenes can also be reflective or refractive, such that they can affect the illumination from one or more of these light sources, or in some cases can act as separate light sources.
[0017] Figure 1AShows an example view 100 that can be generated for the scenario according to various embodiments. In this example, the view 100 is a view of a virtual camera relative to a virtual three-dimensional (3D) environment, so as to obtain a two-dimensional (2D) representation of the environment from a determined viewpoint. In this particular view or image frame, there are three objects 108, 110, 112, which are referred to herein as scene objects. These scene objects can include any object that can be represented in the scene, such as background or foreground objects, player avatars in a game, dynamic or static objects, etc. These objects can have different lighting characteristics, such as the lighting characteristics can include that at least a part of the object is opaque, transparent, reflective, transmissive, or refractive, and other such options. The objects can also have different surface characteristics or textures, and these characteristics or textures may affect their lighting characteristics. There are also three light sources 102, 104, 106 in this scene. These light sources can be any object or element that can generate a certain type of lighting, such as light sources including the sun, lamps, neon lights, etc. In order to render this view of the scene in a realistic manner, the light from the various light sources 102, 104, 106 should illuminate the scene objects 108, 110, 112, similar to the way of illuminating those scene objects in a real-world setting. This includes not only direct lighting, but also aspects such as reflection and shadow creation.
[0018] One method of determining such lighting involves ray tracing. Ray tracing refers to a rendering technique that uses an algorithm to trace the path of light emitted by a light source, and then simulates the way light interacts with the scene objects that the ray "hits" or intersects in a computer-generated environment. Ray tracing can provide realistic shadows and reflections, as well as accurate translucency and scattering. Ray tracing also has the advantage that it can be executed on a graphics processing unit (GPU), such as being executed in a GeForce RTX graphics card produced by NVIDIA Corporation, so as to execute lighting determination in sufficient time, so that accurate rendering can be performed in real time for high frame rate applications.
[0019] Figure 1B As shown, the view 130 of the scene shows example rays from two light sources 102 and 106. Although at least one ray can be projected from each potential light source for each potential pixel position in the scene, this method can be very resource-consuming. Therefore, ray tracing methods usually select or "sample" some of these rays for each frame to determine approximate updated lighting information based on these samples and (at least in some cases) previous lighting information from previous frames. Figure 1BAn example selection of rays from light sources 102 and 106 is shown. As shown, rays from each of these light sources impinge on different objects or points of impact on portions of those objects, the points of impact being at least partially based on the relative positions of the light sources and objects in the scene.
[0020] Methods according to various embodiments can utilize this information to allow for more optimized sampling of rays for a frame or scene. For example, rays from light source 102 impinge on all three objects in this scene, but do not impinge on the right side of scene object 112. Additionally, scene object 112 blocks rays from light source 102 from impinging on most of the frame to the right of scene object 112. Rays from light source 102 primarily illuminate the top and right side of scene object 108 (reference is retained between figures for convenience), possibly all of scene object 110 as well as the top and left side of scene object 112. Then an area 132 of the scene can be determined in which light source 102 has a relatively high importance (greater contribution), as this area includes surfaces of objects that can be impinged upon by rays from light source 102. Similarly, rays from light source 106 impinge on the right side of scene object 112, but scene object 112 blocks these rays from impinging on scene objects 108 and 110. Thus, an area 134 of the scene can be defined in which light source 106 has relative importance. It can be seen that light source 102 contributes less to the illumination in area 134, as rays from this light source are mostly blocked by objects in area 134. Similarly, light source 106 contributes less to area 132, as rays from this light source are also mostly blocked by objects in area 132. When determining which rays to sample for a frame or a portion of a frame, methods according to various embodiments can utilize this importance information, as it may be more valuable to sample only rays from light sources that have a more significant contribution to a given portion or area.
[0021] In at least one embodiment, rays can be sampled only for those areas of light sources that are determined to be at least somewhat important. In some embodiments, this can include up to a maximum number of important light sources. In other embodiments, even for light sources for which little contribution to an area is determined, one or more rays can be sampled per frame to account for changes in the scene. For example, movement of an object can allow a ray to pass through the area, or can reflect a ray into the area, etc. Additionally, sampling can be used to determine contribution determination, so that not every potential ray needs to be analyzed, and thus incomplete data can be used to determine the contribution calculation, and a light source may actually have some undetected importance to the area. This method allows for determination of such lighting effects without dedicating a large amount of resources to light sources that are typically unimportant to an area or portion of the scene.
[0022] Figure 1CThe example view 160 shown depicts sampled rays projected from a light source 104. In this example, rays from the light source 104 directly illuminate portions of the scene objects 108 and 112. As shown, the ray 162 reflected from the scene object 108 also impinges on the scene object 110. In this case, the light source 104 can contribute to each region of the scene and, as such, can sample rays for all regions. In at least some embodiments, the contributions determined for different regions can vary such that the amount or frequency of ray sampling by the light source 104 can vary with the region.
[0023] In at least one embodiment, different types of rays can also be sampled for a given scene. As Figure 2 shown in the view 200, there is a single light source 204 and a single scene object 208 in this scene. A two-dimensional image 206 of the three-dimensional scene is rendered from the viewpoint of the virtual camera 202. In at least some embodiments, the virtual camera 202 can be moved or redirected around the scene for different frames. The image will correspond to a plane that intersects the virtual camera view at a determined location. To render the image 206, the renderer must determine which portions of the scene object 208 are visible from the camera view in order to determine what to represent in the image. Additionally, the renderer must determine the illumination of each such portion in order to realistically render the object relative to the light source. As shown, this can include determining at least two types of rays. A first type of ray (referred to herein as a view ray) corresponds to a visible point or region on the scene object 208 in the virtual view and is illuminated (directly or indirectly) by the light source 204. A second type of ray (referred to herein as a shadow ray) corresponds to a point or region in the scene, such as on the ground or floor in the example environment, onto which an object such as the scene object 208 will cast a shadow at least in part based on light projected from the light source 204. In this case, the shadow ray 212 from the light source 204 to the scene object 208 will be occluded by the scene object 208 and thus will not impinge on the shadow region. However, the ray can be extended to determine the shadow that can be cast by the scene object 208 at that location relative to the light source 204. Thus, light rays can be determined not only for direct illumination but also for shadows, reflections, refractions, diffusions, and other such aspects of light. These different types of rays can be sampled at similar or different sampling rates or frequencies, and in some embodiments, can be sampled at least in part based on their relative importance to the scene or region.
[0024] In a three-dimensional grid, a virtual three-dimensional environment can be composed of multiple individual voxels. In various rendering systems, each voxel will include a virtual cube, triangle, octahedron, or other such shape or geometric object associated with a specific part of the scene. When ray tracing, rays can be emitted for a single voxel to determine any object (e.g., a triangle mesh or higher-order primitive) struck by the ray, and intersection information can be stored for rendering. However, in various situations, it may also be desirable to capture the directional data of the voxel. For example, there may be a material property of the voxel such that a light source from a first direction may be important (contribute more), but a different light source from a second direction may also be important. In such cases, it may be desirable to utilize voxels that can capture at least some of this dimension. In at least one embodiment, octahedral voxels can be used for objects in the scene, as Figure 3A shown in view 300 of Figure 3A . Other geometries can also be used, such as tetrahedral voxels in at least some embodiments. As shown, rays from light source 302 can strike at least two different faces of octahedral voxel 304, and the at least two different faces are the faces visible in the virtual camera view. Also as shown in the figure, there is at least one face of the voxel 304 that the light source 302 does not strike. This directional information helps better determine the importance of the light source 302 relative to the voxel. In at least some embodiments, the octahedral geometry can be selected because it provides a simple and effective way to determine the contributions of various light sources at different positions or directions in the scene.
[0025] Figure 3B An example view 350 of
[0026] shows an octahedral voxel 352 having four side faces or surfaces (A, B, C, D) visible in the virtual camera view. This may be a voxel representing a thin wall or a part of an object, for example, having a first light source 354 on one side and a second light source 356 on the other side. As shown in the figure, the first light source 354 is important for two of the surfaces (A, C), while the second light source 356 is important for the other two surfaces (B, D). In such cases, the importance information can be determined at least in part based on directionality, where it can be determined that each of the light sources is important, but only for the corresponding range of directions or number of faces. In at least one embodiment, the rendering interface can provide the user with an option to turn on or off the octahedral voxel representation of the 3D environment.Each of these voxels can correspond to a respective spatial hash. Different rays can strike different parts of the same spatial hash. Octahedral voxels are used to represent the hash to provide a relatively simple way to determine and store directional light information for the same spatial hash. This method enables shadow and view (or proximity) rays to be traced or projected from different light sources and directions into these voxels and the information to be stored for rendering. The information can also help determine the relative importance and / or contribution of each light source to a given voxel, set of voxels, or scene or image region. However, it should be noted that the spatial hash (figure) with octahedra allows positions and directions in the scene to be associated with positions in memory, where the memory can store light contribution information. However, as previously mentioned, sampling all possible rays of a scene can be too inefficient or resource intensive in many cases, especially for non-point light sources, and thus ray sampling can be utilized. Instead of randomly selecting rays or using pseudo-random sampling, other sampling methods or algorithms can be used, such as those based on quasi-Monte Carlo methods. Such methods can provide low-discrepancy sequences that can be converted into light paths. This method can utilize deterministic, low-discrepancy sequences to generate path segments. To generate a light path segment, the components of the i-th vector of the low-discrepancy sequence can be divided into at least two groups and then used to trace the i-th camera and light path segment.
[0027] According to at least one embodiment, a light path segment can be determined. The method for determining the light path segment first selects an origin (e.g., a light source) and then selects a direction for tracing a ray. At the first point of intersection with a scene surface, another decision can be made with respect to the end point of the path and the scattering direction (if appropriate) for tracing the next ray. This process can be repeated as necessary. An example method for determining a light path utilizes shadow rays and determines whether the end points of the respective path segments of the shadow rays are visible. While shadow rays are sufficient for most diffuse surfaces, shadow rays may be ineffective for light paths that contain specular-reflection-diffuse-reflection-specular segments (e.g., light reflected by a mirror onto a diffuse surface and reflected back by the mirror). To address this inefficiency, photon trajectories can be connected to camera path segments through proximity, also known as photon mapping, which helps capture the effective contribution of shadow rays. Using this path space filtering method, fragments of various light paths can be generated by tracing light rays (or photon trajectories) from a selected light source and tracing the path of a virtual camera. If the end points of these path segments are mutually visible or close enough (e.g., have sufficient proximity) for shadow rays, then the end points of these path segments are connected. In at least one embodiment, path space filtering can be utilized in conjunction with these connection techniques. For the vertex x of the light path i the contribution value c i can be replaced with a smoothed contribution instead, and this smoothed contribution from the average contribution c of vertices within a defined region si+j is obtained. Then this average contribution value can be multiplied by the throughput τ of the path segment towards the camera i and accumulated on the image plane P.
[0028] To ensure consistency, the size of the region should vanish as the number of samples n increases. In at least one embodiment, this can include reducing the radius r of the spherical region. Then the endpoints of the path segments can be connected using photon mapping, where the spacing of the path segment endpoints is less than the specified radius. The radius r(n) decreases as the number of sampled light paths n increases, which can provide a certain degree of consistency as in the limit it is effectively equivalent to shadow ray connection. Similar to stochastic progressive photon mapping, successive batches of light transport paths can be processed. Depending on the low-discrepancy sequence used, certain block sizes may be more preferable than others.
[0029] This method is in fact beneficial at least in the following aspects: Progressive path space filtering is a fast and consistent variance reduction technique that complements shadow rays and progressive photon mapping, enabling sparse sample sets to provide sufficient information for high-quality image generation. Although various methods can be used to select various multiple vertices of the light path, one option is the first vertex along the path towards the camera, where the optical properties of the camera are considered to be sufficiently scattering. A low-discrepancy sequence can be transformed into the sampled path space in successive batches of light transport paths, where for each path, a selected tuple is stored for path space filtering. Since the memory consumption is proportional to the batch size, and given the size of the tuple and the maximum size of the memory block, the maximum natural number can be determined relatively straightforwardly.
[0030] In at least one embodiment, the light data can be temporarily stored or cached for generating an image of the scene. In at least one embodiment, samples of the radiation can be cached and interpolated to increase the efficiency of the light transport simulation implemented in the renderer. In addition to radiation interpolation, path space filtering can also effectively filter discontinuities, such as detailed geometries. This filtering can also overcome excessive ray splitting to reduce noise in the cached samples, thus enabling effective path tracing. Furthermore, consistency-related artifacts present in the frame can be reduced through this calculation without separate computation. The averaging process can be iterated within a batch of light transport paths, thereby further significantly increasing the speed at the cost of some blurred lighting details.
[0031] Although this method can be consistent even without weighting, for larger radii, since the contributions are included in regions that cannot be resolved in the x iIn the average value collected, so the resulting image may appear overly blurred. To reduce this transient artifact of light leakage and benefit from a larger radius to include more contributions in the average value, the weights can be determined to evaluate that through trajectory segmentation, contributions c can be created on vertex xi si+j probability. Blurring can be performed across geometries or across textures using blurring methods. When blurring across geometries, their similarity can be determined by the scalar product of the surface normal and other surface normals. When blurring across textures, if the optical surface properties are evaluated by the contributions of vertices included in the average value, the image may be optimal. For surfaces other than diffuse surfaces (e.g., glossy surfaces), these properties can depend at least in part on the viewing direction, and then the viewing direction can be stored together with the vertices. When the directions are implicitly known to be similar (e.g., for a query position x directly viewed from the camera i ), some of this additional memory can be saved. For blurred shadows, given a point light source, the visibility seen from different vertices can be the same or different for one or more. To avoid clear shadow boundaries from becoming blurred, contributions can be included only when the visibility is the same. For ambient occlusion and illumination caused by environment mapping, blurring can be reduced by comparing the lengths of each ray entering the hemisphere at these vertices (by limiting their differences).
[0032] This path space filtering method can be implemented on top of any sampling-based rendering algorithm and can be implemented with little additional overhead. Progressive algorithms can effectively reduce variance and can guarantee convergence without persistent artifacts due to consistency. This spatial filtering or spatial hashing method can be used to divide the screen into voxels, for example, refer to Figure 3A and 3BThe described octahedral voxels. This method is suitable for applications such as real-time ray tracing and can be performed in at least some embodiments only in one or more GPUs, all of which simultaneously address various deficiencies of prior art methods. This method is also advantageous at least to some extent compared to methods already existing in various ray tracers, as it does not require additional functionality for ray sampling or tracing. Instead, the process can collect statistical information about rays projected onto the scene rays to determine the importance of each light source for a particular part of the scene and determine how many samples to perform in that part of the scene based on factors such as the visibility and material properties of the scene. This method can potentially result in far fewer samples being required than previous methods (e.g., ReSTIR), especially for world space sampling rather than screen space sampling. The need for many additional samples in existing methods often results in approximations being performed, which degrades the image quality of these existing methods. Instead, the method according to various embodiments can determine and cache information about which light sources are important for various regions of the scene and can generate actual samples of these light sources at least partially based on the relative importance of these light sources. In at least some embodiments, at least two steps can be used to determine these samples, which includes determining the importance of the light sources and then using an analytical solution to select or "pick" samples from those light sources in the scene. In at least some embodiments, this can be performed using multiple importance sampling (MIS).
[0033] However, as described above, it is difficult to determine which light sources are important in a complex or dynamic scene. For example, due to occlusion or positioning, or due to movement in the scene, only a small number of rays from a light source may affect a surface. Sampling only a subset of the possible rays in the scene can result in rays from a light source to an object not being sampled for several frames, which can incorrectly result in the computed importance of that light source for that object or region being lower than it should be determined. In some embodiments, the sampling patterns utilized between two regions of the scene may be very different, such that when tracing rays, there may be grid-like artifacts at the boundaries of various voxels. To locate at least this type of artifact, the boundaries between voxels can be blurred by jittering the position of the sampled data. This jittering or offset-based sampling can effectively turn these voxel-like artifacts into noise, which helps improve the overall image quality. In this method, the probability distribution function (PDF) for light selection can be kept consistent from frame to frame, resulting in the noise pattern gradually changing as the light sources move in the scene. In at least one embodiment, for similar reasons, the normal direction can also be jittered.
[0034] As described above, importance values can be determined at the voxel level in at least some embodiments. The sizes of these voxels may be irregular, as the octahedral voxels of the scene can alternatively be arranged as voxels of potentially varying sizes, such as voxels that increase in size with distance from the virtual camera of the scene. In various embodiments, the light information of all the rays that have been projected in the scene can be collected for each voxel, and the average light contribution can be calculated for each voxel. In at least one embodiment, this contribution data can correspond to the radiation data or the directional power quantity from each light source that impinges on the voxel. In at least one embodiment, although other light source abstractions can also be used within the scope of various embodiments, at least for importance determination, a light source such as a spherical light source can be considered a single light source, regardless of the number of triangles or other geometric components that make up the light source. In at least some embodiments, any object that emits light in the scene can be considered a light source.
[0035] These averages can then be used to generate a cumulative distribution function (CDF) for light selection or picking. A discretized CDF can be generated for each frame or every n frames, where n can be selected as a value such as 2, 4, 8, or 10 based on factors such as performance and resource availability. This approach helps to amortize the cost of building the CDF. The average contributions can be used to construct the CDF to determine the relative importance of these light sources, where the average contributions are the average contributions of each light source for each voxel, and these light sources are the light sources for each frame or every n frames relative to these light sources. The CDF of a real-valued random variable X can be given by:
[0036] F X (x) = P(X ≤ x)
[0037] where the right side of the formula represents the probability that the value of the random variable X is less than or equal to x. In different embodiments, the operations for constructing these functions (e.g., determining prefix sums and reduction sums) can be performed serially or in parallel, as these operations can involve high-performance computing primitives that can be executed in parallel on different GPUs. Various other performance enhancements can also be utilized, such as constructing the CDFs of these voxels in parallel on the voxels and lights, or constructing the CDF for groups of voxels rather than individual voxels.
[0038] In various embodiments, the data for the CDF can take the form of an array of floating-point numbers between 0 and 1, where each next consecutive number is greater than or equal to the previous number. A binary search can then be performed using this array. Each time the light source is to be sampled, the process can access the CDF for that voxel and perform a binary search, for example, performing the binary search using a random number. The index corresponding to this binary search in the array can then be utilized. The importance information can also be used to guide adaptive sampling. For example, a greater number or portion of the sampling or casting of shadow rays can be dedicated to regions with multiple important light sources in order to more accurately capture important light information, without having to spend excessive effort on regions that may have only one or two important light sources, and where casting a large number of shadow rays or capturing as many samples is not important. Thus, this method makes it possible to cast a small number of shadow rays into regions that may have only one or two important lights, while a greater number of shadow rays can be applied to scene regions where more important light appears to exist. To improve potential performance, the CDF can be compressed into a standard format. Instead of using 32-bit floating-point arithmetic, a lower precision can be used, which does not bias the result. Instead of storing the CDF values in floating-point form, the value can be stored as an 8-bit value, which can save bandwidth and cache storage.
[0039] In some embodiments, identifying a smaller number of more important lights can allow the use of additional solutions or methods that may otherwise be impractical to use. For example, a method based on linear transformation cosine (LTC) can be used for a smaller number of light sources. Additionally, knowing which lights are the most important in a region and the number of important lights can be useful for guided adaptive sampling, so that the shadow ray budget cannot be evenly utilized across the image to be rendered, but rather a portion of the shadow ray budget is more concentrated where the work is more important. In at least one embodiment, the PDF can be used to evaluate variance, for example, performing light culling. Light culling can effectively cull or remove lights that are not important for a certain region and retain lights that are important for that region or "hero" lights. Then, a potentially relatively expensive evaluation, such as an LTC evaluation, can be performed only on the important lights. Then, in at least one embodiment, the numbers in the CDF (which can include statistical information and standard deviation information) can be used to determine when more shadow rays are needed due to the higher variance.
[0040] In at least one embodiment, an attempt can be made to ensure that the PDF that receives a zero value (indicating that the light source is not important for the region or voxel) has the opportunity to obtain a non-zero value if the light source may actually have some importance for the region or voxel. As previously mentioned, the sampling method may miss rays that are transmitted from the light to the object only through small openings or due to movement in the scene. A method may be needed to capture this importance, which may not have been captured during a period of random sampling. In at least one embodiment, for each frame or subset of frames, the method can be performed by casting a plurality of rays into random light rays in the scene, regardless of their importance or being considered to have relatively low importance. The method can provide an opportunity for ray sampling for those light sources such that the corresponding PDF has a non-zero value due to the detected impact, and the sampled rays are rays that may impact the object or region. In some embodiments, the method can include casting a small number of shadow rays onto a small number of randomly (and uniformly) selected light sources for each voxel, for each frame, or for every n frames. Thus, a small number of shadow rays can be used to attempt to discover light interactions that were previously missed by the sampling method. Another method is a method that avoids allowing the value of the PDF to be 0, thus always allowing certain contributions from the light source. In at least one embodiment, the method can involve setting a minimum allowable PDF value to be assigned to the light source.
[0041] The PDF is determined based at least in part on the amount of light in the scene. For a continuous function, the probability density function (PDF) is the probability that the variable has the value x. For a discrete distribution, the PDF can be given by the following formula:
[0042] f(x) = P[X = x]
[0043] In at least one embodiment, the PDF can correspond to one divided by the amount of light in the scene multiplied by one divided by the number of spatial hash or voxel projection shadow rays. In at least one embodiment, for each frame, there can be a refresh of at least a subset of the light values. The new information obtained for a frame can be used to update the data structure storing the average radiance. At the end of the calculation of a frame, a new CDF can be established, which will have a new probability for a given light source. In at least one embodiment, the number of entries can be fixed, and a replacement strategy is used to minimize the amount of memory required for these operations. Such a method can store only the most important light information about the maximum number. In other methods, light information can be stored only for lights that have an importance value, where the lights with an importance value are lights that reach or exceed a minimum importance threshold.
[0044] As described above, in at least one embodiment, jitter-based blending can be used when sampling voxels. Other blending methods may also be used, possibly related to PDF animations. For example, this can involve linear interpolation animations between PDFs. When a new PDF is obtained at the end of a frame, there is an old PDF in the previous frame, so it may be necessary to animate the transition. The PDF can be animated for this purpose. Other methods can deviate from the mean-based method and instead use a maximum-based method. As mentioned before, there may be light that only occasionally hits a voxel or only hits a small fraction of the rays, so its impact can be minimized by a mean-based method. However, if these rays have high radiance values, this information may need to be retained. The maximum-based method can retain the maximum contribution information, which is the maximum contribution information returned for the light of any voxel, and employ a corresponding blending method between high PDFs and low PDFs, so that the radiance of the high PDF can be retained without over-sampling this light due to its relatively low overall importance.
[0045] In some embodiments, a partial CDF can be generated and stored for use. The partial CDF can only construct a given CDF on a fixed number, selection, or photon set in the scene. In one example, a random variable can be sampled and compared to a threshold b. If the value of this random sample is greater than b, light can be selected from a list of, for example, a fixed size. The partial CDF can be used for this selection, with a probability of:
[0046] Probability = (1 - b) * partial PDF value + b * (1 / n)
[0047] where n is the total light count, and the partial PDF value is constructed from the partial CDF with subtraction. If the random sample is below b, a light can be uniformly selected from all lights, with a probability of b * (1 / n). To update the structure, the maximum light contribution can be accumulated over all uniformly selected lights in the frame, and a light can be selected in the fixed-size structure to replace the maximum contributor. In at least one embodiment, the replacement can be achieved by selecting a random slot to replace, or by another method (such as selecting the lowest PDF and rolling a die) to determine whether to replace.
[0048] Figure 4An example process 400 for rendering a frame or image of a scene that can be performed according to various embodiments is shown. It should be understood that for this and other processes presented herein, unless otherwise specifically stated, additional, fewer, or alternative steps can be performed in a similar or alternative order or at least partially in parallel within the scope of the various embodiments. In this example, a number of rays are sampled 402 for multiple light sources and object rays in the scene. Data from this sampling can be used to determine 404 importance values for these light sources, where the importance values are importance values for respective regions of the scene. The scene can be divided into multiple regions using various methods, such as including dividing the regions into multiple octahedral voxels of different sizes. These importance values can be cached 406, for example, in GPU memory. In at least one embodiment, probability values for selecting each light source can be stored in the cache, where the importance values can be used to calculate the probability values. For a frame to be rendered (which may be part of a sequence of frames), a number of rays to be sampled 408 (from all possible rays that can be projected to all possible light sources) can be selected at least partially based on these cached importance values. For a given region or voxel, the number of rays to be sampled can be at least partially based on multiple important light sources and their relative importance. Once selected, sampling 410 can be performed on those selected rays. This information can be used to update the cached data for the scene. The light information obtained from these sampled rays can also be used to render 412 the frame, for example, to realistically illuminate objects in the scene based on various factors such as the position, shape, and radiance of these light sources relative to the objects in the scene. Other types of regions can also be selected, such as tiles that may involve uniform or different sizes.
[0049] Figure 5 Another example process 500 for rendering a frame that can be rendered according to at least one embodiment is shown. This example process can represent one or more implementation choices for a more general process, such as Figure 4The process 400 shown. In this example, the scene to be rendered is determined (502). For example, this can include: determining a virtual three-dimensional environment and the objects and light sources placed within or near the environment. In this example, the objects of the scene (including both foreground and background objects) can optionally be segmented 504 into a set of octahedral voxels, where other types of regions can also be used within the scope of various embodiments. In some embodiments, spatial hashing allows for determining voxel identifiers or locations in memory for positions as well as directions available during the rendering process. Rays (such as shadow rays) can be projected onto the multiple light sources in the environment. The average light contribution from these rays can be calculated 508 for each voxel. In at least some embodiments, this can include directional considerations. For the current frame in the sequence to be rendered for the scene, these averages can be used to construct 510 a cumulative distribution function (CDF) for light selection. Then, the CDF can be used to select 512 the light to be sampled during the rendering of the current frame. Then, these selected lights can be sampled 514, and in at least one embodiment, the current frame is rendered 516 using light information that is the light information from these sampled lights and the cached light information from previous frames or samples. Unless there are no more frames to render for the sequence, the next frame can be selected 518, and the process can continue with another sampling with an updated CDF.
[0050] In at least one embodiment, aspects of the presentation and rendering of content frames can be performed at various locations, such as on a client device, by a third-party provider, or in the cloud, and other such options. Figure 6An example component 600 is shown that can be used to generate, provide, and present such content according to various embodiments. In at least one embodiment, the client device 602 can use components of the content application 604 on the client device 602 and data locally stored on the client device to generate content for a session, such as a gaming session or a video viewing session. In at least one embodiment, a content application 624 (e.g., a gaming or streaming application) executing on the content server 620 can initiate a session that is at least associated with the client device 602, can also utilize a session manager and user data stored in the user database 636, and can cause the content manager 626 to determine content 632 and render it using the renderer 628 (if this type of content or platform requires), and send it to the client device 602 using an appropriate transmission manager 622, and the sending process is through downloading, streaming, or other such transmission channels. In at least one embodiment, the client device 602 that receives the content can provide the content to the corresponding content application 604, and the content application may also or alternatively include a renderer 608 that is used to render at least some of the content for presentation via the client device 602, such as video content through the display 606 and audio, and audio (e.g., sound and music) through at least one audio playback device (e.g., speakers or headphones). In at least one embodiment, at least some of the content may have been stored, rendered, or accessible on the client device 602, such that at least that portion of the content does not need to be transmitted over the network 640, for example, in the case where that portion of the content may have been previously downloaded or locally stored on a hard drive or optical disc. In at least one embodiment, a transmission mechanism such as data streaming can be used to transmit the content from the content server 620 or the content database 634 to the client device 602. In at least one embodiment, at least a portion of the content can be obtained or streamed from another source, which is another source such as a third-party content service 650, and the third-party content service 650 can also include a content application 652 for generating or providing content. In at least one embodiment, multiple computing devices or multiple processors within one or more computing devices (e.g., a combination that can include a CPU and a GPU) can be used to perform portions of the function.
[0051] In at least one embodiment, the content application 624 includes a content manager 626 that may determine or analyze content before sending the content to the client device 602. In at least one embodiment, the content manager 626 may also include (or work with) other components that are capable of generating, modifying, or enhancing the content to be provided. In at least one embodiment, this may include a renderer 628 for rendering the content, such as aliased content displayed at a first resolution. In at least one embodiment, an upsampling or scaling component 630 may generate at least one additional version of the image at a different resolution (higher or lower) and may perform at least some processing such as anti-aliasing. In at least one embodiment, as discussed herein, a blending module 632 (which may include at least one neural network) may reference one or more previous images and perform blending on one or more of those images. In at least one embodiment, the content manager 626 may then select an image or video frame at an appropriate resolution to send to the client device 602. In at least one embodiment, the content application 604 on the client device 602 may also include components such as a renderer 608 so that any or all of this functionality may be additionally or alternatively performed on the client device 602. In at least one embodiment, the content application 652 on the third-party content service system 650 may also include this functionality. In at least one embodiment, the location where at least some of this functionality is performed may be configurable or may depend on factors such as the type of client device 602 or the availability of a network connection with appropriate bandwidth, among other factors. In at least one embodiment, the upsampling module 630 or the blending module 632 may include one or more neural networks for performing or assisting with this functionality, where those neural networks (or at least the network parameters of those networks) may be provided by the content server 620 or the third-party system 650. In at least one embodiment, the system for content generation may include any suitable combination of hardware and software in one or more locations. In at least one embodiment, the image or video content generated at one or more resolutions may also be provided to (or made available to) other client devices 660, such as for downloading or streaming from a media source that stores a copy of the image or video content. In at least one embodiment, this may include sending images of game content for a multiplayer game, where different client devices may display the content at different resolutions, the different resolutions including one or more super-resolutions.
[0052] Data center
[0053] Figure 7FIG. 0 illustrates an example data center 700 in which at least one embodiment may be used. In at least one embodiment, the data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0054] In at least one embodiment, as Figure 7 shown, the data center infrastructure layer 710 may include a resource coordinator 712, grouped computing resources 714, and node computing resources (“node C.R.”) 716(1)-716(N), where “N” represents any whole positive integer. In at least one embodiment, the node C.R. 716(1)-716(N) may include, but is not limited to, any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (such as dynamic read-only memory), storage devices (such as solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, etc. In at least one embodiment, one or more of the node C.R. 716(1)-716(N) may be a server having one or more of the above computing resources.
[0055] In at least one embodiment, the grouped computing resources 714 may include separate groupings (not shown) of node C.R. housed within one or more racks, or numerous racks (also not shown) within data centers at various geographical locations. The separate groupings of node C.R. within the grouped computing resources 714 may include grouped computing, network, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node C.R. including CPUs or processors may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.
[0056] In at least one embodiment, a job scheduler 722 may configure or otherwise control one or more of the node C.R. 716(1)-716(N) and / or the grouped computing resources 714. In at least one embodiment, the job scheduler 722 may include a software design infrastructure (“SDI”) management entity for the data center 700. In at least one embodiment, the resource coordinator may include hardware, software, or some combination thereof.
[0057] In at least one embodiment, as Figure 7As shown, the framework layer 720 includes a job scheduler 732, a configuration manager 734, a resource manager 736, and a distributed file system 738. In at least one embodiment, the framework layer 720 may include a framework that supports the software 732 of the software layer 730 and / or one or more application programs 742 of the application layer 740. In at least one embodiment, the software 732 or the application program 742 may respectively include web-based service software or application programs, such as the services or application programs provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark that can utilize the distributed file system 738 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 732 may include a Spark driver to facilitate scheduling of the workloads supported by the various layers of the data center 700. In at least one embodiment, the configuration manager 734 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and the distributed file system 738 for supporting large-scale data processing. In at least one embodiment, the resource manager 736 is capable of managing the cluster or grouped computing resources mapped to or allocated for supporting the distributed file system 738 and the job scheduler 732. In at least one embodiment, the cluster or grouped computing resources may include grouped computing resources 714 on the data center infrastructure layer 710. In at least one embodiment, the resource manager 736 may coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.
[0058] In at least one embodiment, the software 732 included in the software layer 730 may include software used by at least a portion of the nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. One or more types of software may include, but are not limited to, Internet web search software, email virus scanning software, database software, and streaming video content software.
[0059] In at least one embodiment, one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of nodes C.R. 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.) or other machine learning applications used in conjunction with one or more embodiments.
[0060] In at least one embodiment, any one of the configuration manager 734, the resource manager 736, and the resource coordinator 712 may implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-modifying actions may relieve the data center operator of the data center 700 from making potentially bad configuration decisions and may avoid underutilization and / or poorly performing parts of the data center.
[0061] In at least one embodiment, the data center 700 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture by using the software and computing resources described above with respect to the data center 700. In at least one embodiment, by using the weight parameters calculated by one or more training techniques described herein, the resources described above with respect to the data center 700 may be used to infer or predict information using the trained machine learning model corresponding to one or more neural networks.
[0062] In at least one embodiment, the data center may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the above resources. In addition, the one or more software and / or hardware resources described above may be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0063] By caching optical importance information and using this information to determine ray sampling of an image or video frame, such components can be used to improve the image quality during the image generation process.
[0064] Computer system
[0065] Figure 8 is a block diagram showing an exemplary computer system in accordance with at least one embodiment. The exemplary computer system can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof that forms a processor 800, which can include execution units to execute instructions. In at least one embodiment, in accordance with the present disclosure, such as the embodiments described herein, computer system 800 can include, but is not limited to, components such as a processor 802, whose execution units include logic to execute algorithms for processing data. In at least one embodiment, computer system 800 can include a processor, such as a processor family, XeonTM, XScaleTM, and / or StrongARMTM, Core TM or Intel microprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) can also be used. In at least one embodiment, computer system 1400 can execute a version of the WINDOWS operating system obtained from Microsoft Corporation of Redmond, Wash., although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces can also be used.
[0066] Embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular telephones, Internet Protocol devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, embedded applications can include microcontrollers, digital signal processors (“DSPs”), systems on a chip, network computers (“NetPCs”), set-top boxes, network hubs, wide area network (“WAN”) switches, or any other system that can execute one or more instructions in accordance with at least one embodiment.
[0067] In at least one embodiment, computer system 800 may include, but is not limited to, a processor 802, which may include, but is not limited to, one or more execution units 808 to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, computer system 8A is a single-processor desktop or server system, but in another embodiment, computer system 8A may be a multi-processor system. In at least one embodiment, processor 802 may include, but is not limited to, a complex instruction set computer (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) microprocessor, a processor implementing an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 802 may be coupled to a processor bus 810, which may transfer data signals between processor 802 and other components in computer system 800.
[0068] In at least one embodiment, processor 802 may include, but is not limited to, a level 1 (“L1”) internal cache 804. In at least one embodiment, processor 802 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, cache memory may reside external to processor 802. Other embodiments may also include a combination of internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, register file 806 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0069] In at least one embodiment, execution units 808, including but not limited to logic for performing integer and floating-point operations, are also located in processor 802. Processor 802 may also include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode for certain macroinstructions. In at least one embodiment, execution units 808 may include logic for processing an encapsulated instruction set 809. In at least one embodiment, by including the encapsulated instruction set 809 in the instruction set of general-purpose processor 802, and the associated circuitry for the instructions to be executed, operations used by many multimedia applications may be performed using the encapsulated data in general-purpose processor 802. In one or more embodiments, operations may be performed on the encapsulated data by using the full width of the processor's data bus to accelerate and more efficiently execute many multimedia applications, which may not require transferring smaller data units on the processor's data bus to perform one or more operations on one data element at a time.
[0070] In at least one embodiment, execution unit 808 can also be used in a microcontroller, an embedded processor, a graphics device, a DSP, and other types of logic circuits. In at least one embodiment, computer system 800 can include, but is not limited to, memory 820. In at least one embodiment, memory 820 can be implemented as a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or other storage devices. In at least one embodiment, memory 820 can store instructions 819 and / or data 821 represented by data signals that can be executed by processor 802.
[0071] In at least one embodiment, the system logic chip can be coupled to processor bus 810 and memory 820. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub (“MCH”) 816, and processor 802 can communicate with MCH 816 via processor bus 810. In at least one embodiment, MCH 816 can provide a high-bandwidth memory path 818 to memory 820 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 816 can initiate data signals among processor 802, memory 820, and other components in computer system 800, and bridge data signals among processor bus 810, memory 1420, and system I / O 822. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 816 can be coupled to memory 820 through high-bandwidth memory path 818, and graphics / video card 812 can be coupled to MCH816 through an Accelerated Graphics Port (“AGP”) interconnect 814.
[0072] In at least one embodiment, the computer system 800 may use the system I / O 822 as a proprietary hub interface bus to couple the MCH 816 to an I / O controller hub (“ICH”) 830. In at least one embodiment, the ICH 830 may provide direct connections to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 820, chipset, and processor 802. Examples may include, but are not limited to, an audio controller 829, a firmware hub (“Flash BIOS”) 828, a wireless transceiver 826, a data storage 824, a legacy I / O controller 823 that includes user input and a keyboard interface, a serial expansion port 827 (e.g., Universal Serial Bus (USB)), and a network controller 834. The data storage 824 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0073] In at least one embodiment, Figure 8 a system including interconnected hardware devices or “chips” is shown, while in other embodiments, Figure 8 an exemplary system-on-a-chip (“SoC”) may be shown. In at least one embodiment, Figure 8 the devices shown therein may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the system 800 use Compute Express Link (CXL) interconnects to interconnect.
[0074] By caching optical importance information and using this information to determine ray sampling of an image or video frame, such components can be used to improve the image quality during the image generation process.
[0075] Figure 9 is a block diagram showing an electronic device 900 for utilizing a processor 910 according to at least one embodiment. In at least one embodiment, the electronic device 900 may be, for example but not limited to, a laptop computer, a tower server, a rack server, a blade server, a laptop, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0076] In at least one embodiment, system 900 may include, but is not limited to, a processor 910 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 910 is coupled using a bus or interface, such as an I2C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, Figure 9 a system is shown that includes interconnected hardware devices or “chips,” while in other embodiments, Figure 9 an exemplary System on a Chip (“SoC”) may be shown. In at least one embodiment, Figure 9 the devices shown therein may be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 9 one or more of the components of are interconnected using Compute Express Link (CXL) interconnects.
[0077] In at least one embodiment, Figure 9 it may include a display 924, a touch screen 925, a touch pad 930, a Near Field Communication unit (“NFC”) 945, a sensor hub 940, a thermal sensor 946, an Embedded Controller (“EC”) 935, a Trusted Platform Module (“TPM”) 938, a BIOS / Firmware / Flash (“BIOS, FW Flash”) 922, a DSP 960, a drive 920 (e.g., a Solid State Disk (“SSD”) or a Hard Disk Drive (“HDD”)), a Wireless Local Area Network unit (“WLAN”) 950, a Bluetooth unit 952, a Wireless Wide Area Network unit (“WWAN”) 956, a Global Positioning System (GPS) 955, a camera (“USB 3.0 camera”) 954 (e.g., a USB 3.0 camera), or a Low Power Double Data Rate (“LPDDR”) memory unit (“LPDDR3”) 915 implemented to, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0078] In at least one embodiment, other components may be communicatively coupled to the processor 910 through the components discussed above. In at least one embodiment, an accelerometer 941, an ambient light sensor (“ALS”) 942, a compass 943, and a gyroscope 944 may be communicatively coupled to the sensor hub 940. In at least one embodiment, a thermal sensor 939, a fan 937, a keyboard 946, and a touchpad 930 may be communicatively coupled to the EC 935. In at least one embodiment, a speaker 963, headphones 964, and a microphone (“mic”) 965 may be communicatively coupled to an audio unit (“audio codec and class-D amplifier”) 964, which may in turn be communicatively coupled to the DSP 960. In at least one embodiment, the audio unit 964 may include, for example but not limited to, an audio encoder / decoder (“codec”) and a class-D amplifier. In at least one embodiment, a subscriber identity module (“SIM”) 957 may be communicatively coupled to the WWAN unit 956. In at least one embodiment, components such as the WLAN unit 950, the Bluetooth unit 952, and the WWAN unit 956 may be implemented in a next-generation form factor (NGFF).
[0079] Such components can improve image quality during image generation by caching light importance information and using that information to determine ray sampling of an image or video frame.
[0080] Figure 10 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, the system 1000 includes one or more processors 1002 and one or more graphics processors 1008, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 1002 or processor cores 1007. In at least one embodiment, the system 1000 is a processing platform integrated within a system-on-chip (SoC) integrated circuit for mobile, handheld, or embedded devices.
[0081] In at least one embodiment, system 1000 can include a game console, or be included within a server-based gaming platform, the game console including a gaming and media console, a mobile gaming console, a handheld gaming console, or an online gaming console. In at least one embodiment, system 1000 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 1000 can also include a wearable device, coupled with or integrated within the wearable device, the wearable device such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, processing system 1000 is a television or a set-top box device, the television or set-top box device having one or more processors 1002 and graphics generated by one or more graphics processors 1008.
[0082] In at least one embodiment, one or more processors 1002 each include one or more processor cores 1007 to process instructions that, when executed, complete operations for system and user software. In at least one embodiment, each of the one or more processor cores 1007 is configured to process a specific instruction set 1009. In at least one embodiment, instruction set 1009 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). In at least one embodiment, processor cores 1007 can each process different instruction sets 1009, the instruction sets 1009 which can include instructions that help to emulate other instruction sets. In at least one embodiment, processor cores 1007 can also include other processing devices, such as a digital signal processor (DSP).
[0083] In at least one embodiment, processor 1002 includes a cache memory 1004. In at least one embodiment, processor 1002 can have a single internal cache memory or multiple levels of internal cache memory. In at least one embodiment, the cache memory is shared among the various components of processor 1002. In at least one embodiment, processor 1002 also uses an external cache, such as a level three (L3) cache or a last level cache (LLC) (not shown), the external cache which can be shared among the cores 1007 within the processor using known cache coherence techniques. In at least one embodiment, processor 1002 further includes a register file 1006, which can include different types of registers for storing different types of data (e.g., integer registers, floating point registers, status registers, and instruction pointer registers). In at least one embodiment, register file 1006 can include general purpose registers or other registers.
[0084] In at least one embodiment, one or more processors 1002 are coupled to one or more interface buses 1010 to transfer communication signals, such as address, data, or control signals, between the processors 1002 and other components in the system 1000. In at least one embodiment, the interface bus 1010 can be a processor bus in one embodiment, such as a version of the Direct Media Interface (DMI) bus. In at least one embodiment, the interface 1010 is not limited to the DMI bus and can also include one or more Peripheral Component Interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 1002 includes an integrated memory controller 1016 and a Platform Controller Hub 1030. In at least one embodiment, the memory controller 1016 facilitates communication between the memory device and other components of the system 1000, while the Platform Controller Hub (PCH) 1030 provides a connection to I / O devices via a local I / O bus.
[0085] In at least one embodiment, the memory device 1020 can be a Dynamic Random Access Memory (DRAM) device, a Static Random Access Memory (SRAM) device, a flash memory device, a Phase Change Memory device, or have suitable performance to be used as a process memory. In at least one embodiment, the memory device 1020 can be used as the system memory of the system 1000 to store data 1022 and instructions 1021 for use when one or more processors 1002 execute an application or process. In at least one embodiment, the memory controller 1016 is also coupled to an optional external graphics processor 1012, which can communicate with one or more graphics processors 1008 in the processor 1002 to perform graphics and media operations. In at least one embodiment, a display device 1011 can be connected to the processor 1002. In at least one embodiment, the display device 1011 can include one or more internal display devices, such as in a mobile electronic device or a laptop device connected via a display interface (e.g., DisplayPort, etc.) or an external display device. In at least one embodiment, the display device 1011 can include a Head-Mounted Display (HMD), such as a stereoscopic display device for Virtual Reality (VR) applications or Augmented Reality (AR) applications.
[0086] In at least one embodiment, the platform controller hub 1030 enables peripheral devices to be connected to the memory device 1020 and the processor 1002 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 1046, a network controller 1034, a firmware interface 1028, a wireless transceiver 1026, a touch sensor 1025, and a data storage device 1024 (e.g., a hard disk drive, a flash memory, etc.). In at least one embodiment, the data storage device 1024 can be connected via a memory interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 1025 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1026 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1028 supports communication with the system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 1034 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1010. In at least one embodiment, the audio controller 1046 is a multi-channel high-definition audio controller. In at least one embodiment, the system 1000 includes an optional legacy I / O controller 1040 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. In at least one embodiment, the platform controller hub 1030 can also be connected to one or more Universal Serial Bus (USB) controllers 1042 that connect input devices, such as a combination of a keyboard and a mouse 1043, a camera 1044, or other USB input devices.
[0087] In at least one embodiment, instances of the memory controller 1016 and the platform controller hub 1030 can be integrated into a discrete external graphics processor, such as the external graphics processor 1012. In at least one embodiment, the platform controller hub 1030 and / or the memory controller 1016 can be external to one or more of the processors 1002. For example, in at least one embodiment, the system 1000 can include an external memory controller 1016 and a platform controller hub 1030, and the system can be configured as a memory controller hub and a peripheral controller hub in a system-on-chip that communicates with the processor 1002.
[0088] By caching light importance information and using this information to determine ray sampling of an image or video frame, such components can be used to improve image quality during image generation.
[0089] Figure 11 is a block diagram of a processor 1100 having one or more processor cores 1102A - 1102N, an integrated memory controller 1114, and an integrated graphics processor 1108 according to at least one embodiment. In at least one embodiment, processor 1100 may include additional cores, the additional cores up to and including additional core 1102N represented by the dashed box. In at least one embodiment, each processor core 1102A - 1102N includes one or more internal cache units 1104A - 1104N. In at least one embodiment, each processor core may also access one or more shared cache units 1106.
[0090] In at least one embodiment, internal cache units 1104A - 1104N and shared cache units 1106 represent the cache memory hierarchy within processor 1100. In at least one embodiment, internal cache units 1104A - 1104N may include at least one level of instruction and data cache within each processor core, as well as one or more shared mid - level caches, such as a second - level (L2), third - level (L3), fourth - level (L4), or other level cache, where the highest - level cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence between the various cache units 1106 and 1104A - 1104N.
[0091] In at least one embodiment, processor 1100 may also include a set of one or more bus controller units 1116 and a system agent core 1110. In at least one embodiment, one or more bus controller units 1116 manage a set of peripheral buses, such as one or more PCI or PCI Express buses. In at least one embodiment, system agent core 1110 provides management functions for various processor components. In at least one embodiment, system agent core 1110 includes one or more integrated memory controllers 1114 to manage access to various external memory devices (not shown).
[0092] In at least one embodiment, one or more processor cores 1102A-1102N include support for concurrent multithreading. In at least one embodiment, the system agent core 1110 includes components for coordinating and operating cores 1102A-1102N during multithreaded processing. In at least one embodiment, the system agent core 1110 may additionally include a power control unit (PCU) that includes logic and components to regulate one or more power states of the processor cores 1102A-1102N and the graphics processor 1108.
[0093] In at least one embodiment, the processor 1100 additionally includes a graphics processor 1108 to perform graphics processing operations. In at least one embodiment, the graphics processor 1108 is coupled to the shared cache unit 1106 and the system agent core 1110, which includes one or more integrated memory controllers 1114. In at least one embodiment, the system agent core 1110 also includes a display controller 1111 to drive the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1111 may also be a separate module coupled to the graphics processor 1108, or may be integrated within the graphics processor 1108, and the coupling is via at least one interconnect.
[0094] In at least one embodiment, a ring-based interconnect unit 1112 is used to couple the internal components of the processor 1100. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1108 is coupled to the ring interconnect 1112 via an I / O link 1113.
[0095] In at least one embodiment, the I / O link 1113 represents at least one of a variety of I / O interconnects, which includes a package I / O interconnect that facilitates communication between various processor components and a high-performance embedded memory module 1118 (such as an eDRAM module). In at least one embodiment, each processor core 1102A-1102N and the graphics processor 1108 use the embedded memory module 1118 as a shared last-level cache.
[0096] In at least one embodiment, the processor cores 1102A-1102N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1102A-1102N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 1102A-1102N execute a common instruction set and one or more other processor cores 1102A-1102N execute a subset of the common instruction set or a different instruction set. In at least one embodiment, in terms of microarchitecture, the processor cores 1102A-1102N are heterogeneous, where one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 1100 may be implemented on one or more chips or as a SoC integrated circuit.
[0097] By caching light importance information and using this information to determine ray sampling of an image or video frame, such components can improve image quality during the image generation process.
[0098] Other variations are within the spirit of the present disclosure. Thus, although the disclosed techniques are susceptible to various modifications and alternative configurations, certain of its illustrated embodiments have been shown in the figures and described in detail above. However, it should be understood that the intention is not to limit the disclosure to the one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative configurations, and equivalents falling within the spirit and scope of the disclosure as defined by the appended claims.
[0099] Other variations are within the spirit of the present disclosure. Thus, although the disclosed techniques are susceptible to various modifications and alternative configurations, certain of its illustrated embodiments have been shown in the figures and described in detail above. However, it should be understood that the intention is not to limit the disclosure to the one or more specific forms disclosed, but on the contrary, the intention is to cover all modifications, alternative configurations, and equivalents falling within the spirit and scope of the disclosure as defined by the appended claims.
[0100] Unless otherwise expressly indicated or clearly inconsistent with the context, conjunctive language such as the phrase "at least one of A, B, and C" or "at least one of A, B, and C" is understood in the context to typically mean that the items, terms, etc. can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in a context having three members, the conjunctive phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunctive language is not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise specified or inconsistent with the context, the term "plural" represents a plural state (e.g., "a plurality of items" means a plurality of items). The number of items in the plural is at least two, but can be more when expressly or indicated by the context. Further, unless otherwise specified or clear from the context, the phrase "based on" means "at least partially based on" rather than "based solely on".
[0101] Unless otherwise indicated herein or otherwise clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable sequence. In at least one embodiment, a process such as those described herein (or variations and / or combinations thereof) is performed under the control of one or more computer systems configured with executable instructions and is implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors by hardware or combinations thereof. In at least one embodiment, the code is stored, for example, in the form of a computer program on a computer-readable storage medium that includes a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that does not include transitory signals (e.g., propagated transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuits (e.g., buffers, caches, and queues) in a transceiver of a transitory signal. In at least one embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, and when executed by one or more processors of a computer system (i.e., as a result of being executed) causes the computer system to perform the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media and one or more individual non-transitory storage media of a plurality of non-transitory computer-readable storage media that lack all code, and the plurality of non-transitory computer-readable storage media jointly store all code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors, e.g., the non-transitory computer-readable storage medium stores instructions and a main central processing unit (“CPU”) executes some instructions while a graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of a computer system have separate processors and different processors execute different subsets of instructions.
[0102] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or jointly perform the operations of the processes described herein, and such a computer system is configured with suitable hardware and / or software capable of implementing the operations. Further, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, is a distributed computer system that includes a plurality of devices operating in different ways such that the distributed computer system performs the operations described herein and such that a single device does not perform all operations.
[0103] The use of any and all examples or exemplary language (e.g., "such as") provided herein is for illustrative purposes only to better clarify the embodiments of the present disclosure and does not limit the scope of the disclosure, unless otherwise stated. Any language in the specification should not be construed as indicating any non-claimed element essential to the practice of the disclosure.
[0104] All references cited herein, including publications, patent applications, and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated by reference.
[0105] In the description and claims, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not necessarily intended as synonyms for each other. Rather, in a particular example, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0106] Unless otherwise specifically stated, it is understood that throughout the specification, terms such as "processing", "computing", "operating", "determining", etc., refer to actions and / or processes of a computer or computing system, or similar electronic computing devices, that process and / or transform data represented as physical quantities (e.g., electrons) in the registers and / or memory of the computing system into other data similarly represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0107] In a similar manner, the term "processor" may refer to any device or part of a device that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Moreover, each process may refer to multiple processes that execute instructions sequentially or in parallel, continuously or intermittently. Since a system may embody one or more methods and a method may be considered a system, the terms "system" and "method" may be used interchangeably herein.
[0108] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. The obtaining, acquiring, receiving, or inputting of analog and digital data may be accomplished in a variety of ways, such as by receiving data as an argument to a function call or a call to an application programming interface. In some embodiments, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting the data via a serial or parallel interface. In another embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data may be accomplished by transmitting the data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data may be implemented by transmitting the data as an input or output argument to a function call, an application programming interface, or an interprocess communication mechanism.
[0109] Although the foregoing discussion sets forth example implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to be within the scope of the present disclosure. Additionally, although specific allocations of responsibility were defined above for discussion purposes, the various functions and responsibilities may be allocated and divided differently depending on circumstances.
[0110] Moreover, although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter claimed in the appended claims need not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
Claims
1. A computer-implemented method, comprising: Determining light information for a first set of rays projected by two or more light sources in a virtual environment; Determining values of the two or more light sources relative to a plurality of spatial regions of the virtual environment, at least in part based on the light information, wherein a spatial hashing algorithm divides the virtual environment into the plurality of spatial regions corresponding to a plurality of non-cubic voxels; Selecting, at least in part based on the values, a second set of rays to be sampled for the two or more light sources relative to the plurality of spatial regions, the second set of rays including a greater number of samples for the light sources having higher values; Sampling the second set of rays to obtain updated lighting information for the plurality of spatial regions; And Using the updated lighting information to render an image of the virtual environment.
2. The computer-implemented method according to claim 1, further comprising: Using the spatial hashing algorithm to determine the plurality of spatial regions of the virtual environment, wherein the spatial regions also provide direction information.
3. The computer-implemented method according to claim 2, wherein the plurality of non-cubic voxels are a plurality of octahedral voxels, the plurality of octahedral voxels having one or more dimensions.
4. The computer-implemented method according to claim 3, further comprising: Determining the values of the selected light sources, at least in part based on the directional information of the selected light sources relative to respective octahedral voxels.
5. The computer-implemented method according to claim 1, further comprising: Constructing a cumulative distribution function CDF for the plurality of spatial regions using the values; And Caching selection probability data determined according to the CDF of the plurality of spatial regions, the selection probability data being used to select the second set of rays.
6. The computer-implemented method according to claim 5, further comprising: Updating the CDF of the plurality of spatial regions for at least one image subset of an image sequence of the virtual environment.
7. The computer-implemented method according to claim 1, further comprising: Selecting at most a maximum number of light sources having the highest values and sampling the second set of rays from the light sources.
8. The computer-implemented method according to claim 1, further comprising: Determining the values of the two or more light sources relative to the plurality of spatial regions, at least in part based on the average light contribution of a single light source relative to a single spatial region.
9. The computer-implemented method according to claim 1, further comprising: Causing two or more light sources having determined values below a threshold to be considered for sampling one or more subsequent images to be rendered.
10. A system, comprising: A processor; And A memory including instructions that, when executed by the processor, cause the system to perform the following operations: Determining light information for a first set of rays projected by two or more light sources in a virtual environment; Determine values for the two or more light sources relative to multiple spatial regions of the virtual environment, at least in part based on the light information, wherein a spatial hashing algorithm divides the virtual environment into the multiple spatial regions corresponding to multiple non-cubic voxels; Select a second set of rays to be sampled for the two or more light sources relative to the multiple spatial regions, at least in part based on the values, the second set of rays including a greater number of samples for light sources having higher values; Sample the second set of rays to obtain updated lighting information for the multiple spatial regions; And Use the updated lighting information to render an image of the virtual environment.
11. The system according to claim 10, wherein executing the instructions further causes the system to: Use the spatial hashing algorithm to determine the multiple spatial regions of the virtual environment.
12. The system according to claim 11, wherein the multiple non-cubic voxels are multiple octahedral voxels, the multiple octahedral voxels having one or more dimensions.
13. The system according to claim 11, wherein executing the instructions further causes the system to: Determine the values of the selected light sources at least in part based on directional information of the selected light sources relative to respective octahedral voxels.
14. The system according to claim 10, wherein executing the instructions further causes the system to: Use the values to construct a cumulative distribution function (CDF) for the multiple spatial regions; and Cache selection probability data determined according to the CDF of the multiple spatial regions, the selection probability data being used to select the second set of rays.
15. The system according to claim 14, wherein executing the instructions further causes the system to: Update the CDF of the multiple spatial regions for at least one image subset of an image sequence of the virtual environment.
16. The system according to claim 10, wherein executing the instructions further causes the system to: Determine the values of the two or more light sources relative to the multiple spatial regions at least in part based on an average light contribution of a single light source relative to a single spatial region.
17. The system according to claim 10, wherein the system comprises at least one of the following: A system for performing graphics rendering operations; A system for performing simulation operations; A system for performing simulation operations to test or validate autonomous machine applications; A system for performing deep learning operations; A system implemented using edge devices; A system comprising one or more virtual machines (VMs); A system implemented at least in part in a data center; or A system implemented at least in part using cloud computing resources.
18. A non-transitory computer-readable storage medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the following operations: Determine light information for a first set of rays projected by two or more light sources in a virtual environment; Determine values for the two or more light sources relative to multiple spatial regions of the virtual environment, at least in part based on the light information, where a spatial hashing algorithm divides the virtual environment into the multiple spatial regions corresponding to multiple non-cubic voxels; Select a second set of rays to sample for the two or more light sources relative to the multiple spatial regions, at least in part based on the values, the second set of rays including a greater number of samples for light sources having higher values; Sample the second set of rays to obtain updated lighting information for the multiple spatial regions; And Use the updated lighting information to render an image of the virtual environment.
19. The non-transitory computer-readable storage medium according to claim 18, wherein executing the instructions further causes the one or more processors to: Use the spatial hashing algorithm to determine the multiple spatial regions of the virtual environment, where the multiple spatial regions are octahedral voxels of one or more dimensions.
20. The non-transitory computer-readable storage medium according to claim 19, wherein executing the instructions further causes the one or more processors to: Determine the values of the selected light sources, at least in part based on directional information of the selected light sources relative to respective octahedral voxels.
Citation Information
Patent Citations
Density coordinate hashing for volumetric data
CN111727462A
Distribution Caching for Direct Lights
US20160071308A1