Adaptive sampling at target sampling rate
By adaptively sampling at the target sampling rate in the graphics system, matching the importance graph and quantizing the sample number using random values, the trade-off problem of image quality and frame rate in the prior art is solved, and higher frame rates and lower power consumption are achieved.
Patent Information
- Application Number
- CN202111281960.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-09
- Filing Date
- 2021-11-01
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-11-01
AI Technical Summary
The prior art is difficult to achieve the optimal trade-off between image quality and frame rate in interactive graphics systems. The step-by-step distribution of adaptive sampling technology cannot effectively match the importance map, resulting in the inability to meet the target image quality and frame rate.
By distributing samples on the image plane to match the importance map while meeting the target average sampling rate per pixel, the sample number is quantified using random values to produce a smoother sampling distribution.
This enables increased frame rate while providing the same visual experience quality, reduces rendering computing power, allows operation at lower power, and better match target frame rate and image quality.
Smart Images

Figure CN114529443B_ABST
Abstract
Description
[0001] Claiming priority
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 108,643, filed on November 2, 2020, entitled “Adaptive Sampling for Ray-Tracing Using a Fixed Budget with Pseudorandom Properties,” and U.S. Provisional Application No. 63 / 127,488, filed on December 18, 2020, entitled “Power-of-Two Adaptive Sampling for Rendering Using a Fixed Budget,” the entire contents of both applications are incorporated herein by reference. Background Art
[0003] Interactive graphics systems provide limited real-time rendering capabilities, forcing a trade-off between image quality and frame rate. Rather than simply increasing the number of samples per pixel (spp) for each pixel to improve image quality, adaptive sampling distributes a fixed budget of samples throughout the rendered image, thereby allocating more samples to pixels that will benefit most from allocating more and fewer samples to other pixels. Conventional adaptive sampling techniques use an importance map (an array as large as the rendered image) to identify pixels that will benefit most from more samples. The importance value of each pixel is usually not an integer, and the resulting number of samples for each pixel is rounded up or truncated to determine an integer number of samples for each pixel. Generating integer values by rounding up or truncating results in a "staircase" distribution. The staircase distribution is usually not averaged over time to equal the importance map, and therefore the desired image quality may not be achieved with respect to a fixed budget. It is necessary to solve these problems and / or other problems associated with the prior art. Summary of the invention
[0004] Embodiments of the present disclosure relate to adaptive sampling at a target sampling rate. Systems and methods are disclosed for distributing samples on an image plane in a manner that matches an importance map while also attempting to meet a target average per-pixel sampling rate (e.g., a fixed budget) on the image so as not to exceed the fixed budget. The target (average) per-pixel sampling rate can be determined based on a desired frame rate or image quality. The number of samples assigned to each pixel is quantized using a random value so that the quantized distribution is smoother than conventional techniques such as those described above.
[0005] Meeting the target sampling rate reduces the computational effort to render each image compared to rendering all pixels at a higher sampling rate. Therefore, meeting the target sampling rate can advantageously allow a rendering system to operate at a higher frame rate while providing the same quality of visual experience. A rendering system that is not frame rate-constrained can advantageously operate at lower power. Adaptive shading is typically applied to rendering graphics primitives, including two-dimensional (2D) graphics primitives and three-dimensional (3D) graphics primitives.
[0006] A method, computer-readable medium, and system for adaptive sampling are disclosed, comprising: determining a total number of samples for an image based on a target per-pixel sampling rate for the image, distributing the total number of samples across pixels included in the image according to an importance map to produce an initial sampling map for the image, and quantizing the initial sampling map using per-pixel random values to produce a sampling map for the image, the sampling map including an integer number of samples for each pixel in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The present system and method for adaptive sampling are described in detail below with reference to the accompanying drawings, in which:
[0008] FIG. 1A shows an importance map and a sampling map generated using prior art techniques.
[0009] Figure 1B A sampling diagram suitable for implementing some embodiments of the present disclosure is shown.
[0010] Figure 1C Importance maps and sampling maps suitable for implementing some embodiments of the present disclosure are shown.
[0011] Figure 1D An adaptive sampling diagram according to an embodiment is shown.
[0012] Figure 1E An adaptive power-of-two sampling diagram suitable for implementing some embodiments of the present disclosure is shown.
[0013] Figure 1F A flow chart of a method suitable for implementing some embodiments of the present disclosure is shown.
[0014] Figure 2A A block diagram of an example adaptive sampling system suitable for implementing some embodiments of the present disclosure is shown.
[0015] Figure 2B A flow chart of another method suitable for implementing some embodiments of the present disclosure is shown.
[0016] Figure 2C A flow chart of a method for adaptive sampling using a power-of-two sampling rate according to an embodiment is shown.
[0017] Figure 3A An importance map and an adaptive sampling map according to an embodiment are shown.
[0018] Figure 3B An adaptive sampling diagram according to an embodiment is shown.
[0019] FIG. 3C shows a prior art sampling graph generated using prior art techniques.
[0020] Figure 4 An example parallel processing unit suitable for implementing some embodiments of the present disclosure is shown.
[0021] Figure 5A For use Figure 4 A conceptual diagram of a PPU-implemented processing system suitable for implementing some embodiments of the present disclosure.
[0022] Figure 5B Exemplary systems are shown in which different architectures and / or functionality of different previous embodiments may be implemented.
[0023] Figure 5C Components of an exemplary system that can be used to train and utilize machine learning in at least one embodiment are shown.
[0024] Fig. 6A is suitable for implementing some embodiments of the present disclosure. Figure 4 Conceptual diagram of the graphics processing pipeline implemented by the PPU.
[0025] Figure 6B An exemplary streaming system suitable for implementing some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0026] Disclosed are systems and methods for adaptive sampling at a target sampling rate. The target sampling rate can be determined based on a desired frame rate or image quality. The target sampling rate is applied to the number of pixels per image (or frame) to calculate the total number of samples. In an embodiment, the target sampling rate is not an integer and is represented as a floating point number to best correspond to the real-time frame rate. The total number of samples is then distributed over the image in a manner based on an importance map, and the total number of samples is quantized to an integer sampling rate for each pixel. Quantization is performed using random values to produce a distribution that closely matches the distribution of the importance map.
[0027] FIG. 1A shows an importance map 100 and a threshold-based sampling map 105 generated using prior art techniques. The importance map is typically as large as the image to be rendered and defines an importance or priority value for each pixel in the image. The values can be used to determine how many samples should be used to produce the final value of each pixel. Brighter areas of the importance map 100 correspond to higher values and darker areas correspond to lower values. An ideal sampling map would have a sampling rate that corresponds to the importance map 100—changing smoothly as the sampling rate and / or importance values increase and / or decrease.
[0028] In an embodiment, the importance map 100 is a downsampled version of the full resolution image, in which case nearest neighbor filtering, or bilinear filtering, or some higher order filter should be used to access the importance map 100. A threshold-based sampling map 105 is generated based on the importance map 100 using conventional techniques.
[0029] In the example, when the target sampling rate specifies an average number of samples per pixel d=4.2, each pixel should receive an average of 4.2 samples, and the samples should be spread over the pixels according to the importance map 100. The threshold-based sampling map 105 has an average sampling rate of 4.22 per pixel. The threshold-based sampling map 105 is calculated based on the importance map 100 using round() or floor(), which produces a "staircased" distribution.
[0030] Figure 1B 1 shows a sampling map suitable for implementing some embodiments of the present disclosure. The threshold-based power-of-two sampling map 110 has an average sampling rate per pixel of 4.05 and includes a threshold value limited to a power of two (2 i , where i is a non-negative integer). The steps are more obvious in the threshold-based power-of-two sampling graph 110 compared to the threshold-based sampling graph 105. When sampling using a low deviation sequence (LDS), power-of-two values may be preferred.
[0031] Thresholding may not always average out to a target (average) per-pixel sampling rate close to d. For example, if all pixels have sampling rate = 2.45 based on the importance map, then each pixel is assigned 2 samples, while if all pixels have sampling rate = 2.55, then each pixel is assigned 3 samples. Failure to match the target per-pixel sampling rate can result in under- or over-utilization of hardware resources, in which case the application may not meet its real-time requirements. When the pixel sampling rate is constrained to a value that is a power of two, then the target per-pixel sampling rate may be even more difficult to match when thresholding is used.
[0032] Figure 1CImportance maps and sampling maps suitable for implementing some embodiments of the present disclosure are shown. The importance values per pixel are 10, 80, 5, and 5. When the target sampling rate is d=5 samples per pixel, the number of samples is calculated as 2, 16, 1, and 1, respectively. When d=2.5, the number of samples is calculated as 1, 8, and two other integer values. When thresholding is used, both other values can be equal to 0 or 1. However, setting one of the other values to 0 and setting the other to 1 will produce an average sampling rate that is closer to the target sampling rate. Unfortunately, determining the sampling rate of the two bottom pixels as two different values is a sequential operation because the value of one depends on the other. In contrast, using thresholding to round up or down to the nearest integer can be performed in parallel for all pixels, but produces two 0s or two 1s, which is undesirable. As further described herein, quantization techniques that can be performed in parallel can also calculate sampling rates that average to values closer to the target sampling rate and closely match the importance map.
[0033] Figure 1D An adaptive sampling map 120 according to an embodiment is shown. The adaptive sampling map 120 is calculated based on the importance map 100 and the random values used to quantize the sample values per pixel. Compared with the threshold-based sampling map 105, which is also calculated based on the importance map 100, the adaptive sampling map 120 is smoother and corresponds more closely to the importance map 100. The sample values per pixel n(x, y) in the adaptive sampling map 120 are calculated by first calculating the total number of samples of the image as the product of the target sampling rate d and the image (and importance map) dimensions w and h. The total number of samples represents the total budget of samples to be distributed across all pixels in the image. The importance values i(x, y) of the image are summed to produce an importance sum s=∑i(x, y). Then, the factor is calculated as the total number of samples (e.g., the product of d, w, and h) divided by the importance sum. Therefore, the factor is dwh / s. Alternatively, the average importance per pixel (ipp) can be calculated as the importance sum divided by the product of the image dimensions, ipp=s / wh. The factor is then calculated as d / ipp = dwh / s. The factor is calculated once for all pixels in the image and can be considered as the first step of the adaptive sampling algorithm. The importance value and / or the sample value per pixel can vary for each image and / or over time. For example, a frame sequence may include a sequence of images, a per-image importance value and / or a per-image random value.
[0034] The second step can be performed in parallel for all pixels. The initial sampling map can be calculated by scaling the importance value of each pixel by a factor to produce an initial sampling rate for each pixel n_initial(x, y) = factor i(x, y). The importance value multiplied by the factor distributes the total number of samples over all pixels and ensures that the target sampling rate of the initial sampling map is not exceeded.
[0035] Typically, all values in the initial sampling map n_initial(x, y) will not be integers, and the final sampling rate should be an integer because rendering is performed using integer sampling rates. Therefore, the non-integer values of the initial sampling map are quantized to produce integer sampling rates in the adaptive sampling map n(x, y). The values can be quantized by rounding each value to the nearest lower integer (e.g., floor operation) or using a conventional rounding function. However, conventional rounding produces results similar to the thresholding shown in the threshold-based sampling map 105 shown in Figure 1A. In an embodiment, the initial sampling rate is quantized using a per-pixel random value instead.
[0036] More specifically, if the random number (between 0 and 1) is less than or equal to the fractional portion, the fractional portion of each initial sampling rate may be removed and additional samples may be added to the remaining integer portion of the floating point value. For example, if the initial sampling rate of a pixel is 5.25, a random number r∈[0,1] is used to determine whether to use 5 or 6 samples as the sampling rate of the pixel. The random number may be calculated on the fly or looked up in a texture containing random values. If r>0.25, the initial sampling rate is quantized to 5 samples, otherwise, the initial sampling rate is quantized to 6 samples.
[0037] Figure 1E An adaptive power-of-two sampling map 125 suitable for implementing some embodiments of the present disclosure is shown. Compared to the adaptive sampling map 120, the initial sampling rate is quantized to a power-of-two value using a per-pixel random value. Compared to the threshold-based power-of-two sampling map 110, the adaptive power-of-two sampling map 125 is smoother and corresponds more closely to the importance map 100.
[0038] Advantageously, the initial sampling rate for each pixel can be quantized independently, and therefore the initial sampling rates can be quantized in parallel. Furthermore, any error between the total number of samples and the sum of the quantized sampling rates for all pixels decreases over time because the fractional portion acts as a probability value that controls the distribution of the quantization results (increasing or not). In other words, for any particular initial sampling rate value, there will be a distribution of quantized sampling rates having a value equal to the integer portion of the initial sampling rate value and a value equal to the integer portion incremented by one. More specifically, for a fractional portion of 0.25, the integer portion increases on average 25% of the time, and does not increase 75% of the time.
[0039] In contrast, using conventional thresholding techniques (where all sampling rates equal to 5.25 are quantized to 5 or 6), quantization using random numbers causes most initial sampling rates equal to 5.25 to become 5 and other initial sampling rates to become 6. In another example, quantizing all sampling rates equal to 2.45 to 2 using conventional thresholding may result in underutilization of hardware resources, while quantizing all sampling rates equal to 2.55 to 3 may result in failure to meet a target frame rate. Quantizing the initial sampling map using random values produces a distribution that better matches the distribution of the importance map, and may more easily meet a target frame rate and / or target sampling rate. Because the generation of the initial sampling map may be performed in parallel for all pixels and quantization may be performed in parallel for all pixels, adaptive sampling may be generated quickly - achieving real-time frame rates.
[0040] Table 1 includes pseudo code for an adaptive sampling technique that can be performed in parallel for each pixel. After calculating the ipp and factors for the importance map in the first step, the initial sampling map is calculated in the second step. The initial sampling map includes SPP_float values for each pixel. The SPP_float values are separated into an integer part and a fractional part for quantization. The integer part (whole_samples) is equal to the truncated SPP_float (under truncation). The fractional part (fraction) is the truncated part of the SPP_float. The fractional part of each pixel is compared with a random number between zero and one (including zero and one).
[0041] Table 1: Adaptive sampling pseudo code
[0042]
[0043] Compared to conventional thresholding techniques, the integer number of samples for each pixel is incremented or not incremented based on the comparison of the fractional part of n_initial(x, y) (or SPP_float(x, y)) with a random value. In an embodiment, the per-pixel random values are generated by a random function or orthogonal array sampling. In an embodiment, the per-pixel random values include a random texture, a multi-jitter texture, a blue noise texture, a red noise texture, etc.
[0044] Now more illustrative information about different optional architectures and features that can implement the aforementioned framework will be described according to the user's desire. It should be strongly noted that the following information is described for illustrative purposes and should not be interpreted as limiting in any way. Any of the following features can be optionally combined with or without excluding other features described.
[0045] Figure 1FA flow chart of a method 150 suitable for implementing some embodiments of the present disclosure is shown. Each box of the method 150 described herein includes a computing process that can be performed using any combination of hardware, firmware and / or software. For example, different functions can be performed by a processor that executes instructions stored in a memory. The method 150 can also be embodied as computer-usable instructions stored on a computer storage medium. To name a few, the method can be provided by a stand-alone application, a service or a hosted service (independently or in combination with another hosted service), or a plug-in for another product. In addition, the method 150 can be performed in addition or alternatively by any one system or any combination of systems, including but not limited to those described herein. In addition, it will be understood by a person of ordinary skill in the art that any system that performs the method 150 is within the scope and spirit of the embodiments of the present disclosure.
[0046] At step 155, the total number of samples of the image is determined based on the target per-pixel sampling rate (d) of the image. In an embodiment, the target per-pixel sampling rate is calculated based on the rendering frame rate and the target frame rate. In an embodiment, the total number of samples is the product of the target per-pixel sampling rate and the size of the image. In an embodiment, the total number of samples is calculated as dwh, where w is the width of the image and h is the height of the image.
[0047] In step 160, the total number of samples is distributed across the pixels included in the image according to the importance map to generate an initial sampling map of the image. In an embodiment, step 160 is performed in parallel for at least a portion of the pixels in the image. In an embodiment, the per-pixel values in the importance map are summed to generate an importance sum s=∑i(x,y). The factor for distributing the total number of samples can be calculated by dividing the total number of samples by s. In an embodiment, the average importance per pixel ipp is calculated as s / wh and the factor is calculated as d / ipp. In an embodiment, d, s, ipp and the factor are floating point numbers. In an embodiment, the total number of samples is distributed across each pixel in the image by multiplying the factor by the per-pixel value in the importance map of the pixel to generate an initial sampling rate for the pixel. In an embodiment, a minimum number of samples is allocated to each pixel before the remainder of the total number of samples is allocated across the pixels. In an embodiment, the remainder of the total number of samples is distributed over the pixels according to the importance map. For example, a minimum number of samples equal to one may be distributed to each pixel.
[0048] In an embodiment, the minimum sampling rate is greater than zero and less than the target sampling rate. In an embodiment, the total number of samples n(x, y) for the sampling graph ideally satisfies the following:
[0049] ∑n(x,y)=dwh=mwh+(dm)wh Equation (1)
[0050] Where mwh is the minimum number of samples assigned to a pixel and (dm)wh is the remainder assigned to the pixel according to the importance map. Summing and normalizing the importance over the image yields
[0051]
[0052] The variable k is defined as the remainder divided by the sum of the importance k = (dm)wh / s, where k is equal to a factor of m = 0, and substituting k into equation (2) results in ∑ki(x, y) = (dm)wh.
[0053] In step 165, the initial sampling map is quantized using the per-pixel random values to produce a sampling map of the image, the sampling map including an integer number of samples for each pixel in the image. In an embodiment, the integer is a value that is a power of two. In an embodiment, step 165 is performed in parallel for at least a portion of the pixels in the image. In an embodiment, the image is a frame included in a frame sequence (e.g., a video), and the per-pixel random values vary for each frame in the sequence. In an embodiment, the per-pixel random values are read from a texture map or generated by a function. In an embodiment, the scene is rendered according to the sampling map to produce a frame. In an embodiment, the scene is rendered using ray tracing, and the sampling map controls the number of rays cast for each pixel in the frame.
[0054] In an embodiment, quantizing the initial sampling map includes: for each pixel in the image, dividing the initial sampling rate of the pixel included in the initial sampling map into an integer part and a fractional part. When the random value is less than the fractional part, the integer part is incremented to generate a quantized sampling rate included in the sampling map of the pixel. When the random value is not less than the fractional part, the quantized sampling rate is set equal to the integer part.
[0055] In an embodiment, the initial sampling rate n(x, y) is quantized using the following equation:
[0056]
[0057] where frac() returns the fractional part of its argument, and (a?b:c) returns b if a is true, and c otherwise. In the example, if ki(x,y)=5.25, i.e., the ideal number of samples for a pixel, then the first part of the expression in equation (3) is the minimum number of samples m plus The fractional part frac(5.25)=0.25, which means that there is a probability of 0.25 to increase by 1, and a probability of 0.75 to increase by 0. The random value r(x, y, t) per pixel may vary with time (t).
[0058] To quantize the initial sampling rate to a power of 2 using rounding without any random values, the following equation can be used:
[0059] n(x, y) = b + Δround(f) Equation (4)
[0060] where b = 2 p is a power of 2 calculated separately for each pixel, and f ∈ [0, 1]. In most cases, Δ = b, which means that n(x, y) is either b or 2b, i.e., a power of 2. In an embodiment, the minimum number m of samples is a power of 2. Depending on whether the initial sampling rate is zero or greater than zero, the variables f, b, and Δ are calculated differently:
[0061] If
[0062] If
[0063]
[0064] Note that when c is zero, f is reduced to f = ki(x, y), because m must be zero and b = 0 and Δ = 1. When the quantized sampling rate is limited to a power of 2, it is assumed that the function h(n) returns the first high-order bit (e.g., leading one) in the positive integer n. When c > 0, b becomes equal to the largest power of 2 less than or equal to c, and rounding is used to quantize n(x, y) to b or 2b. For example, if m = 0 and c = 14.2, then h(14.2) = 3 and thus b = 2 3 = 8. Thus, n = 16, because Δ = 8 and f = (14.2 - 8) / 8 = 0.775 and round(0.775) = 1.
[0065] However, using a rounding operation to quantize the sampling rate typically results in a stepped quantization, such as Figure 1B shown in the threshold-based power-of-2 sampling of FIG. 110. Equation (6) can be used to quantize the sampling rate to a power-of-2 value using a per-pixel random number, resulting in a more accurate sampling map, such as Figure 1E shown in the adaptive power-of-2 sampling of FIG. 125. The values of b, f, and Δ are calculated according to Equation (5).
[0066] n(x, y) = b + (r(x, y, t) < f? Δ : 0) Equation (6)
[0067] Figure 2AA block diagram of an example adaptive sampling system suitable for implementing some embodiments of the present disclosure is shown. It should be understood that this and other arrangements described herein are set forth as examples only. In addition to or in place of those arrangements and elements shown, other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used, and some elements may be omitted together. Further, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and implemented in any suitable combination and position. The different functions described herein as being performed by an entity may be performed by hardware, firmware, and / or software. For example, the different functions may be performed by a processor that executes instructions stored in a memory. In addition, it will be appreciated by those of ordinary skill in the art that any system that performs the operation of the adaptive sampling system 200 is within the scope and spirit of the embodiments of the present disclosure.
[0068] The adaptive sampling system 200 includes a memory 210, a sampling map generation unit 220, and a processor 225. In an embodiment, the operations performed by the sampling map generation unit 220 may be performed by the processor 225. The memory 210 stores the random texture 230, the importance map 205, and the sampling map 215. The sampling map generation unit 220 receives a sampling control, which may include one or more of a target sampling rate (d), a target frame rate, image dimensions (w and h), a minimum sampling rate (m), and power-of-two sampling enable / disable. The sampling control may also specify a random texture, a random function, or other source of random numbers to be used to quantize the sampling map.
[0069] In an embodiment, the random texture 230 is smaller than the image to be rendered and is repeated on the image using REPEAT as a texture wrapping (warp) mode. The random texture 230 can be generated in any manner and may include pure random values, jitter, multi-jitter, orthogonal array sampling, blue noise random values, red noise random values, etc. In an embodiment, the random texture 230 is a three-dimensional random texture, where the third dimension is the number of frames or time. Therefore, the distribution of random numbers will change frame by frame, so that the adaptive sampling map 215 is averaged toward the importance map as the number of frames increases. For example, a 3D blue noise texture can be used with dimensions 256x256x128 pixels, where 256x256 is the spatial dimension and 128 is the frame (or time) dimension. For the first frame, the first 256X256 slice of the 3D texture is used by the sampling map generation unit 220 to quantize the sampling map. For the second frame, the second 256x 256 slice is used, and so on. When the last frame is reached, the sampling map generation unit 220 can repeat the sequence starting from the first slice.
[0070] In an embodiment, the sampling map generation unit 220 determines the total number of samples based on the target per-pixel sampling rate defined by the sampling control. In an embodiment, the target per-pixel sampling rate is a floating point number. The total number of samples is then distributed across the pixels according to the importance map to generate an initial sampling map. In an embodiment, the minimum number of samples is assigned to all pixels before the remainder of the total number of samples less than the product of the minimum number and the number of pixels is distributed across the pixels according to the importance map. The initial sampling map may be stored in the memory 210 as a sampling map 215. The sampling map generation unit 220 then quantizes the initial sampling map using the per-pixel random values in the random texture 230 to generate the sampling map.
[0071] In an embodiment, the initial sampling map (ism) includes a per-pixel floating-point value that specifies the target number of samples ism(x, y) for each pixel. When the random number of the pixel is less than the fractional part of ism(x, y), that is, random(x, y) <frac(ism(x, y)), the per-pixel random value between 0 and 1 is effectively used to select floor(ism(x, y)). Otherwise, the sampling rate of the pixel is quantized to ceil(ism(x, y)). The sampling map can replace the initial sampling map and be stored in the memory 210 as a sampling map 215 or as a separate sampling map. The quantized sampling map stores integer values for each pixel, and when power-of-two sampling is enabled, the integer values are constrained to powers of two.
[0072] Processor 225 receives the sampling map and 3D data of a scene rendered using the sampling map to produce a frame. In an embodiment, the scene is rendered using ray tracing, and the sampling map controls the number of rays cast for each pixel in the frame. In an embodiment, the scene is rendered using rasterization or other techniques, including blending techniques.
[0073] Figure 2B A flow chart of a method 240 for adaptive sampling according to an embodiment is shown. Each block of the method 240 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, different functions can be performed by a processor executing instructions stored in a memory. The method 240 can also be embodied as computer-usable instructions stored on a computer storage medium. To name a few, the method 240 can be provided by a stand-alone application, a service or a hosted service (standalone or in combination with another hosted service), or a plug-in for another product. In addition, by way of example, with respect to Figure 2A Method 240 is described with reference to a system. However, this method may additionally or alternatively be performed by any system or any combination of systems, including but not limited to the systems described herein. In addition, one of ordinary skill in the art will appreciate that any system that performs method 240 is within the scope and spirit of the embodiments of the present disclosure.
[0074] Method 240 includes steps 155 and 165 from method 150. After step 155, method 240 determines whether a minimum number of samples per pixel is required. The minimum number of samples per pixel may be specified by a user and / or an application. If a minimum number of samples per pixel is required at step 255, then at step 260, the minimum number of samples is allocated to each pixel included in the image. At step 265, the remainder of the total number of samples is distributed across the pixels included in the image according to the importance map to generate an initial sampling map for the image. When the minimum number of samples is zero, the remainder of the total number of samples is equal to the total number of samples. After step 265, the method proceeds to step 165 to quantize the initial sampling map.
[0075] Figure 2C A flow chart of a method 250 for adaptive sampling using a power-of-two sampling rate according to an embodiment is shown. Each block of the method 250 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, different functions can be performed by a processor executing instructions stored in a memory. The method 250 can also be embodied as computer-usable instructions stored on a computer storage medium. To name a few, the method 250 can be provided by a stand-alone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in for another product. In addition, by way of example, with respect to Figure 2A The method 250 is described with reference to a system. However, this method may be performed additionally or alternatively by any system or any combination of systems, including but not limited to the systems described herein. In addition, one of ordinary skill in the art will understand that any system that performs the method 250 is within the scope and spirit of the embodiments of the present disclosure.
[0076] Method 250 includes step 155 from method 150. At step 270, the total number of samples is distributed across pixels included in the image according to the importance map to generate an initial sampling map for the image. In an embodiment, the initial sampling rate calculated for each pixel is
[0077] At step 275, if random value quantization is not used, then at step 280, the initial sampling map is quantized using equation (4) to quantize the initial sampling rate to a power of two using rounding without any random values. Otherwise, at step 285, the initial sampling map is quantized using per-pixel random values according to equation (6) to produce a sampling map for the image that includes a power-of-two integer number of samples for each pixel in the image. For steps 280 and 285, the values of b, f, and Δ for the initial sampling rate per pixel can be calculated according to equation (5).
[0078] Figure 3AAn importance map 300 and an adaptive sampling map 320 according to an embodiment are shown. The lighter areas of the importance map 300 indicate pixels that need more samples than the darker areas of the importance map 300. The importance map 300 can be generated by feeding the variance from a denoising algorithm, or an exponentially weighted variance can be calculated over time. In an embodiment, the importance map is generated based on gaze or eye tracking data to focus on samples where the user is looking. The adaptive sampling map 320 is generated by quantizing the initial sampling map using a multi-jitter random texture and generating an integer number of samples for each pixel. In an embodiment, the adaptive sampling map 320 is generated using method 150 or 250.
[0079] Figure 3B An adaptive sampling map according to an embodiment is shown. The adaptive sampling map 350 is generated by quantizing an initial sampling map using a blue noise texture and producing an integer number of samples per pixel. In an embodiment, the adaptive sampling map 350 is generated using method 150 or 250 .
[0080] FIG3C shows a prior art sampling map generated using prior art techniques. A threshold-based sampling map 360 is generated using conventional techniques. The "staircase" distribution of samples in the threshold-based sampling map 360 contrasts with the smooth distribution of the adaptive sampling maps 320 and 350. In the threshold-based sampling map 360, black corresponds to zero samples per pixel (SPP), dark gray corresponds to 1SPP, light gray corresponds to 2SPP, and white corresponds to 3SPP. The threshold used to quantize the floating-point sample values is fixed for the entire image and for the image sequence. In contrast, using random values to quantize the floating-point sample values not only produces a smoother transition between the quantized sample values, but also better matches the per-pixel target sampling rate. In particular, when compared to the number of samples per pixel for the adaptive sampling map generated using a 3D random texture, the sampling map generated using thresholding is generally not averaged over time to match the input importance map 300.
[0081] The threshold-based sampling map 360 results in lower quality sample placement, which may severely degrade the quality of the rendered image. Additionally, the threshold-based sampling map 360 may correspond to a lower or higher number of samples per pixel than the target number of samples per pixel. A lower number of samples per pixel may result in underutilization of rendering resources. A higher number of samples per pixel may result in overutilization of rendering resources, which may result in not meeting the target frame rate. Quantizing the sample per pixel values using a random number produces an adaptive sampling map that more closely matches the target number of samples per pixel on average and over time.
[0082] Parallel processing architecture
[0083] Figure 4A parallel processing unit (PPU) 400 is shown according to an embodiment. The PPU 400 may be used to implement the adaptive sampling system 200. The PPU 400 may be used to implement one or more of the processor 225 and the sampling map generation unit 220 in the adaptive sampling system 200. The PPU 400 may be configured to perform the methods 150 and / or 250.
[0084] In one embodiment, PPU 400 is a multi-threaded processor implemented on one or more integrated circuit devices. PPU 400 is a latency-hiding architecture designed for processing many threads in parallel. A thread (i.e., an execution thread) is an instance of an instruction set configured to be executed by PPU 400. In one embodiment, PPU 400 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data to generate two-dimensional (2D) image data for display on a display device. In other embodiments, PPU 400 can be used to perform general-purpose calculations. Although an exemplary parallel processor is provided herein for illustrative purposes, it should be specifically noted that the processor is described only for illustrative purposes, and any processor can be used to supplement and / or replace the processor.
[0085] One or more PPUs 400 can be configured to accelerate thousands of high-performance computing (HPC), data center, cloud computing, and machine learning applications. PPU 400 can be configured to accelerate numerous deep learning systems and applications for autonomous vehicles, simulations, computational graphics (such as ray or path tracing), deep learning, high-precision speech, image and text recognition systems, intelligent video analysis, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analysis, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.
[0086] like Figure 4As shown, the PPU 400 includes an input / output (I / O) unit 405, a front-end unit 415, a scheduler unit 420, a work distribution unit 425, a hub 430, a crossbar switch (Xbar) 470, one or more general processing clusters (GPCs) 450, and one or more memory partition units 480. The PPU 400 can be connected to a host processor or other PPUs 400 via one or more high-speed NVLink 410 interconnects. The PPU 400 can be connected to a host processor or other peripheral devices via interconnect 402. The PPU 400 can also be connected to a local memory 404 including multiple memory devices. In one embodiment, the local memory can include multiple dynamic random access memory (DRAM) devices. The DRAM device can be configured as a high bandwidth memory (HBM) subsystem, in which multiple DRAM dies are stacked within each device.
[0087] The NVLink 410 interconnect enables the system to scale and include one or more PPUs 400 in conjunction with one or more CPUs, supporting cache coherency between the PPU 400 and the CPU, and CPU mastering. Data and / or commands may be sent by the NVLink 410 through the hub 430 to or from other units of the PPU 400, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). Figure 5B NVLink 410 is described in more detail.
[0088] I / O unit 405 is configured to send and receive communications (e.g., commands, data, etc.) from a host processor (not shown) via interconnect 402. I / O unit 405 may communicate with the host processor directly via interconnect 402, or through one or more intermediate devices (such as a memory bridge). In one embodiment, I / O unit 405 may communicate with one or more other processors (e.g., one or more PPUs 400) via interconnect 402. In one embodiment, I / O unit 405 implements a peripheral component interconnect express (PCIe) interface for communicating over a PCIe bus, and interconnect 402 is a PCIe bus. In alternative embodiments, I / O unit 405 may implement other types of known interfaces for communicating with external devices.
[0089] The I / O unit 405 decodes data packets received via the interconnect 402. In one embodiment, the data packets represent commands configured to cause the PPU 400 to perform various operations. The I / O unit 405 sends the decoded commands to various other units of the PPU 400 as specified by the commands. For example, some commands may be sent to the front end unit 415. Other commands may be sent to the hub 430 or other units of the PPU 400, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, the I / O unit 405 is configured to route communications between and among the various logical units of the PPU 400.
[0090] In one embodiment, a program executed by a host processor encodes a command stream in a buffer that provides a workload to the PPU 400 for processing. The workload may include many instructions and data to be processed by those instructions. A buffer is an area in memory that is accessible (e.g., read / write) by both the host processor and the PPU 400. For example, the I / O unit 405 may be configured to access a buffer in a system memory connected to the interconnect 402 via a memory request transmitted through the interconnect 402. In one embodiment, the host processor writes a command stream to the buffer and then sends a pointer to the start of the command stream to the PPU 400. The front end unit 415 receives a pointer to one or more command streams. The front end unit 415 manages one or more streams, reads commands from the streams, and forwards the commands to the various units of the PPU 400.
[0091] The front end unit 415 is coupled to a scheduler unit 420, which configures the various GPCs 450 to process the tasks defined by one or more streams. The scheduler unit 420 is configured to track state information related to the various tasks managed by the scheduler unit 420. The state may indicate which GPC 450 the task is assigned to, whether the task is active or inactive, a priority associated with the task, etc. The scheduler unit 420 manages the execution of multiple tasks on one or more GPCs 450.
[0092] Scheduler unit 420 is coupled to work distribution unit 425, which is configured to dispatch tasks for execution on GPC 450. Work distribution unit 425 may keep track of a number of scheduled tasks received from scheduler unit 420. In one embodiment, work distribution unit 425 manages a pending task pool and an active task pool for each GPC 450. When a GPC 450 completes execution of a task, the task is evicted from the active task pool of GPC 450, and one of the other tasks from the pending task pool is selected and scheduled for execution on GPC 450. If an active task on GPC 450 has become idle, such as while waiting for a data dependency to be resolved, then the active task may be evicted from GPC 450 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on GPC 450.
[0093] In one embodiment, the host processor executes a driver kernel that implements an application programming interface (API) that enables one or more applications to be executed on the host processor to schedule operations for execution on the PPU 400. In one embodiment, multiple computing applications are executed simultaneously by the PPU 400, and the PPU 400 provides isolation, quality of service (QoS), and independent address spaces for multiple computing applications. The application can generate instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks to be executed by the PPU 400. The driver kernel outputs the task to one or more streams being processed by the PPU 400. Each task can include one or more related thread groups, referred to herein as warps. In one embodiment, a warp includes 32 related threads that can be executed in parallel. A cooperative thread can refer to multiple threads that include instructions to execute a task and can exchange data through a shared memory. The task can be assigned to one or more processing units within the GPC 450, and the instruction is scheduled for execution by at least one warp.
[0094] Work distribution unit 425 communicates with one or more GPCs 450 via XBar 470. XBar 470 is an interconnect network that couples many of the units of PPU 400 to other units of PPU 400. For example, XBar 470 may be configured to couple work distribution unit 425 to a particular GPC 450. Although not explicitly shown, one or more other units of PPU 400 may also be connected to XBar 470 via hub 430.
[0095] Tasks are managed by a scheduling unit 420 and dispatched to GPCs 450 by a work distribution unit 425. GPCs 450 are configured to process tasks and produce results. Results may be consumed by other tasks within the GPC 450, routed to different GPCs 450 via XBar 470, or stored in memory 404. Results may be written to memory 404 via a memory partition unit 480, which implements a memory interface for reading data to and writing data from memory 404. Results may be transferred to another PPU 400 or CPU via an NV link 410. In an embodiment, a PPU 400 includes a number U of memory partition units 480, which is equal to the number of separate and distinct memory devices coupled to the memory 404 of the PPU 400. Each GPC 450 may include a memory management unit to provide virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the memory management unit provides one or more translation lookaside buffers (TLBs) for performing translations of virtual addresses to physical addresses in memory 404 .
[0096] In one embodiment, the memory partition unit 480 includes a raster operation (ROP) unit, a second level (L2) cache memory, and a memory interface coupled to the memory 404. The memory interface can implement a 32, 64, 128, 1024-bit data bus, etc., for high-speed data transmission. The PPU 400 can be connected to up to Y memory devices, such as a high bandwidth memory stack or a graphics double data rate, version 5, synchronous dynamic random access memory, or other types of persistent storage. In an embodiment, the memory interface implements an HBM2 memory interface, and Y is equal to half U. In an embodiment, the HBM2 memory stack is located on the same physical package as the PPU 400, providing significant power and area savings compared to traditional GDDR5 SDRAM systems. In an embodiment, each HBM2 stack includes four memory dies and Y is equal to 4, wherein each HBM2 stack includes two 128-bit channels per die (8 channels total) and a data bus width of 1024 bits.
[0097] In an embodiment, memory 404 supports single error correction double error detection (SECDED) error correction code (ECC) to protect data. ECC provides higher reliability for computing applications that are sensitive to data corruption. Reliability is particularly important in large-scale cluster computing environments where PPU 400 processes very large data sets and / or runs applications for extended periods of time.
[0098] In an embodiment, the PPU 400 implements a multi-level memory hierarchy. In one embodiment, the memory partition unit 480 supports unified memory to provide a single unified virtual address space for the CPU and PPU 400 memory, thereby enabling data sharing between virtual memory systems. In an embodiment, the frequency of PPU 400 accesses to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU 400 where pages are accessed more frequently. In one embodiment, NVLink 410 supports address translation services, allowing the PPU 400 to directly access the CPU's page table and providing the PPU 400 with full access to the CPU memory.
[0099] In an embodiment, the copy engine transfers data between multiple PPUs 400 or between a PPU 400 and a CPU. The copy engine may generate a page fault that is not mapped to an address in a page table. The memory partition unit 480 may then service the page fault, map the address into a page table, and thereafter the copy engine may perform the transfer. In conventional systems, for multiple copy engine operations between multiple processors, memory is pinned (e.g., non-pageable), thereby greatly reducing the available memory. In the event of a hardware page fault, the address may be passed to the copy engine without concern for whether the memory page is resident, and the copy process is transparent.
[0100] Data from memory 404 or other system memory may be retrieved by memory partition unit 480 and stored in L2 cache memory 460, which is located on chip and shared between different GPCs 450. As shown, each memory partition unit 480 includes a portion of the L2 cache memory associated with the corresponding memory 404. Lower-level caches may then be implemented in different units within GPC 450. For example, each processing unit in GPC 450 may implement a level 1 (L1) cache. The L1 cache is a private memory dedicated to a specific processing unit. The L2 cache 460 is coupled to the memory interface 470 and XBar 470, and data from the L2 cache may be retrieved and stored in each of the L1 caches for processing.
[0101] In an embodiment, the processing unit within each GPC 450 implements a SIMD (single instruction, multiple data) architecture, wherein each thread in a group of threads (e.g., a warp) is configured to process different data sets based on the same instruction set. All threads in a thread group execute the same instruction. In another embodiment, the processing unit implements a SIMT (single instruction, multiple thread) architecture, wherein each thread in a group of threads is configured to process different data sets based on the same instruction set, but wherein individual threads in a group of threads are allowed to fork during execution. In one embodiment, a program counter, call stack, and execution state are maintained for each warp, thereby achieving concurrency between the warp and the serial execution within the warp when the threads within the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby achieving equal concurrency between all threads, within the warp, and between the warps. When the execution state is maintained for each individual thread, threads executing the same instruction can be converged and executed in parallel for maximum efficiency.
[0102] Cooperative Groups is a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads are communicating, enabling the expression of richer and more efficient decompositions of parallelism. The cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. Conventional programming models provide a single simple structure for synchronizing cooperative threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, programmers often want to define thread groups at a granularity smaller than the thread block granularity and synchronize within the defined group, enabling higher performance, design flexibility, and software reuse in the form of a collective group-wide function interface.
[0103] Cooperative Groups enables programmers to explicitly define thread groups at sub-block (e.g., as small as a single thread) and multi-block granularity and perform collective operations such as synchronization on threads in a cooperative group. The programming model supports clean composition across software boundaries so that libraries and utility functions can safely synchronize in their local environment without making assumptions about convergence. Cooperative Groups primitives enable new patterns of cooperative parallelism, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire grid of thread blocks.
[0104] Each processing unit includes a large number (e.g., 128, etc.) of different processing cores (e.g., functional units), which may include fully pipelined, single precision, double precision, and / or mixed precision, and which include a floating point arithmetic logic unit and an integer arithmetic logic unit. In one embodiment, the floating point arithmetic logic unit implements the IEEE 754-2008 standard for floating point operations. In one embodiment, the core includes 64 single precision (32-bit) floating point cores, 64 integer cores, 32 double precision (64-bit) floating point cores, and 8 tensor cores.
[0105] The tensor cores are configured to perform matrix operations. Specifically, the tensor cores are configured to perform deep learning matrix operations, such as GEMM (matrix-matrix multiplication) for convolution operations during neural network training and inference. In one embodiment, each tensor core operates on a 4×4 matrix and performs a matrix multiplication and accumulation operation D=A×B+C, where A, B, C, and D are 4×4 matrices.
[0106] In an embodiment, the matrix multiplication inputs A and B may be integer, fixed-point or floating-point matrices, and the accumulation matrices C and D may be integer, fixed-point or floating-point matrices of equal or higher bit width. In an embodiment, the tensor core operates on one, four or eight-bit integer input data with 32-bit integer accumulation. 8-bit integer matrix multiplication requires 1024 operations and results in full-precision products, which are then accumulated with other intermediate products of 8x8x16 matrix multiplication using 32-bit integer addition. In one embodiment, the tensor core operates on 16-bit floating-point input data and 32-bit floating-point accumulation. 16-bit floating-point multiplication requires 64 operations, produces full-precision products, which are then accumulated using 32-bit floating-point additions with other intermediate products of 4×4×4 matrix multiplication. In practice, tensor cores are used to perform larger two-dimensional or higher-dimensional matrix operations established by these smaller elements. APIs such as the CUDA 9 C++ API expose specialized matrix load, matrix multiply and accumulate, and matrix store operations to efficiently use tensor cores from CUDA-C++ programs. At the CUDA level, the warp-level interface assumes that a 16×16 size matrix spans all 32 threads of a warp.
[0107] Each processing unit also includes M special function units (SFUs) that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, the SFU may include a tree traversal unit that is configured to traverse a hierarchical tree data structure. In one embodiment, the SFU may include a texture unit configured to perform texture mapping filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texture pixels) from memory 404 and sample the texture map to generate sampled texture values for use in a shader program executed by the processing unit. In one embodiment, the texture map is stored in a shared memory that may include or contain an L1 cache. The texture unit implements texture operations, such as filtering operations using mip maps (i.e., texture maps of different levels of detail). In one embodiment, each processing unit includes two texture units.
[0108] Each processing unit also includes N access units (LSUs) that implement load and store operations between the shared memory and the register file. Each processing unit includes an interconnect network that connects each core to the register file and connects the LSU to the register file and the shared memory. In one embodiment, the interconnect network is a crossbar switch that can be configured to connect any core to any register in the register file and to connect the LSU to the register file and the memory location in the shared memory.
[0109] Shared memory is an on-chip memory array that allows data storage and communication between processing units and between threads in a processing unit. In one embodiment, shared memory includes 128KB of storage capacity and is in the path from each processing unit to memory partition unit 480. Shared memory can be used for cache reads and writes. One or more of shared memory, L1 cache, L2 cache, and memory 404 is a backing store.
[0110] Combining data cache and shared memory functionality into a single memory block provides the best overall performance for both types of memory accesses. This capacity can be used by programs as a cache that does not use the shared memory. For example, if the shared memory is configured to use half of the capacity, texture and load / store operations can use the remaining capacity. The integration within the shared memory enables the shared memory to function as a high-throughput pipeline for streaming data, while providing high-bandwidth and low-latency access to frequently reused data.
[0111] When configured for general parallel computing, a simpler configuration can be used compared to graphics processing. Specifically, the fixed function graphics processing unit is bypassed, creating a simpler programming model. In the general parallel computing configuration, the work distribution unit 425 assigns and distributes thread blocks directly to the processing units within the GPC 450. The threads execute the same program, use unique thread IDs in the calculation to ensure that each thread generates a unique result, use the processing unit to execute the program and perform the calculation, use shared memory to communicate between threads, and use LSUs to read and write global memory through shared memory and memory partition unit 480. When configured for general parallel computing, the processing unit can also write commands that the scheduler unit 420 can use to start new work on the processing unit.
[0112] The PPU 400 may each include and / or be configured to perform the following functions, one or more processing cores and / or components thereof, such as a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a ray tracing (RT) core, a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.
[0113] The PPU 300 may be included in a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., wireless, handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In one embodiment, the PPU 300 is included on a single semiconductor substrate. In another embodiment, the PPU 300 is included on a system on a chip (SoC) along with one or more other devices (such as an additional PPU 300, a memory 304, a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog converter (DAC), etc.).
[0114] In one embodiment, PPU 300 may be included on a graphics card that includes one or more memory devices 304. The graphics card may be configured to interface with a PCIe slot on a motherboard of a desktop computer. In yet another embodiment, PPU 300 may be an integrated graphics processing unit (iGPU) or parallel processor included in a chipset of the motherboard. In yet another embodiment, PPU 400 may be implemented in reconfigurable hardware. In yet another embodiment, portions of PPU 400 may be implemented in reconfigurable hardware.
[0115] Exemplary Computing System
[0116] Systems with multiple GPUs and CPUs are used in a variety of industries as developers expose and exploit more parallelism in applications such as artificial intelligence computing. High-performance GPU-accelerated systems with tens to thousands of computing nodes are deployed in data centers, research institutions, and supercomputers to solve larger problems. As the number of processing devices within high-performance systems increases, communication and data transmission mechanisms need to scale to support this increased bandwidth.
[0117] Figure 5A According to one embodiment, the use Figure 4 A conceptual diagram of a processing system 500 implemented by a PPU 400. An exemplary system 565 may be configured to implement Figure 1F The method 150 shown in Figure 2B The method 240 shown in and / or Figure 2C The processing system 500 includes a CPU 530 , a switch 510 , and a plurality of PPUs 400 and corresponding memories 404 .
[0118] NVLink 410 provides a high-speed communication link between each PPU 400. Figure 5B 402 connections, but the number of connections connected to each PPU 400 and CPU 530 may vary. Switch 510 interfaces between interconnect 402 and CPU 530. PPU 400, memory 404, and NVLink 410 may be located on a single semiconductor platform to form a parallel processing module 525. In one embodiment, switch 510 supports two or more protocols that interface between various different connections and / or links.
[0119] In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between each PPU 400 and CPU 530, and switch 510 interfaces between interconnect 402 and each PPU 400. PPU 400, memory 404, and interconnect 402 may be located on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), interconnect 402 provides one or more communication links between each PPU 400 and CPU 530, and switch 510 interfaces between each PPU 400 using NVLink 410 to provide one or more high-speed communication links between PPUs 400. In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between PPU 400 and CPU 530 through switch 510. In yet another embodiment (not shown), interconnect 402 provides one or more communication links directly between each PPU 400. One or more NVLink 410 high-speed communication links may be implemented as a physical NVLink interconnect or as an on-chip or on-die interconnect using the same protocol as NVLink 410.
[0120] In the context of this specification, a single semiconductor platform may refer to a unique single semiconductor-based integrated circuit manufactured on a bare die or chip. It should be noted that the term single semiconductor platform may also refer to a multi-chip module with increased connectivity that simulates on-chip operations and is substantially improved by utilizing conventional bus implementations. Of course, various circuits or devices may also be placed separately or in various combinations of semiconductor platforms, depending on the needs of the user. Optionally, the parallel processing module 525 may be implemented as a circuit board substrate, and each of the PPU 400 and / or memory 404 may be a packaged device. In one embodiment, the CPU 530, switch 510, and parallel processing module 525 are located on a single semiconductor platform.
[0121] In one embodiment, the signaling rate of each NVLink 410 is 20 to 25 Gbit / s, and each PPU 400 includes six NVLink 410 interfaces (e.g., Figure 5A As shown, each PPU 400 includes five NVLink 410 interfaces. Each NVLink 410 provides a data transfer rate of 25 Gbit / s in each direction, with six links providing 400 Gbit / s. When the CPU 530 also includes one or more NVLink 410 interfaces, the NVLink 410 can be used exclusively for example Figure 5A PPU to PPU communication as shown, or some combination of PPU to PPU and PPU to CPU.
[0122] In one embodiment, NVLink 410 allows direct load / store / atomic access from CPU 530 to memory 404 of each PPU 400. In one embodiment, NVLink 410 supports coherency operations, allowing data read from memory 404 to be stored in the cache hierarchy of CPU 530, reducing cache access latency of CPU 530. In one embodiment, NVLink 410 includes support for address translation services (ATS), allowing PPU 400 to directly access page tables within CPU 530. One or more NVLinks 410 may also be configured to operate in a low power mode.
[0123] Figure 5B An exemplary system 565 is shown in which various architectures and / or functions of various previous embodiments may be implemented. The exemplary system 565 may be configured to implement Figure 1F The method shown in 150, Figure 2B The method 240 shown in and / or Figure 2C The method 250 shown in .
[0124] As shown, a system 565 is provided, which includes at least one central processing unit 530 connected to a communication bus 575. The communication bus 575 can directly or indirectly couple one or more of the following devices: main memory 540, network interface 535, one or more CPUs 530, one or more display devices 545, one or more input devices 560, switch 510, and parallel processing system 525. The communication bus 575 can be implemented using any suitable protocol and can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The communication bus 575 may include one or more bus or link types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standard association (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect high speed (PCIe) bus, a hypertransport, and / or another type of bus or link. In some embodiments, there is a direct connection between the components. For example, the CPU 530 can be directly connected to the main memory 540. Further, the CPU 530 can be directly connected to the parallel processing system 525. In the case where there is a direct or point-to-point connection between components, the communication bus 575 may include a PCIe link for performing the connection. In these examples, the PCI bus need not be included in the system 565.
[0125] although Figure 5CThe various blocks of are shown as being connected with wires via communication bus 575, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, presentation components such as display device 545 may be considered to be I / O components such as input device 560 (e.g., if the display is a touch screen). As another example, CPU 530 and / or parallel processing system 525 may include memory (e.g., main memory 540 may represent a storage device in addition to parallel processing system 525, CPU 530, and / or other components). In other words, Figure 5C The computing devices described herein are illustrative only. No distinction is made between such categories as "workstations," "servers," "laptops," "desktop computers," "tablet computers," "client devices," "mobile devices," "handheld devices," "game consoles," "electronic control units (ECUs)," "virtual reality systems," and / or other device or system types, as all are considered Figure 5C within the range of computing devices.
[0126] The system 565 also includes a main memory 540. Control logic (software) and data are stored in the main memory 540, which can take the form of various computer-readable media. Computer-readable media can be any available media that can be accessed by the system 565. Computer-readable media can include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media.
[0127] Computer storage media may include both volatile and nonvolatile media and / or both removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, main memory 540 may store computer readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by system 565. As used herein, computer storage media does not include signals themselves.
[0128] Computer storage media may embody computer readable instructions, data structures, program modules, and / or other data types as a modulated data signal (such as a carrier wave or other transmission mechanism) and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in a manner that encodes the information in the signal. By way of example and not limitation, computer storage media may include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared and other wireless media). Combinations of any of the above should also be included within the scope of computer readable media.
[0129] The computer program enables the system 565 to perform different functions when executed. The CPU 530 may be configured to execute at least some of the computer-readable instructions to control one or more components of the system 565 to perform one or more of the methods and / or processes described herein. The CPU 530 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing multiple software threads simultaneously. The CPU 530 may include any type of processor, and may include different types of processors (e.g., a processor with fewer cores for mobile devices and a processor with more cores for servers) depending on the type of system 565 implemented. For example, depending on the type of system 565, the processor may be an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (such as math coprocessors), the system 565 may also include one or more CPUs 530.
[0130] In addition to or in lieu of the CPU 530, the parallel processing module 525 may be configured to execute at least some of the computer-readable instructions to control one or more components of the system 565 to perform one or more of the methods and / or processes described herein. The parallel processing module 525 may be used by the system 565 to render graphics (e.g., 3D graphics) or perform general purpose computations. For example, the parallel processing module 525 may be used for general purpose computations on a GPU (GPGPU). In embodiments, one or more CPUs 530 and / or parallel processing modules 525 may perform any combination of methods, processes, and / or portions thereof, either discretely or in combination.
[0131] The system 565 also includes one or more input devices 560, a parallel processing system 525, and one or more display devices 545. The one or more display devices 545 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more display devices 545 may receive data from other components (e.g., the parallel processing system 525, the CPU 530, etc.) and output the data (e.g., as an image, video, sound, etc.).
[0132] The network interface 535 can enable the system 565 to be logically coupled to other devices, including input devices 560, one or more display devices 545, and / or other components, some of which can be built into (e.g., integrated into) the system 565. Illustrative input devices 560 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. The input device 560 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and near the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of the system 565. The system 565 may include a depth camera for gesture detection and recognition, such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations of these. Additionally, system 565 may include an accelerometer or gyroscope that enables detection of motion (e.g., as part of an inertial measurement unit (IMU)). In some examples, system 565 may use the output of the accelerometer or gyroscope to render an immersive augmented reality or virtual reality.
[0133] Further, the system 565 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) (such as the Internet), a peer-to-peer network, a cable network, etc.) for communication purposes through the network interface 535. The system 565 can be included in a distributed network and / or cloud computing environment.
[0134] The network interface 535 may include one or more receivers, transmitters, and / or transceivers that enable the system 565 to communicate with other computing devices via an electronic communications network (including wired and / or wireless communications). The network interface 535 may include components and functionality for enabling communications via any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., via Ethernet or InfiniBand communications), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0135] The system 565 may also include an auxiliary storage device (not shown). The auxiliary storage device includes, for example, a hard disk drive and / or a removable storage drive, representing a floppy disk drive, a tape drive, a compact disk drive, a digital versatile disk (DVD) drive, a recording device, a universal serial bus (USB) flash memory. The removable storage drive reads from and / or writes to a removable storage unit in a known manner. The system 565 may also include a hardwired power supply, a battery power supply, or a combination thereof (not shown). The power supply can provide power to the system 565 to enable the components of the system 565 to operate.
[0136] Each of the aforementioned modules and / or devices may even be located on a single semiconductor platform to form system 565. Alternatively, different modules may also be located individually or in different combinations of semiconductor platforms according to the needs of the user. Although various embodiments have been described above, it should be understood that these embodiments are presented by way of example only and not limitation. Thus, the breadth and scope of the preferred embodiments should not be limited by any of the above exemplary embodiments, but should only be defined in accordance with the appended claims and their equivalents.
[0137] Example network environment
[0138] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be configured to communicate with the client device, server, or other device types. Figure 5A The processing system 500 and / or Figure 5B The exemplary system 565 may be implemented on one or more instances of the exemplary system 565 of the processing system 500—for example, each device may include similar components, features, and / or functionality of the exemplary system 565 and / or the processing system 500.
[0139] The components of the network environment can communicate with each other via one or more networks, which can be wired, wireless, or both. The network can include multiple networks or one of multiple networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or a public switched telephone network (PSTN), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connections.
[0140] Compatible network environments may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to the server may be implemented on any number of client devices.
[0141] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, the servers may include one or more core network servers and / or edge servers. The framework layer may include a framework for software supporting the software layer and / or one or more applications of the application layer. The software or application may include network-based service software or application programs, respectively. In an embodiment, one or more client devices may use network-based service software or applications (e.g., by accessing service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, the type of free and open source software network application framework, such as a distributed file system that may be used for large-scale data processing (e.g., "big data").
[0142] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across a state, region, country, globe, etc.). If the connection to the user (e.g., client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0143] Client devices may include Figure 5B The example processing system 500 and / or Figure 5C At least some of the components, features, and functions of the exemplary system 565 of the embodiment of the present invention. By way of example and not limitation, the client device may be implemented as a personal computer (PC), a laptop computer, a mobile device, a smart phone, a tablet computer, a smart watch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a boat, a spacecraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of the depicted devices, or any other suitable device.
[0144] Machine Learning
[0145] Deep Neural Networks (DNNs) developed on processors such as the PPU 400 have been used in a variety of use cases: from self-driving cars to faster drug development, from automatic image captioning in online image databases to intelligent real-time language translation in video chat applications. Deep learning is a technology that models the neural learning process of the human brain, constantly learning, constantly getting smarter, and delivering more accurate results faster over time. A child is initially taught by an adult to correctly identify and classify various shapes, and eventually is able to recognize shapes without any coaching. Similarly, deep learning or neural learning systems need to be trained in object recognition and classification in order to become smarter and more efficient in recognizing basic objects, occluded objects, etc., while also assigning context to objects.
[0146] At the simplest level, neurons in the human brain look at the various inputs they receive, assign a level of importance to each of these inputs, and pass outputs to other neurons for processing. An artificial neuron, or perceptron, is the most basic model of a neural network. In one example, a perceptron can receive one or more inputs that represent various features of an object that the perceptron is being trained to recognize and classify, and each of these features is given a certain weight based on the importance of that feature when defining the shape of the object.
[0147] Deep neural network (DNN) models include multiple layers of many connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.), which can be trained with large amounts of input data to solve complex problems quickly and accurately. In one example, the first layer of a DNN model breaks down an input image of a car into its parts and looks for basic patterns (such as lines and angles). The second layer assembles lines to find higher-level patterns, such as wheels, windshields, and mirrors. The next layer identifies the type of vehicle, and the last few layers generate labels for the input image, identifying the model of a specific car brand.
[0148] Once a DNN is trained, it can be deployed and used to recognize and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include recognizing handwritten numbers on a check deposited into an ATM, identifying images of friends in photos, providing movie recommendations to over 50 million users, recognizing and classifying different types of cars, pedestrians, and road hazards in a self-driving car, or translating human speech in real time.
[0149] During training, data flows through the DNN in a forward propagation phase until a prediction is produced, which indicates a label corresponding to the input. If the neural network does not correctly label an input, the error between the correct label and the predicted label is analyzed, and the weights are adjusted for each feature during a backward propagation phase until the DNN correctly labels that input and other inputs in the training data set. Training complex neural networks requires a large amount of parallel computing performance, including floating-point multiplications and additions supported by the PPU 400. Inference is less computationally intensive than training and is a latency-sensitive process in which a trained neural network is applied to new inputs it has not seen before to classify images, detect emotions, identify recommendations, recognize and translate speech, and generally reason about new information.
[0150] Neural networks rely heavily on matrix math operations, and complex multi-layer networks require a lot of floating point performance and bandwidth for efficiency and speed. With thousands of processing cores optimized for matrix math operations and delivering tens to hundreds of TFLOPS of performance, the PPU 400 is a computing platform capable of delivering the performance required for deep neural network-based artificial intelligence and machine learning applications.
[0151] In addition, images generated using one or more of the techniques disclosed herein can be used to train, test, or certify DNNs for recognizing objects and environments in the real world. Such images may contain scenes of roads, factories, buildings, urban environments, rural environments, people, animals, and any other physical objects or real-world environments. Such images can be used to train, test, or certify DNNs used in machines or robots to manipulate, process, or modify physical objects in the real world. In addition, such images can be used to train, test, or certify DNNs used in autonomous vehicles to navigate and move vehicles in the real world. In addition, images generated using one or more of the techniques disclosed herein can be used to convey information to users of such machines, robots, and vehicles.
[0152] Figure 5C Components of an exemplary system 555 that may be used to train and utilize machine learning in accordance with at least one embodiment are shown. As will be discussed, the various components may be provided by different combinations of computing devices and resources or a single computing system, which may be under the control of a single entity or multiple entities. Further, various aspects may be triggered, initiated, or requested by different entities. In at least one embodiment, the training of the neural network may be directed by a provider associated with the provider environment 506, while in at least one embodiment, the training may be requested by a customer or other user accessing the provider environment through a client device 502 or other such resource. In at least one embodiment, the training data (or data to be analyzed by the trained neural network) may be provided by a provider, a user, or a third-party content provider 524. In at least one embodiment, the client device 502 may be a vehicle or object navigating on behalf of a user, for example, which may submit requests and / or receive instructions to assist the device in navigating.
[0153] In at least one embodiment, the request can be submitted across at least one network 504 to be received by the provider environment 506. In at least one embodiment, the client device can be any suitable electronic and / or computing device that enables a user to generate and send such a request, such as, but not limited to, a desktop computer, a notebook computer, a computer server, a smart phone, a tablet computer, a game console (portable or otherwise), a computer processor, computing logic, and a set-top box. The network 504 may include any suitable network for transmitting the request or other such data, such as an ad hoc network that may include the Internet, an intranet, an Ethernet network, a cellular network, a local area network (LAN), a wide area network (WAN), a personal area network (PAN), a direct wireless connection between peers, etc.
[0154] In at least one embodiment, the request may be received at the interface layer 508, which in this example may forward the data to the training and reasoning manager 532. The training and reasoning manager 532 may be a system or service including hardware and software for managing requests and services corresponding to data or content, and in at least one embodiment, the training and reasoning manager 532 may receive a request to train a neural network and may provide the requested data to the training module 512. In at least one embodiment, the training module 512 may select an appropriate model or neural network to use (if the request is not specified), and may use the relevant training data to train the model. In at least one embodiment, the training data may be a batch of data stored in the training data repository 514, received from the client device 502, or obtained from the third party provider 524. In at least one embodiment, the training module 512 may be responsible for the training data. The neural network may be any appropriate network, such as a recurrent neural network (RNN) or a convolutional neural network (CNN). Once the neural network is trained and successfully evaluated, the trained neural network may be stored in, for example, a model repository 516 that may store different models or networks for users, applications, or services, etc. In at least one embodiment, there may be multiple models for a single application or entity, as may be utilized based on a number of different factors.
[0155] In at least one embodiment, at a later point in time, a request for content (e.g., path determination) or data determined or influenced at least in part by a trained neural network may be received from the client device 502 (or another such device). This request may include, for example, input data to be processed using the neural network to obtain one or more inferences or other output values, classifications, or predictions, or for at least one embodiment, the input data may be received by the interface layer 508 and directed to the inference module 518, although different systems or services may also be used. In at least one embodiment, if not already locally stored to the inference module 518, the inference module 518 may obtain an appropriate trained network from the model repository 516, such as a trained deep neural network (DNN) as discussed herein. The inference module 518 may provide data as input to the trained network, which may then generate one or more inferences as output. This may include, for example, a classification of an instance of the input data. In at least one embodiment, the inferences may then be transmitted to the client device 502 for display or other communication to a user. In at least one embodiment, the user's contextual data may also be stored in a user contextual data repository 522, which may include data about the user, which may be used as input to the network when generating inferences or determining data to be returned to the user after obtaining an instance. In at least one embodiment, related data that may include at least some of the input or inference data may also be stored in a local database 534 for processing future requests. In at least one embodiment, a user may use account information or other information to access resources or functions of a provider environment. In at least one embodiment, if permitted and available, user data may also be collected and used to further train the model to provide more accurate inferences for future requests. In at least one embodiment, a request for a machine learning application 526 executed on a client device 502 may be received through a user interface, and the results may be displayed through the same interface. The client device may include resources (such as a processor 528 and a memory 562) for generating requests and processing results or responses, and at least one data storage element 552 for storing data of the machine learning application 526.
[0156] In at least one embodiment, the processor 528 (or the processor of the training module 512 or the reasoning module 518) will be a central processing unit (CPU). However, as mentioned, resources in such an environment may utilize GPUs to process data for at least certain types of requests. With thousands of cores, GPUs (such as PPU 300) are designed to handle substantially parallel workloads, and have therefore become popular in deep learning for training neural networks and generating predictions. Although the use of GPUs for offline construction has enabled faster training of larger and more complex models, generating predictions offline means that input features cannot be used when requested, or predictions must be generated for all permutations of features and stored in lookup tables in order to serve real-time requests. If the deep learning framework supports CPU mode and the model is small and simple enough to perform feedforward on the CPU with reasonable latency, the service on the CPU instance can host the model. In this case, training can be done offline on the GPU, and reasoning can be done in real time on the CPU. If the CPU approach is not feasible, the service can be run on a GPU instance. However, because GPUs have different performance and cost characteristics than CPUs, running a service that offloads runtime algorithms to the GPU may require designing the GPU differently from a CPU-based service.
[0157] In at least one embodiment, video data may be provided from client device 502 for enhancement in provider environment 506. In at least one embodiment, video data may be processed for enhancement on client device 502. In at least one embodiment, video data may be streamed from third party content provider 524 and enhanced by third party content provider 524, provider environment 506, or client device 502. In at least one embodiment, video data may be provided from client device 502 for use as training data in provider environment 506.
[0158] In at least one embodiment, supervised and / or unsupervised training can be performed by the client device 502 and / or the provider environment 506. In at least one embodiment, a set of training data 514 (e.g., classified or labeled data) is provided as input to be used as training data. In at least one embodiment, the training data may include instances of at least one type of object for which the neural network is to be trained, and information identifying the type of object. In at least one embodiment, the training data may include a set of images each including a representation of a type of object, wherein each image also includes or is associated with a label, metadata, classification, or other information identifying the type of object represented in the corresponding image. Various other types of data may also be used as training data, and may include text data, audio data, video data, etc. In at least one embodiment, the training data 514 is provided to the training module 512 as training input. In at least one embodiment, the training module 512 may be a system or service including hardware and software, such as one or more computing devices that execute a training application for training a neural network (or other model or algorithm, etc.). In at least one embodiment, the training module 512 receives an instruction or request indicating the type of model to be used for training. In at least one embodiment, the model can be any appropriate statistical model, network or algorithm for such a purpose, such as artificial neural networks, deep learning algorithms, learning classifiers, Bayesian networks, etc. In at least one embodiment, the training module 512 can select an initial model or other untrained model from an appropriate repository 516 and train the model using training data 514 to generate a trained model (e.g., a trained deep neural network) that can be used to classify similar types of data, or generate other such inferences. In at least one embodiment in which training data is not used, an appropriate initial model can still be selected for training on the input data of each training module 512.
[0159] In at least one embodiment, the model can be trained in a variety of different ways, as may depend in part on the type of model selected. In at least one embodiment, a machine learning algorithm may be provided with a training data set, where the model is a model artifact created by the training process. In at least one embodiment, each instance of the training data contains a correct answer (e.g., a classification), which may be referred to as a target or target attribute. In at least one embodiment, the learning algorithm finds patterns in the training data that map input data attributes to targets, the answers to be predicted, and outputs a machine learning model that captures these patterns. In at least one embodiment, the machine learning model can then be used to obtain predictions about new data for which a target is not specified.
[0160] In at least one embodiment, the training and inference manager 532 can select from a set of machine learning models including binary classification, multi-class classification, generative, and regression models. In at least one embodiment, the type of model to be used may depend at least in part on the type of target to be predicted.
[0161] Graphics Processing Pipeline
[0162] In one embodiment, PPU 400 includes a graphics processing unit (GPU). PPU 400 is configured to receive commands specifying a shader for processing graphics data. Graphics data may be defined as a set of primitives, such as points, lines, triangles, quadrilaterals, triangle strips, etc. Typically, a primitive includes data specifying a plurality of vertices of the primitive (e.g., in a model space coordinate system) and attributes associated with each vertex of the primitive. PPU 400 may be configured to process the primitives to generate a frame buffer (e.g., pixel data for each of the pixels of a display).
[0163] The application writes the model data (e.g., a collection of vertices and attributes) of the scene into a memory (such as system memory or memory 404). The model data defines each of the objects that may be visible on the display. The application then makes an API call to the driver kernel, which requests the model data to be rendered and displayed. The driver kernel reads the model data and writes commands to one or more streams to perform operations to process the model data. These commands may refer to different shading programs to be implemented on the processing units within the PPU 400, including one or more of vertex shading, hull shading, domain shading, geometry shading, and pixel shading. For example, one or more of the processing units may be configured to execute a vertex shading program that processes multiple vertices defined by the model data. In one embodiment, different processing units may be configured to execute different shading programs simultaneously. For example, a first subset of processing units may be configured to execute a vertex shading program, and a second subset of processing units may be configured to execute a pixel shading program. The first subset of processing units processes the vertex data to generate processed vertex data, and writes the processed vertex data to the L2 cache 460 and / or the memory 404. After the processed vertex data is rasterized (e.g., converted from three-dimensional data to two-dimensional data in screen space) to generate fragment data, a second subset of processing units performs pixel shading to generate processed fragment data, which is then blended with other processed fragment data and written to a frame buffer in memory 404. Vertex shading programs and pixel shading programs can be executed simultaneously, processing different data from the same scene in a pipelined manner until all model data for the scene has been rendered to the frame buffer. The contents of the frame buffer are then transmitted to a display controller for display on a display device.
[0164] Fig. 6A According to one embodiment, Figure 4 4. A conceptual diagram of a graphics processing pipeline 600 implemented by PPU 400 of FIG. 4. Graphics processing pipeline 600 is an abstract flow chart of processing steps implemented to generate a 2D computer-generated image from 3D geometric data. As is well known, pipeline architectures can perform long latency operations more efficiently by breaking the operations into multiple stages, where the output of each stage is coupled to the input of the next consecutive stage. Thus, graphics processing pipeline 600 receives input data 601 that is passed from one stage of graphics processing pipeline 600 to the next stage to generate output data 602. In one embodiment, graphics processing pipeline 600 may represent a graphics processing pipeline composed of API-defined graphics processing pipeline. Alternatively, graphics processing pipeline 600 can be implemented in the functional and architectural context of the previous figures and / or one or more of any subsequent figures.
[0165] like Fig. 6A As shown, the graphics processing pipeline 600 includes a pipeline architecture including multiple stages. These stages include, but are not limited to, a data assembly stage 610, a vertex shading stage 620, a primitive assembly stage 630, a geometry shading stage 640, a viewport scale, cull, and clip (VSCC) stage 650, a rasterization stage 660, a fragment shading stage 670, and a raster operation stage 680. In one embodiment, input data 601 includes commands that configure a processing unit to implement the stages of the graphics processing pipeline 600 and configure geometric primitives (e.g., points, lines, triangles, quadrilaterals, triangle strips or fans, etc.) to be processed by these stages. Output data 602 may include pixel data (i.e., color data), which is copied to a frame buffer or other type of surface data structure in memory.
[0166] The data assembly stage 610 receives input data 601, which specifies vertex data for high-order surfaces, primitives, etc. The data assembly stage 610 collects the vertex data in temporary storage or queues, such as by receiving a command from a host processor including a pointer to a buffer in memory and reading the vertex data from the buffer. The vertex data is then passed to the vertex shading stage 620 for processing.
[0167] The vertex shading stage 620 processes vertex data by executing a set of operations (e.g., a vertex shader or program) once for each vertex. A vertex may be specified, for example, as a 4-coordinate vector (e.g.,<x,y,z,w> ). The vertex shading stage 620 can manipulate various vertex attributes, such as position, color, texture coordinates, etc. In other words, the vertex shading stage 620 performs operations on vertex coordinates or other vertex attributes associated with the vertex. These operations typically include lighting operations (e.g., modifying the color attribute of the vertex) and transformation operations (e.g., modifying the coordinate space of the vertex). For example, a vertex can be specified using coordinates in an object coordinate space, which is transformed by multiplying the coordinates by a matrix that converts the coordinates from the object coordinate space to world space or normalized-device-coordinate (NCD) space. The vertex shading stage 620 generates transformed vertex data that is transmitted to the primitive assembly stage 630.
[0168] The primitive assembly stage 630 collects the vertices output by the vertex shading stage 620 and groups the vertices into geometric primitives for processing by the geometry shading stage 640. For example, the primitive assembly stage 630 may be configured to group every three consecutive vertices into geometric primitives (e.g., triangles) for transmission to the geometry shading stage 640. In some embodiments, particular vertices may be reused for consecutive geometric primitives (e.g., two consecutive triangles in a triangle strip may share two vertices). The primitive assembly stage 630 transmits the geometric primitives (e.g., a collection of associated vertices) to the geometry shading stage 640.
[0169] The geometry shading stage 640 processes geometric primitives by performing a set of operations (e.g., geometry shaders or programs) on the geometric primitives. A tessellation operation can generate one or more geometric primitives from each geometric primitive. In other words, the geometry shading stage 640 can subdivide each geometric primitive into a finer grid of two or more geometric primitives for processing by the rest of the graphics processing pipeline 600. The geometry shading stage 640 transmits the geometric primitives to the viewport SCC stage 650.
[0170] In one embodiment, the graphics processing pipeline 600 can operate within a streaming multiprocessor and vertex shading stage 620, primitive assembly stage 630, geometry shading stage 640, fragment shading stage 670 and / or hardware / software associated therewith, and can perform processing operations sequentially. Once the sequential processing operations are completed, in one embodiment, the viewport SCC stage 650 can utilize the data. In one embodiment, primitive data processed by one or more stages in the graphics processing pipeline 600 can be written to a cache (e.g., an L1 cache, a vertex cache, etc.). In this case, in one embodiment, the viewport SCC stage 650 can access the data in the cache. In one embodiment, the viewport SCC stage 650 and the rasterization stage 660 are implemented as fixed function circuits.
[0171] The viewport SCC stage 650 performs viewport scaling, culling, and clipping of geometric primitives. Each surface being rendered is associated with an abstract camera position. The camera position represents the position of the viewer who is viewing the scene and defines a view cone that surrounds the objects of the scene. The view cone may include a viewing plane, a back plane, and four clipping planes. Any geometric primitives that are completely outside the view cone may be culled (e.g., discarded) because they will not contribute to the final rendered scene. Any geometric primitives that are partially within the view cone and partially outside the view cone may be clipped (e.g., converted to new geometric primitives that are enclosed within the view cone). In addition, each geometric primitive may be scaled based on the depth of the view cone. All potentially visible geometric primitives are then transferred to the rasterization stage 660.
[0172] The rasterization stage 660 converts 3D geometric primitives into 2D fragments (e.g., capable of being used for display, etc.). The rasterization stage 660 can be configured to use the vertices of the geometric primitives to set a set of plane equations from which various attributes can be interpolated. The rasterization stage 660 can also calculate a coverage mask for multiple pixels, which indicates whether one or more sample positions of a pixel intercept the geometric primitive. In one embodiment, a z test can also be performed to determine whether the geometric primitive is occluded by other geometric primitives that have been rasterized. The rasterization stage 660 generates fragment data (e.g., interpolated vertex attributes associated with a specific sample position of each covered pixel), which is transmitted to the fragment shading stage 670.
[0173] The fragment shading stage 670 processes the fragment data by executing a set of operations (e.g., a fragment shader or program) on each of the fragments. The fragment shading stage 670 may generate pixel data (e.g., color values) for the fragments, such as by performing lighting operations or sampling texture maps using interpolated texture coordinates of the fragments. The fragment shading stage 670 generates pixel data, which is sent to the raster operations stage 680.
[0174] Raster operations stage 680 may perform various operations on the pixel data, such as performing alpha tests, stencil tests, and blending the pixel data with other pixel data corresponding to other fragments associated with the pixel. When raster operations stage 680 has completed processing the pixel data (e.g., output data 602), the pixel data may be written to a rendering object, such as a frame buffer, a color buffer, etc.
[0175] It should be appreciated that one or more additional stages may be included in the graphics processing pipeline 600 in addition to or in place of one or more of the above stages. Various implementations of the abstract graphics processing pipeline may implement different stages. Furthermore, in some embodiments, one or more of the above stages may be excluded from the graphics processing pipeline (such as the geometry shading stage 640). Other types of graphics processing pipelines are considered to be within the scope of the present disclosure. Furthermore, any stage of the graphics processing pipeline 600 may be implemented by one or more dedicated hardware units within a graphics processor (such as PPU 400). Other stages of the graphics processing pipeline 600 may be implemented by programmable hardware units (such as processing units within PPU 400).
[0176] The graphics processing pipeline 600 may be implemented via an application program executed by a host processor (such as a CPU). In one embodiment, a device driver may implement an application programming interface (API) that defines various functions that may be utilized by an application program to generate graphics data for display. A device driver is a software program that includes a plurality of instructions that control the operation of the PPU 400. The API provides an abstraction for programmers that allows programmers to utilize dedicated graphics hardware (such as the PPU 400) to generate graphics data without requiring the programmer to utilize a specific instruction set of the PPU 400. An application may include an API call that is routed to a device driver of the PPU 400. The device driver interprets the API call and performs various operations in response to the API call. In some cases, the device driver may perform operations by executing instructions on the CPU. In other cases, the device driver may perform operations at least in part by initiating operations on the PPU 400 using an input / output interface between the CPU and the PPU 400. In one embodiment, the device driver is configured to implement the graphics processing pipeline 600 using the hardware of the PPU 400.
[0177] Various programs may be executed within the PPU 400 to implement the various stages of the graphics processing pipeline 600. For example, a device driver may launch a kernel on the PPU 400 to perform the vertex shading stage 620 on one processing unit (or multiple processing units). The device driver (or the initial kernel executed by the PPU 400) may also launch other kernels on the PPU 400 to perform other stages of the graphics processing pipeline 600, such as the geometry shading stage 640 and the fragment shading stage 670. In addition, some of the stages of the graphics processing pipeline 600 may be implemented on fixed unit hardware, such as a rasterizer or data assembler implemented within the PPU 400. It should be appreciated that the results from one kernel may be processed by one or more intermediate fixed function hardware units before being processed by a subsequent kernel on a processing unit.
[0178] The image generated by applying one or more of the technologies disclosed herein can be displayed on a monitor or other display device. In some embodiments, the display device may be directly coupled to a system or processor that generates or renders an image. In other embodiments, the display device may be indirectly coupled to the system or processor, for example, via a network. Examples of such networks include the Internet, mobile telecommunications networks, WIFI networks, and any other wired and / or wireless networking systems. When the display device is indirectly coupled, the image generated by the system or processor may be streamed to the display device via a network. Such streaming allows, for example, video games or other applications that render images to be executed on a server, in a data center, or in a cloud-based computing environment, and the rendered image will be transmitted and displayed on one or more user devices (such as computers, video game consoles, smart phones, other mobile devices, etc.) that are physically separated from the server or data center. Therefore, the technology disclosed herein can be applied to enhanced streaming images and services that enhance streaming images, such as NVIDIA GeForceNow (GFN), Google Stadia, etc.
[0179] Example streaming system
[0180] Figure 6B is an example system diagram of a streaming system 605 according to some embodiments of the present disclosure. In one embodiment, the streaming system 605 is a game streaming system. Figure 6B One or more servers 603 (which may include Figure 5A The example processing system 500 and / or Figure 5B ), one or more client devices 604 (which may include components, features, and / or functions similar to the exemplary system 565 of Figure 5A The example processing system 500 and / or Figure 5B) and one or more networks 606 (which may be similar to one or more networks described herein). In some embodiments of the present disclosure, system 605 may be implemented.
[0181] In the system 605, for a game session, one or more client devices 604 may receive input data only in response to input to one or more input devices, transmit the input data to one or more servers 603, receive encoded display data from one or more servers 603, and display the display data on a display 624. In this way, more computationally intensive calculations and processing are offloaded to one or more servers 603 (e.g., rendering, specifically ray or path tracing, for one or more GPUs of one or more servers 603 to perform graphics output for the game session). In other words, the game session is streamed from one or more servers 603 to one or more client devices 604, thereby reducing the requirements of one or more client devices 604 for graphics processing and rendering.
[0182] For example, with respect to an example of a game session, the client device 604 may display a frame of the game session on a display 624 based on receiving display data from one or more servers 603. The client device 604 may receive input from one of the one or more input devices and generate input data in response. The client device 604 may transmit the input data to one or more servers 603 via the communication interface 621 and via one or more networks 606 (e.g., the Internet), and the one or more servers 603 may receive the input data via the communication interface 618. The CPU may receive the input data, process the input data, and transmit the data to the GPU so that the GPU generates a rendering of the game session. For example, the input data may represent the movement of a character of a user in a game, firing a weapon, reloading, passing a ball, turning a vehicle, etc. The rendering component 612 may render the game session (e.g., representing the result of the input data), and the rendering capture component 614 may capture the rendering of the game session as display data (e.g., capturing image data of a rendered frame of the game session). The rendering of the game session may include ray or path tracing lighting and / or shadow effects calculated using one or more parallel processing units (such as GPUs), which may further employ one or more dedicated hardware accelerators or processing cores to perform ray or path tracing techniques of one or more servers 603. The encoder 616 may then encode the display data to produce encoded display data, and the encoded display data may be transmitted to the client device 604 via the network 606 via the communication interface 618. The client device 604 may receive the encoded display data via the communication interface 621, and the decoder 622 may decode the encoded display data to generate the display data. The client device 604 may then display the display data via the display 624.
[0183] It is noted that the techniques described herein may be implemented in executable instructions stored in a computer-readable medium for use by or in conjunction with a processor-based instruction execution machine, system, device, or apparatus. Those skilled in the art will appreciate that for some embodiments, different types of computer-readable media may be included for storing data. As used herein, "computer-readable media" includes one or more of any suitable media for storing executable instructions of a computer program, so that an instruction execution machine, system, device, or apparatus can read (or obtain) instructions from the computer-readable medium and execute instructions for implementing the described embodiments. Suitable storage formats include one or more of electronic formats, magnetic formats, optical formats, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer disks; random access memory (RAM); read-only memory (ROM); erasable programmable read-only memory (EPROM); flash memory devices; and optical storage devices, including portable compact disks (CDs), portable digital video disks (DVDs), and the like.
[0184] It should be understood that the arrangement of the parts shown in the drawings is for illustrative purposes and other arrangements are possible. For example, one or more of the elements described herein may be implemented as an electronic hardware assembly in whole or in part. Other elements may be implemented in software, hardware, or a combination of software and hardware. In addition, some or all of these other elements may be combined, some elements may be completely omitted, and additional components may be added while still implementing the functions described herein. Thus, the subject matter described herein may be embodied in many different variations, and all such variations are contemplated to be within the scope of the claims.
[0185] For ease of understanding of the subject matter described herein, many aspects are described with respect to action sequences. It will be appreciated by those skilled in the art that different actions may be performed by a dedicated circuit or circuit, by a program instruction executed by one or more processors, or by a combination of the two. The description of any action sequence herein is not intended to imply that the described particular order for executing the sequence must be followed. Unless otherwise indicated herein or context clearly contradicts, all methods described herein may be performed in any suitable order.
[0186] The use of the terms "a" and "the" and similar references in the context of describing the subject matter (particularly in the context of the following claims) should be interpreted to cover both the singular and the plural, unless otherwise specified herein or clearly contradicted by the context. The use of the term "at least one" followed by a list of one or more items (e.g., "at least one of A and B") should be interpreted to mean an item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise specified herein or clearly contradicted by the context. In addition, the foregoing description is for illustrative purposes only, not for limiting purposes, because the scope of protection sought is defined by the claims set forth below and any equivalents thereof. The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate the subject matter and does not limit the scope of the subject matter, unless otherwise required. The use of the term "based on" and other similar phrases indicating conditions that cause a result in the claims and written description is not intended to exclude any other conditions that cause the result. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention as claimed.
Claims
1. A computer-implemented method for adaptive sampling, comprising: determining a total number of samples for the image based on a target per-pixel sampling rate for the image; distributing the total number of samples across pixels included in the image according to the importance map to generate an initial sampling map for the image; as well as The initial sampling map is quantized using per-pixel random values to produce a sampling map for the image that includes a quantized sampling rate that includes an integer number of samples for each pixel in the image, wherein the quantized sampling rate is based on a comparison between a fractional portion of an initial sampling rate included in the initial sampling map for a pixel and the per-pixel random value associated with the pixel. 2 . The computer-implemented method of claim 1 , wherein the integer quantity is a value that is a power of two. 3 . The computer-implemented method of claim 1 , wherein the distributing and quantizing are performed in parallel for each pixel in the image. 4 . The computer-implemented method of claim 1 , wherein the image is included in a sequence of images, and the per-pixel random value varies for each image in the sequence of images.
5. The computer-implemented method of claim 1, wherein the per-pixel random value is read from a texture map or generated by a function.
6. The computer-implemented method of claim 1 , further comprising: The scene is rendered according to the sampling graph to produce a frame. 7 . The computer-implemented method of claim 6 , wherein the scene is rendered using ray tracing and the sampling map controls a number of rays cast for each pixel in the frame.
8. The computer-implemented method of claim 1, wherein the total number of samples is the product of the target per-pixel sampling rate and the size of the image.
9. The computer-implemented method of claim 1 , further comprising: summing each pixel value in the importance map to produce a significance sum; as well as The factor of the image is calculated as the total number of samples divided by the sum of the importances.
10. The computer-implemented method of claim 9, wherein distributing the total number of samples comprises: For each pixel in the image, the factor is multiplied by the per-pixel value in the importance map for the pixel to produce an initial sampling rate for the pixel.
11. The computer-implemented method of claim 1 , wherein quantizing the initial sampling map comprises: For each pixel in the image: dividing the initial sampling rate of the pixel into an integer part and the fractional part; as well as When the fractional portion is less than the per-pixel random value associated with the pixel, incrementing the integer portion to generate the quantized sampling rate for the pixel, or When the fractional portion is not less than the per-pixel random value associated with the pixel, the quantized sampling rate is set equal to the integer portion.
12. The computer-implemented method of claim 1, wherein the target per-pixel sampling rate is calculated based on a rendering frame rate and a target frame rate.
13. The computer-implemented method of claim 1 , wherein distributing the total number of samples comprises: A minimum number of samples are distributed to each pixel before distributing the remainder of the total number of samples across the pixels.
14. The computer-implemented method of claim 1, wherein the steps of determining, distributing, and quantizing are repeated for at least one additional image, additional importance map, and additional per-pixel random value.
15. The computer-implemented method of claim 1, wherein the steps of determining, distributing, and quantifying are performed within a cloud computing environment.
16. The computer-implemented method of claim 1, wherein the steps of determining, distributing, and quantizing are performed on a server or in a data center and the image is streamed to a user device.
17. The computer-implemented method of claim 1, wherein the images are used to train, test, or validate a neural network employed in a machine, robot, or autonomous vehicle.
18. The computer-implemented method of claim 1, wherein the steps of determining, distributing, and quantizing are performed on a virtual machine that includes a portion of a graphics processing unit.
19. A system comprising: The processor is configured as: determining a total number of samples for the image based on a target per-pixel sampling rate for the image; distributing the total number of samples across pixels included in the image according to the importance map to generate an initial sampling map for the image; as well as The initial sampling map is quantized using per-pixel random values to produce a sampling map for the image that includes a quantized sampling rate that includes an integer number of samples for each pixel in the image, wherein the quantized sampling rate is based on a comparison between a fractional portion of an initial sampling rate included in the initial sampling map for a pixel and the per-pixel random value associated with the pixel.
20. The system of claim 19, wherein the distributing and quantizing are performed in parallel for each pixel in the image.
21. A non-transitory computer readable medium storing computer instructions for adaptive sampling, the computer instructions, when executed by one or more processors, causing the one or more processors to perform the following steps: determining a total number of samples for the image based on a target per-pixel sampling rate for the image; distributing the total number of samples across pixels included in the image according to the importance map to generate an initial sampling map for the image; and The initial sampling map is quantized using per-pixel random values to produce a sampling map for the image comprising a quantized sampling rate, the quantized sampling rate comprising an integer number of samples for each pixel in the image, wherein The quantized sampling rate is based on a comparison between a fractional portion of an initial sampling rate included in the initial sampling map for a pixel and the per-pixel random value associated with the pixel.
Citation Information
Patent Citations
Computing resource request-driven adaptive cloud rendering method for three-dimensional scene
CN110717968A
Image processing apparatus and method
US20100277478A1
Adaptive multi-resolution for graphics
US20180284872A1
Mobile Cleaning Robot Artificial Intelligence for Situational Awareness
US20190213438A1