Method, system and storage medium for denoising a dynamic ray tracing scene using temporal accumulation
By introducing a fast history buffer and current frame buffer into the image generation system, combining exponential moving average and mixed weights, the problems of time lag and artifacts in dynamic scenes are solved, and efficient image denoising and smoothing processing are achieved.
Patent Information
- Application Number
- CN202110901632.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-09
- Filing Date
- 2021-08-06
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-08-06
AI Technical Summary
The prior art is difficult to effectively reduce time lag and artifacts when processing images of dynamic scenes, and conventional methods increase computational complexity and resource requirements.
Time smoothing and denoising of images are performed by using a fast history buffer and current frame buffer, combining exponential moving averages and mixed weights. Fast history buffers use higher mix weights, respond quickly and reduce time lag.
It realizes the effect of reducing time lag and artifacts in dynamic scenes, while reducing computational complexity and resource requirements, and improving the smoothness and reality of the image.
Smart Images

Figure CN114332250B_ABST
Abstract
Description
Background Art
[0001] As the quality of display devices - along with user expectations - continues to increase, there is a need to continually increase the quality of the content to be displayed. This may include tasks such as removing noise and reducing artifacts in rendered images, such as content frames that may correspond to video games or animations. Some conventional methods utilize processes such as temporal accumulation to attempt to denoise different effects (such as shadows, reflections, ambient occlusion, and direct illumination for ray tracing applications). However, in many cases, this may result in time lags in dynamic scenes, which may result in noticeable ghosting due to the inability of this accumulation to quickly and accurately account for changes in dynamic scenes. Existing methods for managing time lags either fail to adequately reduce undesirable effects, or have added undesirable additional computational requirements and complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Various embodiments according to the present disclosure will be described with reference to the accompanying drawings, in which:
[0003] Figure 1A , 1B , 1C and 1D illustrate images rendered for a dynamic scene according to at least one embodiment;
[0004] Figure 2 An example image generation system according to at least one embodiment is shown;
[0005] Figure 3A , 3B , 3C, 3D, and 3E illustrate stages in an example clamping process according to at least one embodiment;
[0006] Figure 4 An example history clamping process is shown in accordance with at least one embodiment;
[0007] Figure 5 An example image generation system including a clamped perceptual blur capability is shown in accordance with at least one embodiment;
[0008] Figure 6 A process for applying clamped perceptual blur to an image is shown in accordance with at least one embodiment;
[0009] Figure 7 An example data center system is shown in accordance with at least one embodiment;
[0010] Figure 8 A computer system according to at least one embodiment is shown;
[0011] Fig. 9 A computer system according to at least one embodiment is shown;
[0012] Fig.10illustrates at least a portion of a graphics processor according to one or more embodiments; and
[0013] Fig.11 At least a portion of a graphics processor is shown in accordance with one or more embodiments. DETAILED DESCRIPTION
[0014] Methods according to various embodiments can overcome defects in existing image generation methods. Specifically, various embodiments can provide improved denoising of image artifacts (such as artifacts that can be introduced by ray tracing or other image generation or rendering techniques). In a system for generating images or video frames of dynamic scenes, there may be different artifacts caused by changes in the scene. Temporal accumulation can be used to try to minimize the presence of at least some of these artifacts. The temporal accumulation method can retain information from previously generated frames in a sequence to try to provide at least a certain amount of temporal smoothing, where the color of the pixel in the current frame is mixed with the color from the previous frame to try to minimize ghosting and other such artifacts, and provide a smoother transition of the color in the scene. When mixing the colors from the currently rendered frame and the historical frame, it may be desirable to achieve an appropriate balance when determining the mixing weights. If the color of the current frame is weighted too heavily, the effectiveness of smoothing can be reduced, resulting in increased artifacts. On the contrary, if the weighting of the historical color is too heavy, there may be an undesirable time lag.
[0015] As an example, consider Figures 1A to 1D Such as Figure 1A In image 100 of FIG. 1 , there is a light source 102 on the wall, which may correspond to a window, an illuminated poster, a mirror, or other such light source. Visible on the nearby floor is an area 104 where the emitted or reflected light is brighter, causing the light source to color those pixels brighter. In some embodiments, this may correspond to a dynamic scene where the position of an object in the current view may change. For example, in Figure 1B In the image 120 of FIG. 1 , there may be a pan motion of the virtual camera or other such motion that causes the position of the light source 102 in the image to move relative to the previous frame. Many rendering engines will detect this motion or provide information about this motion, such as may be related to motion vectors of different objects in the scene. However, this motion information may not indicate other relevant changes - such as may correspond to, for example, but not limited to, shadows, lighting, reflections, or ambient occlusion - which may change due to this motion, such as may be due to ray tracing involving one or more moving objects in the scene. If temporal smoothing is applied, the influence of historical data may cause changes in pixels of this area or object (e.g., bright area 104) to not move with the corresponding object 102. Figure 1B As shown, in the initial frame where the light source 102 moves to the right, the bright area 104 may hardly move. Figure 1C As shown in the image 140 of FIG. 1 , there will be some time lag, where the bright area 104 will follow the movement of the light source 102 with some delay until the bright area reaches a position such as Figure 1D 160. Factors such as the blending weights utilized and the amount of historical data may affect the extent of this lag, which may be quite significant in some cases and may distract the viewer, or may at least reduce the perceived realism of the scene. It is desirable to retain the advantages of temporal smoothing and accumulation in at least some systems while reducing the effects of lag on dynamic scenes.
[0016] One of the most common temporal accumulation techniques is temporal anti-aliasing (TAA). One technique developed to improve TAA is to use history clamping as a way to handle dynamic events. Traditional temporal accumulation methods typically involve the use of two buffers. The first buffer is a history buffer, which contains a large number of frames, such as 30 frames for a given application, but can be in the range of about 10 to 100 frames or more for other applications. These frames can be accumulated over time using an exponential moving average. The second buffer is a current buffer containing data for the current frame (e.g., the most recent frame received from the rendering engine). In conventional methods, this history buffer (h) will be reprojected to the camera position of the current frame (c) and synthesized with the data of the current frame using (for example) an exponential moving average blending weight (w). In at least one implementation, the pixel value (p) of the output frame can be given by the following formula:
[0017] p=w*c+(1-w)*h
[0018] To handle time lags, historical values are typically clamped to a min / max window of a pixel neighborhood (e.g., a 3x3 pixel window) from the current frame before temporal anti-distortion. Other metrics may be used instead of min / max, which may involve, for example, calculating the mean and variance of the neighborhood and clamping the historical values to that distribution. This prior art approach is often insufficient for very noisy signals or dynamic scenes.
[0019] Another existing method that attempts to manage lag is A-SVGF, a spatiotemporal variance-guided filtering technique. This process is actually a filtering technique that can take a noisy frame as input and reconstruct it into a full image with reduced noise. A-SVGF can produce the desired results in many instances, but the time required to produce these results may be too long for a ray tracer operating in real time at modern frame rates (such as at least 60 frames per second (fps)). A-SVGF uses temporal gradients to minimize lag while denoising the path-traced image; however, since it must be deeply integrated into the renderer, it increases the renderer cost and complexity. Some games use the speed of the occluder to guide the temporal accumulation of shadows. Although this is a relatively simple solution, it brings limitations and corner cases that are difficult to solve. For example, if both the occluder and the receiver are moving, guiding the temporal accumulation becomes complicated. Moreover, when denoising ray tracing effects, the current frame signal is usually very noisy and is not suitable for calculating the neighborhood clamping window because the window is ultimately noisy.
[0020] Methods according to various embodiments utilize a responsive or "fast" history buffer in conjunction with a conventional history buffer and a current frame buffer. The fast history buffer can utilize much higher blending weights than for a conventional or "full" history buffer, such as an order of magnitude higher, so that fewer history frames contribute to the fast history. This fast history can be used to determine a clamping window for the current frame in order to clamp normal history values before or after reprojection. In at least one embodiment, the fast history clamping blending weights can be used as an intuitive knob to determine the appropriate balance between the amount of noise and the amount of lag in the scene. In at least one embodiment, a full history is maintained to provide more accurate temporal smoothing in the absence of conditions that cause the application of clamping.
[0021] Figure 2 Components of an example image generation system 200 that can be utilized in accordance with various embodiments are shown. In at least one embodiment, content such as video game content or animation can be generated using a renderer 202, a rendering engine, or other such content generation system or component. The renderer 202 can receive an input of one or more frames of a sequence, and can generate images or frames of a video using stored content 204 that is modified at least in part based on the input. In at least one embodiment, the renderer 202 can be part of a rendering pipeline that can provide functionality such as deferred shading, global illumination, illuminated translucency, post-processing, and graphics processing unit (GPU) particle simulation using vector fields.
[0022] In some embodiments, the amount of processing necessary to generate such complex, high-resolution images may make it difficult to render these video frames to meet current frame rates, such as at least sixty frames per second (fps). In at least one embodiment, the renderer 202 may be used to generate rendered images at a resolution lower than one or more final output resolutions in order to meet timing requirements and reduce processing resource requirements. The renderer may alternatively render the current image (or may otherwise obtain the current image) at the same resolution as the target output image, so that no upscaling or super-resolution procedures are needed or utilized. In at least one embodiment, if the current rendered image has a lower resolution, the low-resolution rendered image may be processed using an (optional) scaler 206 to generate an upscaled image that represents the content of the low-resolution rendered image at a resolution equal to (or at least more closely approximating) the target output resolution.
[0023] The currently rendered image (whether or not enlarged) may be provided as an input to an image reconstruction module 208 that may generate a high-resolution, anti-distorted output image using data of the current image and one or more previously generated images, which may be at least temporarily stored in a full history buffer 216 or other such location. In some embodiments, the previously generated image may be a single historical image in which pixel (e.g., color) values are accumulated over multiple previous frames using, for example, an exponential moving average. In at least one embodiment, the image reconstruction module 208 may include a blending component 210, such as one or more neural networks. In at least one embodiment, this may include at least a first optical flow network (OFN) for generating motion vectors or other information indicating movement between adjacent frames in a sequence. In at least one embodiment, this may include external regression, image pre-reconstruction, unsupervised optical flow networks. In at least one embodiment, this may also include at least a first image reconstruction network (RN) that utilizes these motion vectors to correlate positions in the current image and previous (historical) images and infer an output image from a blend of those images. In at least one embodiment, this blending of the current image with the historical image frames can help temporally converge to a good, clear, high-resolution output image, which can then be provided for presentation via display 212 or other such presentation mechanism. In at least one embodiment, a copy of the output image can also be provided to history manager 214, which can use an accumulation factor to accumulate pixel values with values from previous historical frames, so that the accumulated historical data can be represented by a single frame to save memory and reduce processing requirements. In at least one embodiment, the weighting factor can cause the pixel values from the older frames to contribute less to the accumulated pixel values. The historical image frame can then be stored in the full history buffer 216 or another such storage location for blending with subsequently generated images in the sequence.
[0024] As mentioned, real-time reconstruction of images utilizes information from one or more previous frames after a certain warping to align the image generated for the current frame. In at least one embodiment, such warping is utilized at least in part because image reconstruction is simplified when pixel information in these images is aligned. However, in at least one embodiment, appropriate image deformation utilizes not only information from previous frames, but also additional information about how objects move between these frames. In at least one embodiment, this may include computer vision or optical flow data, which may be represented by a set of motion vectors. In at least one embodiment, this may include motion vectors for each pixel position, or at least for pixel positions where there is movement. This motion information may help better warp image information and align corresponding pixels or objects. Motion vector information may be provided by a rendering engine for games or applications, but as mentioned, this may not take into account corresponding changes in aspects such as reflections or lighting that may be generated by a ray tracing process. This blending may then result in lag as previously discussed. Different existing methods may apply clamping to try to reduce lag, but with different drawbacks presented previously.
[0025] Thus, in at least one embodiment, the history manager 214 may also generate a response or "fast" history frame that may be stored in the fast history buffer 218. The fast frame may be generated using different cumulative weights, which in some cases may be approximately an order of magnitude greater than the cumulative weights used to generate the full history image. In one example, the cumulative weight of the fast history frame is approximately 0.5, while the cumulative weight of the full history frame is approximately 0.05. In at least one embodiment, this may result in the fast history frame including data accumulated over the most recent two to four frames, while the full history frame may include data accumulated over the most recent twenty to one hundred frames. These weights may be learned over time or set by a user, and may be configurable through one or more interfaces and other such options.
[0026] In at least one embodiment, when the image reconstruction module 208 is to generate the next output frame in the sequence, this fast history frame can be pulled from the fast history buffer 218. As mentioned, it may be desirable to blend the newly rendered current frame with the full history frame to provide at least some temporal smoothing of the image to reduce the presence of artifacts when displayed. However, instead of clamping based on the full history frame, a fast history frame can be used to make clamp determinations, which will contain historical data accumulated only over a small number of previous frames (e.g., the previous two to four frames in the sequence). The clamp module 220 can analyze the number of pixels in the area around the pixel location to be analyzed, such as the pixels in a 3X3 pixel neighborhood of the fast history image. A neighborhood larger than 3x3 can be used, but additional spatial bias can be introduced for at least some dynamic scenes. The blending module can then determine the distribution of expected pixel (e.g., color) values for the pixels. This expected distribution can then be compared with the value of the corresponding pixel in the full history image. If the full history pixel value is outside the distribution of expected values, the pixel value can be "clamped" to, for example, the value that is closest to the historical pixel value within the distribution of expected values. Instead of clamping to the current value, which may cause ghosting, noise, or other artifacts, this method can clamp to an intermediate value determined using the fast history frame. The blending module can then take the clamped or otherwise value from the full history frame and blend accordingly with the pixels of the current frame as discussed herein. This new image can then be processed by the history manager 214 to generate an updated history image to be stored in the history buffers 216, 218 for use in reconstructing subsequent images.
[0027] about FIG. 3A to FIG. 3E This clamping process can be better understood. As discussed, a fast history frame can be used to perform clamping analysis. This history frame is generated by accumulating the history information of individual pixels using an identified accumulation factor or blending weight. As mentioned, although the accumulation factor of a complete history frame can be on the order of about 0.05, the accumulation factor of a fast frame can be much larger, such as on the order of 0.5, so that the contribution from older frames is minimized more quickly. Minimum additional effort is required to accumulate and reproject fast history in a system that has accumulated complete or "long" history information. In this example, blending weights are used together with an exponential moving average of past frame data to avoid storing data for each of those past frames in memory. The accumulated data in the history is multiplied by (1-weight) and then combined with the data in the current frame, which can be multiplied by the blending weight. In this way, a single history frame can be stored in each buffer, wherein the contribution of older frames is reduced according to a recursive accumulation method. In at least one embodiment, this cumulative weight may be automatically adjusted based on any of a number of factors, such as the current frame rate being provided, the total variance, or the amount of noise in the generated frame.
[0028] When performing clamp analysis, data for points in a surrounding neighborhood (e.g., a 3x3 neighborhood) of each pixel position in the fast history frame can be determined. These pixel values can each be viewed as a color point in a three-dimensional color space, such as Figure 3A 300. Although a red-green-blue (RGB) color space may be utilized in various embodiments, other color spaces having other numbers of dimensions (e.g., YIQ, CMYK (cyan, magenta, yellow, and black), YCoCg, or HSL (hue, saturation, brightness value)) may be utilized in other embodiments. Figure 3A As shown, these points from the neighborhood are located in a region of color space. When determining the pixel values of the historical frame that can be reasonably expected based on these points, different methods can be used to determine these expected values. The expected values can be located at or within the volume defined by the points, or a reasonable amount of distance outside of this volume, as is configurable and can depend at least in part on the method taken. Figure 3B In the example method of , the curve graph 320 shows the expected area 322 around these points. Any one of a plurality of projection or prediction algorithms or networks can be used to determine or infer the size and shape of this expected area. In one embodiment, a convex hull-based method can be used. Figure 3C Another example method is shown in which a bounding box 342 may be determined for those points in the color space as shown in the graph 340. The bounding box may be determined using many different bounding algorithms that may include different amounts of buffering around the points in each direction. In at least one embodiment, the expected area may be determined using a mean and variance distribution. Various other areas, boxes, ranges, or determinations may also be used within the scope of different embodiments.
[0029] Once this desired range or region is determined from the fast history frame, the corresponding pixels from the full history frame can be identified. Figure 3D The graph 360 of FIG. 360 shows historical pixel values 362 in color space relative to an expected region. The historical pixel can be compared to the expected region to determine whether the pixel falls within the expected region or falls outside the expected region. If the pixel value is within the expected region, the corresponding full historical pixel value can be used and no clamping is applied. However, it may be the case that the historical point 362 may be outside the region, such as Figure 3D In this case, clamping can be applied to the historical values. Figure 3EThe graph 380 shows that pixel values can be "clamped" or adjusted so that they fall within an expected range. In this example, the clamp value 382 to be used for the full history frame is the "clamped" value in color space that is closest to the expected range of fast history pixel values. In one embodiment, this can involve clamping by applying a min / max analysis along each dimension of the color space to determine a new clamp value.
[0030] Figure 4 An example process 400 for performing clamping of historical data that can be performed according to various embodiments is shown. It should be understood that for this and other processes presented herein, additional, fewer or alternative steps that may be performed in a similar or alternative order or at least partially in parallel may exist within the scope of different embodiments, unless otherwise specifically stated. In this example, a current frame is received 402 or otherwise obtained from a rendering engine. This may be a current frame or image in a series of frames or images, such as may be generated for animation, games, virtual reality (VR), augmented reality (AR), video, or other such content. The current frame may be mixed 404 with a corresponding fast history frame using an appropriate blending factor. Pixel values from the corresponding pixel neighborhood may be determined 406 for individual pixels of this fast history frame. The range of expected pixel values may then be determined 408 based on those neighborhood values. The corresponding pixel value from the complete history frame may be determined and compared 410 with this expected range. If it is determined 412 that the fast history value is within the expected range, the actual history value from the complete history frame may be used for this pixel position. If the fast history value is outside the range, a determination may be made to clamp 416 the full history value to the closest value within the expected range, such as by using a min / max or projection-based approach as discussed herein. Once such a determination is made for all relevant pixels, the value of the current frame may be blended 418 with the clamped value or actual value of the corresponding pixel of the history frame. Once completed, the reconstructed image may be provided 420 for presentation, such as through a display as part of a gaming experience. Further, updated fast and full history frames may be generated 422 using temporal accumulation with this newly reconstructed image, and corresponding buffers may be used to store these history frames for use in reconstructing the next image or frame in the sequence.
[0031] In at least one embodiment, an alternative process may be performed in which blending may occur earlier in the process, such as at step 404 rather than step 418. In such an embodiment, the values may be clamped for a complete history of frames that already contain the current frame values. While both approaches may produce acceptable results, in some cases one approach may be easier to implement.
[0032] To further improve the appearance of the generated images, a certain amount of blur (e.g., Gaussian blur) or other spatial filtering can be applied to dynamic scenes. This blur can help smooth the images in the sequence to provide more natural motion, and can also help reduce the presence of spatial sampling bias and artifacts (e.g., noise or flicker). It can be difficult to determine the appropriate amount of blur to apply to the image, because too much blur will reduce the clarity of the image, and too little blur may not be enough to remove these and other such artifacts. Further, there may be portions of a scene with significant motion, while other portions of the image are dynamic, such that it may be desirable to apply blur to the portions with motion and not apply blur to the static portions. However, as discussed herein, it can be difficult to identify motion associated with shadows, reflections, and other aspects that may come from processes such as ray tracing, making it difficult in existing methods to determine how much blur to apply to areas of the image associated with such aspects.
[0033] Thus, methods according to various embodiments may utilize information such as the clamp determination presented herein to determine the application of blurring, spatial filtering, or other such image processing. In the clamp process presented with respect to at least one embodiment, clamping is applied when the amount of motion or change in the image causes the historical pixel values from the fast image frame to fall outside the expected range or region. Using this process, a determination may be made for each pixel whether the pixel corresponds to a static portion of the image, a portion with a motion or change amount within the expected range, or a motion or change amount outside the expected range. In some embodiments, the amount of blurring may be applied based on the pixel difference, regardless of whether the distance will result in clamping. The amount of blurring or spatial filtering may be applied, which may be different for any or all of these situations. In another example method, a determination may be made for each pixel regarding the difference between the pixel values of the fast history frame and the full history frame. The difference in pixel values or the distance in color space may be used to determine the amount of blurring or the weighting of the spatial filter to be applied. This difference may be determined before or after the temporal accumulation clamping, and may therefore be based on the original value or the clamp value. For pixels with large differences or where clamping is applied, a larger spatial filter may be applied to those pixels. For smaller differences, a smaller spatial filter may be applied. If the fast and full history pixel values are the same, within an allowable deviation, no (or minimal) spatial filtering may be applied in some embodiments. In some embodiments, a minimal amount of spatial filtering may be applied to the entire image, where this history-based spatial filtering acts as an additional filter for pixels with larger degrees of motion or change.
[0034] In at least one embodiment, the difference between the colors of the full history and the colors of the fast history, or between the colors of the full history and the clamped full history, can be calculated. It should be understood that in this case, the "full" history refers to the multiple history buffers accumulated in this image generation process at any given time, and data from all previously generated images or frames in the sequence is not required. This difference can be multiplied by a constant or scalar to determine the weight of the spatial filter to be applied. In at least one embodiment, this weighting can be used to determine the blur radius to be applied, such as a radius of 0, 1, 2, or 3 pixels in any or all directions. This method can be used to apply only the amount of blur required for specific areas of the image, which can minimize the total amount of blur applied to the image and therefore produce a clearer image.
[0035] In at least one embodiment, time accumulation can be performed to generate a fast history frame and a full history frame using corresponding blending weights. The determined neighborhood of the fast history frame can then be used to perform historical clamping for each pixel position. In this example, a historical confidence [0.0, 1.0] can be calculated, which is a function of the difference between the full history value and the clamped full history value at a given pixel position. In at least one embodiment, this confidence value can be used as an indication of which pixels are affected by the historical clamping and the degree to which these pixels are affected. In this instance, a default amount of spatial filtering (e.g., cross-bilateral spatial filtering) can be applied to the pixel positions in the image. The depth / normal bilateral weight can be set to a minimum value, where the effective radius is calculated at least in part based on some form of noise estimation (within a predefined maximum radius), which can involve temporal variation, spatial variation, or total variation and other such options. In at least one embodiment, an additional spatial filter can be added, which is a function of the historical confidence calculated based on the historical clamping. In at least one embodiment, this may be a linear interpolation ("lerp") given by the interpolation lerp(MaxRadius, EffectiveRadius, HistoryConfidence). In at least one embodiment, HistoryConfidence may be given by HistoryConfidence = saturate(abs(FullAccumulatedHistory - ClampedFullAccumulatedHistory) * ScalingFactor, where ScalingFactor may be any arbitrary scaling factor that has been determined to provide acceptable results for a given signal or implementation. In this example, abs() is an absolute value function, and saturate() clamps this value to the range [0.0, 1.0]. If there is no clamping for a pixel position, then no additional spatial filtering may be applied. In at least some embodiments, the amount of additional spatial filtering applied may be a factor of the difference between the pixel value of a given pixel in the fast history frame and the full history frame.
[0036] Figure 5 An example system 500 is shown that can be utilized in accordance with various embodiments. Figure 2, but it should be understood that this is for simplicity of explanation and should not be construed as a limitation on the scope or variability of the various embodiments. In such a system, a fast history frame and a full history frame can be generated by the history manager component 214, as previously discussed. The fast history frame can be compared with the current frame to determine whether to apply clamping to a particular pixel position or region of the full history frame. As mentioned, information from this clamp determination process can be used to determine the amount of spatial filtering to be applied during image reconstruction. In this example, the clamp information can be passed from the clamp module 220 to the spatial filter module 502. In other embodiments, the spatial filter module 502 can act directly on the fast history frame and the full history frame. The spatial filter module 502 can determine information such as whether to apply clamping and the difference between the pixel values of corresponding pixel positions in the fast history image and the full history image. The spatial filter module 502 may then determine the size of the spatial filter or amount of blur to be applied to each pixel of the image generated by the reconstruction module 208, which may include a single per-pixel determination, or may include a default filter amount plus any additional filtering determined from the clamp or historical confidence data. The reconstructed image with this additional filtering applied may then be provided for display and provided to the history manager 214 for accumulation into updated full and fast history frames for use in reconstructing the next image or frame in the sequence.
[0037] Figure 6 An example process 600 for determining spatial filtering to be applied to an image during reconstruction that can be utilized according to various embodiments is shown. In this example, a clamp-based approach will be described, but as discussed herein, other information may also be used to determine historical confidence values. A fast history frame and a full history frame are obtained 602 from a temporal accumulation process. A determination 604 whether to apply clamping to individual pixels of a full history frame may be made based at least in part on the values of pixels from a corresponding pixel neighborhood of the fast history frame (e.g., whether the pixel value of the current frame falls within an expected range based on a given pixel neighborhood). A historical confidence value for each pixel position may then be calculated 606 based on the full history value and the clamped history value. If a clamp value is not used, there may be a high confidence in the full history pixel value. For a pixel of an image, a determination 608 to apply a default spatial filter during reconstruction may be made. Further, any additional spatial filtering to be applied may be determined 610 based at least in part on the historical confidence value, where a lower confidence corresponding to a larger difference in pixel value may result in a larger amount of spatial filtering, such as using a larger offset radius. Then, the default and additional spatial biases may be applied 612 to corresponding pixels during image reconstruction. In this way, spatial filtering may be minimized for more static portions of the image.
[0038] As mentioned, this approach can be advantageous when dealing with denoising of dynamic scenes. History clamping as discussed herein can help detect stale history, and in the event that stale history is detected, a temporary collision blending factor can be determined during the temporal accumulation, so that a heavier weight is assigned to the most recent data. While this can adequately address time lags, limiting the number of frames in the temporal accumulation can introduce an unacceptable amount of noise. Because the system cannot rely on temporal data to remove this noise, a spatial bias can be added in the locations where stale history is detected in order to provide enhanced denoising.
[0039] In other embodiments, other methods for determining and applying spatial biases may also be used, such as for determining temporal or spatial gradients, which may be independent of historical clamping. Such gradients may be determined for the A-SGVF process discussed previously. In another example, occlusion speed may be used as a factor for determining whether to apply additional or increased amounts of spatial filtering during image reconstruction, because ambient occlusion may provide another confidence measure for anything that uses historical pixel data. Ambient occlusion is a feature that may be used in processes such as ray tracing in a high-definition rendering pipeline (HDRP). Ambient occlusion may be used to calculate the exposure of each point in a scene to ambient lighting. Other types of occlusion determinations may also be used for other light sources, etc. The speed of determining occlusion or shading may indicate whether to trust historical pixel data and how much to trust. Light from a given light source may be traced to determine whether the corresponding geometry is moving, changing, or static, and the amount of movement or change. In some cases, only a binary decision based on occlusion speed regarding whether a pixel is static or dynamic may be provided. Where speed data is provided or determined, the blur radius or weight of the spatial filter may increase as speed increases. Further, in some embodiments where HistoryConfidence may be slightly unstable in time, this value may be reprojected and slightly increased over a few frames in order to make the spatial filter bias smoother in time. This approach may help denoise the otherwise improved image quality in situations where stale history data is detected.
[0040] Data Center
[0041] Figure 7 An example data center 700 is shown in which at least one embodiment may be used. In at least one embodiment, data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.
[0042] In at least one embodiment, Figure 7As shown, the data center infrastructure layer 710 may include a resource coordinator 712, group computing resources 714, and node computing resources ("node CRs") 716 (1)-716 (N), where "N" represents any positive integer. In at least one embodiment, the node CRs 716 (1)-716 (N) may include, but are not limited to, any number of central processing units ("CPUs") or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memories), storage devices (e.g., solid-state drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VMs"), power modules and cooling modules, etc. In at least one embodiment, one or more of the node CRs 716 (1)-716 (N) may be a server having one or more of the above computing resources.
[0043] In at least one embodiment, the grouped computing resources 714 may include separate groups (not shown) of node CRs housed in one or more racks, or many racks (also not shown) housed in data centers at various geographic locations. The separate groups of node CRs within the separate grouped computing resources 714 may include computing, networks, memory, or storage resources that can be configured or allocated to support groupings of one or more workloads. In at least one embodiment, several node CRs including a CPU or processor may be grouped in one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0044] In at least one embodiment, resource coordinator 712 may configure or otherwise control one or more nodes CR 716(1)-716(N) and / or grouped computing resources 714. In at least one embodiment, resource coordinator 712 may include a software design infrastructure ("SDI") management entity for data center 700. In at least one embodiment, resource coordinator 1012 may include hardware, software, or some combination thereof.
[0045] In at least one embodiment, Figure 7As shown, the framework layer 720 includes a job scheduler 722, a configuration manager 724, a resource manager 726, and a distributed file system 728. In at least one embodiment, the framework layer 720 may include a framework that supports software 732 of the software layer 730 and / or one or more applications 742 of the application layer 740. In at least one embodiment, the software 732 or the application 742 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 may be, but is not limited to, a free and open source software web application framework, such as Apache Spark, which may utilize the distributed file system 728 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 732 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 724 may be able to configure different layers, such as the software layer 730 and the framework layer 720 including Spark and a distributed file system 728 for supporting large-scale data processing. In at least one embodiment, the resource manager 726 can manage cluster or group computing resources mapped to or allocated to support the distributed file system 728 and the job scheduler 722. In at least one embodiment, the cluster or group computing resources may include group computing resources 714 on the data center infrastructure layer 710. In at least one embodiment, the resource manager 726 can coordinate with the resource coordinator 712 to manage these mapped or allocated computing resources.
[0046] In at least one embodiment, software 732 included in software layer 730 may include software used by at least a portion of node CRs 716(1)-716(N), grouped computing resources 714, and / or distributed file system 728 of framework layer 720. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0047] In at least one embodiment, one or more applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the node CRs 716(1)-716(N), the grouped computing resources 714, and / or the distributed file system 728 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0048] In at least one embodiment, any of the configuration manager 724, resource manager 726, and resource coordinator 712 can implement any number and type of self-modification actions based on any number and type of data acquired in any technically feasible manner. In at least one embodiment, the self-modification actions can relieve a data center operator of the data center 700 from making potentially bad configuration decisions and can avoid underutilized and / or poorly performing portions of the data center.
[0049] In at least one embodiment, the data center 700 may include tools, services, software, or other resources to train one or more machine learning models or use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 700. In at least one embodiment, by using weight parameters calculated by one or more training techniques described herein, information may be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to the data center 700.
[0050] In at least one embodiment, the data center can use a CPU, an application-specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to use the above resources to perform training and / or reasoning. In addition, one or more of the above software and / or hardware resources can be configured as a service to allow users to train or perform information reasoning, such as image recognition, speech recognition, or other artificial intelligence services.
[0051] Such components may be used to improve image quality using fast history-based clamping and disparity determination during image reconstruction.
[0052] Computer Systems
[0053] Figure 8800 is a block diagram illustrating an exemplary computer system according to at least one embodiment, which may be a system of interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor, which may include an execution unit to execute instructions. In at least one embodiment, according to the present disclosure, such as the embodiments described herein, the computer system 800 may include, but is not limited to, components, such as a processor 802, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, the computer system 800 may include a processor, such as a processor available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM microprocessor, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.) may also be used. In at least one embodiment, computer system 800 may execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.
[0054] Embodiments may be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (Internet Protocol) devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area network ("WAN") switch, or any other system that can execute one or more instructions according to at least one embodiment.
[0055] In at least one embodiment, the computer system 800 may include, but is not limited to, a processor 802, which may include, but is not limited to, one or more execution units 808 to perform machine learning model training and / or reasoning according to the techniques described herein. In at least one embodiment, the computer system 800 is a single-processor desktop or server system, but in another embodiment, the computer system 800 may be a multi-processor system. In at least one embodiment, the processor 802 may include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor that implements an instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 802 may be coupled to a processor bus 810, which may transmit data signals between the processor 802 and other components in the computer system 800.
[0056] In at least one embodiment, processor 802 may include, but is not limited to, a level 1 ("L1") internal cache memory ("cache") 804. In at least one embodiment, processor 802 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 802. Other embodiments may also include a combination of internal and external caches, depending on the particular implementation and needs. In at least one embodiment, register file 806 may store different types of data in various registers, including, but not limited to, integer registers, floating point registers, status registers, and instruction pointer registers.
[0057] In at least one embodiment, an execution unit 808, including but not limited to logic to perform integer and floating point operations, is also located in the processor 802. In at least one embodiment, the processor 802 may also include a microcode ("ucode") read-only memory ("ROM") for storing microcode for certain macroinstructions. In at least one embodiment, the execution unit 808 may include logic for processing a packed instruction set 809. In at least one embodiment, by including the packed instruction set 809 in the instruction set of a general-purpose processor, and the associated circuitry to execute the instructions, operations used by many multimedia applications may be performed using packed data in the general-purpose processor 802. In one or more embodiments, many multimedia applications may be executed faster and more efficiently by using the full width of the processor's data bus to perform operations on packed data, which may not require the transfer of smaller units of data on the processor's data bus to perform one or more operations one data element at a time.
[0058] In at least one embodiment, execution unit 808 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 800 may include, but is not limited to, memory 820. In at least one embodiment, memory 820 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other storage device. In at least one embodiment, memory 820 may store instructions 819 and / or data 821 represented by data signals that may be executed by processor 802.
[0059] In at least one embodiment, the system logic chip can be coupled to the processor bus 810 and the memory 820. In at least one embodiment, the system logic chip can include, but is not limited to, a memory controller hub ("MCH") 816, and the processor 802 can communicate with the MCH 816 via the processor bus 810. In at least one embodiment, the MCH 816 can provide a high-bandwidth memory path 818 to the memory 820 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, the MCH 816 can initiate data signals between the processor 802, the memory 820, and other components in the computer system 800, and bridge data signals between the processor bus 810, the memory 820, and the system I / O interface 822. In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 816 can be coupled to the memory 820 via a high-bandwidth memory path 818, and the graphics / video card 812 can be coupled to the MCH 816 via an Accelerated Graphics Port ("AGP") interconnect 814.
[0060] In at least one embodiment, the computer system 800 may use a system I / O interface 822, which is a proprietary hub interface bus, to couple the MCH 816 to an I / O controller hub ("ICH") 830. In at least one embodiment, the ICH 830 may provide direct connection to certain I / O devices through a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus used to connect peripheral devices to the memory 820, the chipset, and the processor 802. Examples may include, but are not limited to, an audio controller 829, a firmware hub ("Flash BIOS") 828, a wireless transceiver 826, a data store 824, a traditional I / O controller 823 including a user input and keyboard interface, a serial expansion port 827 (e.g., a universal serial bus (USB) port), and a network controller 834. The data store 824 may include a hard drive, a floppy drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0061] In at least one embodiment, Figure 8 The system is shown as comprising interconnected hardware devices or "chips", while in other embodiments, Figure 8 A system on a chip (SoC) may be shown. In at least one embodiment, the devices may be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 800 are interconnected using a compute express link (CXL) interconnect.
[0062] Such components may be used to improve image quality using fast history-based clamping and disparity determination during image reconstruction.
[0063] Fig. 9 is a block diagram illustrating an electronic device 900 for utilizing a processor 910 according to at least one embodiment. In at least one embodiment, the electronic device 900 may be, for example but not limited to, a notebook computer, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet computer, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0064] In at least one embodiment, system 900 may include, but is not limited to, a processor 910 communicatively coupled to any suitable number or variety of components, peripherals, modules, or devices. In at least one embodiment, processor 910 is coupled using a bus or interface, such as an I 2C bus, system management bus ("SMBus"), low pin count (LPC) bus, serial peripheral interface ("SPI"), high-definition audio ("HDA") bus, serial advanced technology attachment ("SATA") bus, universal serial bus ("USB") (version 1, 2, 3, etc.), or universal asynchronous receiver / transmitter ("UART") bus. In at least one embodiment, Fig. 9 A system is shown that includes interconnected hardware devices or "chips", while in other embodiments, Fig. 9 An exemplary system on chip (SoC) may be shown. In at least one embodiment, Fig. 9 The devices shown in can be interconnected with a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, Fig. 9 One or more components of the system are interconnected using Compute Express Link (CXL) interconnect lines.
[0065] In at least one embodiment, Fig. 9 The display 924, touch screen 925, touch pad 930, near field communication unit ("NFC") 945, sensor hub 940, thermal sensor 946, fast chipset ("EC") 935, trusted platform module ("TPM") 938, BIOS / firmware / flash memory ("BIOS, FW Flash") 922, DSP 960, drive 920 (e.g., solid state disk ("SSD") or hard disk drive ("HDD")), wireless local area network unit ("WLAN") 950, Bluetooth unit 952, wireless wide area network unit ("WWAN") 956, global positioning system (GPS) unit 955, camera ("USB 3.0 camera") 954 (e.g., USB 3.0 camera) and / or low power double data rate ("LPDDR") memory unit ("LPDDR3") 915 implemented in, for example, the LPDDR3 standard may be included. Each of these components may be implemented in any suitable manner.
[0066] In at least one embodiment, other components may be communicatively coupled to the processor 910 via the components described above herein. In at least one embodiment, an accelerometer 941, an ambient light sensor (“ALS”) 942, a compass 943, and a gyroscope 944 may be communicatively coupled to the sensor hub 940. In at least one embodiment, a thermal sensor 939, a fan 937, a keyboard 936, and a touchpad 930 may be communicatively coupled to the EC 935. In at least one embodiment, a speaker 963, an earphone 964, and a microphone (“mic”) 965 may be communicatively coupled to an audio unit (“audio codec and class D amplifier”) 962, which in turn may be communicatively coupled to the DSP 960. In at least one embodiment, the audio unit 962 may include, for example, but not limited to, an audio encoder / decoder (“codec”) and a class D amplifier. In at least one embodiment, a SIM card (“SIM”) 957 may be communicatively coupled to the WWAN unit 956. In at least one embodiment, components such as the WLAN unit 950 and the Bluetooth unit 952 and the WWAN unit 956 may be implemented as a next generation form factor (NGFF).
[0067] Such components may be used to improve image quality using fast history-based clamping and disparity determination during image reconstruction.
[0068] Fig.10 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, system 1000 includes one or more processors 1002 and one or more graphics processors 1008, and may be a single processor desktop system, a multi-processor workstation system, or a server system with a large number of processors 1002 or processor cores 1007. In at least one embodiment, system 1000 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in a mobile, handheld, or embedded device.
[0069] In at least one embodiment, the system 1000 may include or be incorporated into a server-based gaming platform, including a gaming console, a mobile gaming console, a handheld gaming console, or an online gaming console for gaming and media consoles. In at least one embodiment, the system 1000 is a mobile phone, a smart phone, a tablet computing device, or a mobile Internet device. In at least one embodiment, the processing system 1000 may also include a wearable device coupled to or integrated in a wearable device, such as a smart watch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device. In at least one embodiment, the processing system 1000 is a television or set-top box device having one or more processors 1002 and a graphical interface generated by one or more graphics processors 1008.
[0070] In at least one embodiment, one or more processors 1002 each include one or more processor cores 1007 to process instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 1007 is configured to process a specific instruction set 1009. In at least one embodiment, the instruction set 1009 can facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or calculate by very long instruction words (VLIW). In at least one embodiment, the processor cores 1007 can each process different instruction sets 1009, and the instruction sequence can include instructions that help emulate other instruction sets. In at least one embodiment, the processor core 1007 can also include other processing devices, such as a digital signal processor (DSP).
[0071] In at least one embodiment, the processor 1002 includes a cache memory 1004. In at least one embodiment, the processor 1002 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared between various components of the processor 1002. In at least one embodiment, the processor 1002 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), which can be shared between the processor cores 1007 using known cache coherence techniques. In at least one embodiment, the processor 1002 additionally includes a register file 1006, and the processor may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, the register file 1006 may include general registers or other registers.
[0072] In at least one embodiment, one or more processors 1002 are coupled to one or more interface buses 1010 to transmit communication signals, such as address, data, or control signals, between the processor 1002 and other components in the system 1000. In at least one embodiment, the interface bus 1010 can be a processor bus in one embodiment, such as a version of a direct media interface (DMI) bus. In at least one embodiment, the interface bus 1010 is not limited to a DMI bus, and can include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor 1002 includes an integrated memory controller 1016 and a platform controller hub 1030. In at least one embodiment, the memory controller 1016 facilitates communication between memory devices and other components of the processing system 1000, while the platform controller hub (PCH) 1030 provides connections to input / output (I / O) devices through a local I / O bus.
[0073] In at least one embodiment, the memory device 1020 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or have appropriate performance to be used as a processor memory. In at least one embodiment, the memory device 1020 may be used as a system memory of the processing system 1000 to store data 1022 and instructions 1021 for use when one or more processors 1002 execute applications or processes. In at least one embodiment, the memory controller 1016 is also coupled to an optional external graphics processor 1012, which may communicate with one or more graphics processors 1008 in the processor 1002 to perform graphics and media operations. In at least one embodiment, the display device 1011 may be connected to the processor 1002. In at least one embodiment, the display device 1011 may include one or more of the internal display devices, such as in a mobile electronic device or laptop device or an external display device connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 1011 may include a head mounted display (HMD), such as a stereoscopic display device used in virtual reality (VR) applications or augmented reality (AR) applications.
[0074] In at least one embodiment, the platform controller hub 1030 enables peripheral devices to be connected to the storage device 1020 and the processor 1002 via a high-speed I / O bus. In at least one embodiment, the I / O peripherals include, but are not limited to, an audio controller 1046, a network controller 1034, a firmware interface 1028, a wireless transceiver 1026, a touch sensor 1025, a data storage device 1024 (e.g., a hard drive, flash memory, etc.). In at least one embodiment, the data storage device 1024 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnect bus (e.g., PCI, PCIe). In at least one embodiment, the touch sensor 1025 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 1026 can be a Wi-Fi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or long-term evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1028 enables communication with the system firmware and can be, for example, a unified extensible firmware interface (UEFI). In at least one embodiment, the network controller 1034 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 1010. In at least one embodiment, the audio controller 1046 is a multi-channel high-definition audio controller. In at least one embodiment, the processing system 1000 includes an optional legacy I / O controller 1040 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system 1000.
[0075] In at least one embodiment, the platform controller hub 1030 may also be connected to one or more Universal Serial Bus (USB) controllers 1042 that connect input devices such as a keyboard and mouse 1043 combination, camera 1044, or other USB input devices.
[0076] In at least one embodiment, instances of memory controller 1016 and platform controller hub 1030 may be integrated into a discrete external graphics processor, such as external graphics processor 1012. In at least one embodiment, platform controller hub 1030 and / or memory controller 1016 may be external to one or more processors 1002. For example, in at least one embodiment, system 1000 may include external memory controller 1016 and platform controller hub 1030, which may be configured as a memory controller hub and a peripheral controller hub in a system chipset that communicates with processor 1002.
[0077] Such components may be used to improve image quality using fast history-based clamping and disparity determination during image reconstruction.
[0078] Fig.11 is a block diagram of a processor 1100 having one or more processor cores 1102A-1102N, an integrated memory controller 1114, and an integrated graphics processor 1108 in accordance with at least one embodiment. In at least one embodiment, the processor 1100 may include additional cores, up to and including the additional core 1102N represented by the dashed box. In at least one embodiment, each processor core 1102A-1102N includes one or more internal cache units 1104A-1104N. In at least one embodiment, each processor core may also have access to one or more shared cache units 1106.
[0079] In at least one embodiment, the internal cache units 1104A-1104N and the shared cache unit 1106 represent a cache memory hierarchy within the processor 1100. In at least one embodiment, the cache memory units 1104A-1104N may include at least one level of instruction and data cache within each processor core and one or more levels of cache in a shared mid-level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is categorized as LLC. In at least one embodiment, cache coherency logic maintains coherency between the various cache units 1106 and 1104A-1104N.
[0080] In at least one embodiment, the processor 1100 may also include a set of one or more bus controller units 1116 and a system agent core 1110. In at least one embodiment, the one or more bus controller units 1116 manage a set of peripheral buses, such as one or more PCI or PCIe buses. In at least one embodiment, the system agent core 1110 provides management functions for various processor components. In at least one embodiment, the system agent core 1110 includes one or more integrated memory controllers 1114 to manage access to various external memory devices (not shown).
[0081] In at least one embodiment, one or more processor cores 1102A-1102N include support for multiple threads simultaneously. In at least one embodiment, system agent core 1110 includes components for coordinating and operating cores 1102A-1102N during multithreaded processing. In at least one embodiment, system agent core 1110 may additionally include a power control unit (PCU) that includes logic and components for regulating one or more power states of processor cores 1102A-1102N and graphics processor 1108.
[0082] In at least one embodiment, the processor 1100 also includes a graphics processor 1108 for performing graphics processing operations. In at least one embodiment, the graphics processor 1108 is coupled to a shared cache unit 1106 and a system agent core 1110 including one or more integrated memory controllers 1114. In at least one embodiment, the system agent core 1110 also includes a display controller 1111 for driving the graphics processor output to one or more coupled displays. In at least one embodiment, the display controller 1111 may also be a separate module coupled to the graphics processor 1108 via at least one interconnect, or may be integrated within the graphics processor 1108.
[0083] In at least one embodiment, a ring-based interconnect unit 1112 is used to couple the internal components of the processor 1100. In at least one embodiment, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other technologies. In at least one embodiment, the graphics processor 1108 is coupled to the ring interconnect 1112 via an I / O link 1113.
[0084] In at least one embodiment, I / O link 1113 represents at least one of a variety of I / O interconnects, including packaged I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 1118 (e.g., eDRAM modules). In at least one embodiment, each of processor cores 1102A-1102N and graphics processor 1108 uses embedded memory modules 1118 as a shared last level cache.
[0085] In at least one embodiment, the processor cores 1102A-1102N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 1102A-1102N are heterogeneous in terms of instruction set architecture (ISA), wherein one or more processor cores 1102A-1102N execute a common instruction set, while one or more other processor cores 1102A-1102N execute a subset or a different instruction set of the common instruction set. In at least one embodiment, the processor cores 1102A-1102N are heterogeneous in terms of microarchitecture, wherein one or more cores with relatively high power consumption are coupled with one or more power cores with lower power consumption. In at least one embodiment, the processor 1100 can be implemented on one or more chips or implemented as a SoC integrated circuit.
[0086] Such components may be used to improve image quality using fast history-based clamping and disparity determination during image reconstruction.
[0087] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described in detail above. However, it should be understood that there is no intention to limit the disclosure to one or more specific forms disclosed, but on the contrary, it is intended to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of the present disclosure as defined by the appended claims.
[0088] Unless otherwise noted or clearly contradictory to the context, in the context of describing the disclosed embodiments (particularly in the context of the appended claims), the use of the terms "one" and "an" and "the" and similar references should be interpreted as covering the singular and plural, rather than as definitions of terms. Unless otherwise noted, the terms "include", "have", "include" and "contain" should be interpreted as open terms (meaning "including but not limited to") unless otherwise noted. The term "connected" (which refers to a physical connection when unmodified) should be interpreted as partially or completely included, attached to or connected together, even if there are some interventions. Unless otherwise noted herein, references to numerical ranges herein are intended only to be used as a shorthand method of referring to each individual value falling within the range, respectively, and each individual value is incorporated into the specification as if it were individually described herein. Unless otherwise noted or contradictory to the context, the use of the term "set" (e.g., "item set") or "subset" should be interpreted as a non-empty set including one or more members. Furthermore, unless otherwise indicated or contradicted by context, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather a subset and a corresponding set may be equivalent.
[0089] Unless expressly indicated otherwise or clearly contradicted by context, conjunctions such as phrases of the form "at least one of A, B, and C" or "at least one of A, B and C" are understood in context to be generally used to refer to an item, clause, or the like that may be A or B or C, or any non-empty subset of the set A and B and C. For example, in the illustrative example of a set having three members, the conjunction phrases "at least one of A, B, and C" and "at least one of A, B and C" refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Thus, such conjunction language is not generally intended to imply that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. In addition, unless expressly indicated otherwise or contradicted by context, the term "plurality" refers to a plural state (e.g., "plurality of items" means a plurality of items). The number of items in a plurality of items is at least two items, but may be more if expressly indicated or indicated by context. Further, the phrase "based on" means "based at least in part on" rather than "based solely on" unless otherwise specified or clear from context.
[0090] Unless otherwise indicated herein or clearly contradictory to the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that are jointly executed on one or more processors by hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of, for example, a computer program that includes a plurality of instructions that can be executed by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagated transient electrical or electromagnetic transmissions), but includes non-transitory data storage circuits (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of being executed), causes the computer system to perform the operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes a plurality of non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the plurality of non-transitory computer-readable storage media lacks all the code, but the plurality of non-transitory computer-readable storage media stores all the code together. In at least one embodiment, the executable instructions are executed so that different instructions are executed by different processors, for example, a non-transitory computer-readable storage medium stores instructions, and a main central processing unit ("CPU") executes some instructions, while a graphics processing unit ("GPU") executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and different processors execute different subsets of instructions.
[0091] Thus, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform the operations of the processes described herein, and such a computer system is configured with applicable hardware and / or software that enables the implementation of the operations. In addition, a computer system that implements at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system that includes multiple devices that operate in different ways, so that the distributed computer system performs the operations described herein, and so that a single device does not perform all operations.
[0092] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate embodiments of the present disclosure and does not limit the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any non-claimed element is essential to practicing the disclosure.
[0093] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0094] In the specification and claims, the terms "coupled" and "connected," as well as their derivatives, may be used. It should be understood that these terms may not be intended as synonyms for each other. On the contrary, in specific examples, "connected" or "coupled" may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0095] Unless explicitly stated otherwise, it is to be understood that throughout the specification, terms such as “processing”, “computing”, “calculating”, “determining” and the like refer to the actions and / or processes of a computer or computing system or similar electronic computing device that processes and / or converts data represented as physical quantities (e.g., electronic) in registers and / or memories of the computing system into other data similarly represented as physical quantities in the memories, registers or other such information storage, transmission or display devices of the computing system.
[0096] In a similar manner, the term "processor" may refer to any device or part of a memory that processes electronic data from registers and / or memory and converts the electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" may be a CPU or a GPU. A "computing platform" may include one or more processors. As used herein, a "software" process may include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process may refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently. The terms "system" and "method" may be used interchangeably herein, as long as a system may embody one or more methods, and a method may be considered a system.
[0097] In this document, reference may be made to obtaining, acquiring, receiving or inputting analog or digital data into a subsystem, a computer system or a computer-implemented machine. The process of obtaining, acquiring, receiving or inputting analog and digital data can be accomplished in a variety of ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In some implementations, the process of obtaining, acquiring, receiving or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending or presenting analog or digital data can be accomplished by transmitting data as an input or output parameter of a function call, an application programming interface or an interprocess communication mechanism.
[0098] Although the above discussion sets forth example implementations of the described techniques, other architectures may be used to implement the described functionality and are intended to fall within the scope of the present disclosure. In addition, although specific responsibilities are defined above for discussion purposes, various functions and responsibilities may be allocated and divided in different ways, depending on the circumstances.
[0099] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as exemplary forms of implementing the claims.
Claims
1. A method, include: Generate fast history frames for rendering frames; determining, from the fast history frame, a set of fast pixel values corresponding to pixel locations of the rendered frame; determining a range of expected pixel values based at least in part on the fast pixel value; determining that a historical pixel value from a complete historical frame is outside of the range of expected pixel values; clamping the historical pixel values from the full historical frame to clamped pixel values within the range of expected pixel values; as well as As part of an image reconstruction process, pixel values at the pixel locations of the rendered frame are blended with clamped values.
2. The method according to claim 1, further comprising: include: The images obtained from the image reconstruction process are provided to a time accumulation process, which is used to update the fast history frame and the full history frame and stored in respective buffers for the image reconstruction process.
3. The method according to claim 2, further comprising: include: A first time accumulation weight for generating the fast history frame and a second time accumulation weight for generating the complete history frame are determined, the first time accumulation weight being higher than the second time accumulation weight.
4. The method according to claim 2, further comprising: include: An initial blend of the rendered frame and the fast history frame is performed.
5. The method according to claim 2, further comprising: include: The images resulting from the image reconstruction process are provided for display as part of a sequence of images.
6. The method of claim 1, wherein the set of fast pixel values is determined from a neighborhood of pixels surrounding the pixel location in the fast history image.
7. The method according to claim 1, further comprising: include: determining a second set of fast pixel values corresponding to pixel locations from a second rendered frame; determining a second range of expected pixel values based at least in part on the second set of fast pixel values; determining that an updated historical pixel value from the fast history frame is within the second range of expected pixel values; and The updated historical pixel value from the full history frame is used when blending with a second current pixel value at the pixel location of the second rendered frame.
8. The method according to claim 1, in, The expected range corresponds to a region in a color space, the region being one of a boundary shape, a convex hull, or an amorphous region around a position of the current pixel value in the color space.
9. The method of claim 8, wherein clamping the historical pixel values comprises at least one of applying a blending / maximum threshold for each color component, determining a closest value in the color space, or determining an intersection of a color vector in the color space with a boundary of the region.
10. A system, include: processor; as well as a memory comprising instructions that, when executed by the processor, cause the system to: Generate fast history frames for rendering frames; determining a set of fast pixel values corresponding to pixel locations from the fast history frame; determining a range of expected pixel values based at least in part on the fast pixel value; determining that a historical pixel value from a complete historical frame is outside of the range of expected pixel values; clamping the historical pixel values from the full historical frame to clamped pixel values within the range of expected pixel values; as well as As part of an image reconstruction process, current pixel values at the pixel locations of the rendered frame are blended with clamped values.
11. The system according to claim 10, in, The instructions, when executed, further cause the system to: The images obtained from the image reconstruction process are provided to a time accumulation process, which is used to update the fast history frame and the full history frame and stored in respective buffers for the image reconstruction process.
12. The system according to claim 11, in, The instructions, when executed, further cause the system to: A first time accumulation weight for generating the fast history frame and a second time accumulation weight for generating the complete history frame are determined, the first time accumulation weight being higher than the second time accumulation weight.
13. The system according to claim 10, in, The set of pixel values is determined from a neighborhood of pixels around the pixel location in the fast history image.
14. The system according to claim 10, in, The instructions, when executed, further cause the system to: determining a second set of fast pixel values corresponding to pixel locations from a second rendered frame; determining a second range of expected pixel values based at least in part on the second set of fast pixel values; determining that an updated historical pixel value from the complete historical frame is within the second range of expected pixel values; and The updated historical pixel value from the full history frame is used when blending with a second current pixel value at the pixel location of the second rendered frame.
15. The system according to claim 10, in, The expected range corresponds to a region in a color space, the region being one of a boundary shape, a convex hull, or an amorphous region around the location of the current pixel value in the color space, and wherein clamping the second historical pixel value comprises at least one of applying a blending / maximum threshold for each color component, determining a closest value in the color space, or determining an intersection of a color vector in the color space and a boundary of the region.
16. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processor, cause the processor to: Generate fast history frames for rendering frames; determining a set of fast pixel values corresponding to pixel locations from the fast history frame; determining a range of expected pixel values based at least in part on the fast pixel value; determining that a historical pixel value from a complete historical frame is outside of the range of expected pixel values; clamping the historical pixel values from the full historical frame to clamped pixel values within the range of expected pixel values; as well as As part of an image reconstruction process, current pixel values at the pixel locations of the rendered frame are blended with clamped values.
17. The non-transitory computer readable storage medium of claim 16, in, The instructions, when executed, further cause the processor to: The image obtained by the image reconstruction process is provided to a time accumulation process, and the time accumulation process is used to update the fast history frame and the complete history frame and store them in respective buffers used for the image reconstruction process.
18. The non-transitory computer readable storage medium of claim 17, in, The instructions, when executed, further cause the processor to: A first time accumulation weight for generating the fast history frame and a second time accumulation weight for generating the complete history frame are determined, the first time accumulation weight being higher than the second time accumulation weight.
19. The non-transitory computer readable storage medium of claim 16, in, The instructions, when executed, further cause the processor to: determining a second set of fast pixel values corresponding to pixel locations from a second rendered frame; determining a second range of expected pixel values based at least in part on the second set of fast pixel values; determining that an updated historical pixel value from the complete historical frame is within the second range of expected pixel values; and The updated historical pixel value from the full history frame is used when blending with a second current pixel value at the pixel location of the second rendered frame.
20. The non-transitory computer-readable storage medium of claim 16, wherein the expected range corresponds to a region in a color space, the region being one of a boundary shape, a convex hull, or an amorphous region around the location of the current color value in the color space, and wherein clamping the second historical pixel value comprises at least one of applying a blending / maximum threshold for each color component, determining a closest value in the color space, or determining an intersection of a color vector in the color space with a boundary of the region.
Citation Information
Patent Citations
Shared k-SVD dictionary-based DVS visual video denoising method
CN107610069A
History-aware selective pixel shifting
CN108694904A