A locking mechanism for image classification.

JP2025509472A5Active Publication Date: 2026-03-06ADVANCED MICRO DEVICES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024554140
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-28
Filing Date
2023-03-17
Publication Date
2026-03-06
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

Traditional spatial amplifiers require high-quality anti-aliasing source images when improving image resolution, resulting in additional anti-aliasing in scenarios such as games that do not use anti-aliasing, which increases integration difficulty and performance consumption. Furthermore, when the source pixel resolution is extremely low, insufficient detail reconstruction leads to emerging artistic ifacts such as flicker and inadequate edge reconstruction.

Method used

Super-resolution amplification is achieved by reconstructing high-resolution images using time feedback, combined with image processing techniques such as backprojection accumulation and color correction. This method uses low-resolution rendered frames and previous enlarged frames to generate high-resolution images.

Benefits of technology

Improves image quality, reduces dependence on anti-aliasing, enhances the ability to rebuild details at low source pixel resolution, and reduces emerging art ifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A first frame of the video stream is obtained, the first frame being defined by a plurality of pixels associated with a set of color data, and a determination is made that a pixel of the plurality of pixels includes high frequency information, and in response to determining that the pixel includes high frequency information, a pixel lock is generated for the pixel such that color data associated with the pixel is maintained during a color accumulation process for at least one of the first frame or a second frame of the video stream subsequent to the first frame.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Typically, spatial upscalers implemented in graphics pipelines receive frames rendered at lower than native resolution to improve performance, and upscale these frames to match the native resolution of the display. Although easy to integrate, spatial upscalers also have drawbacks. For example, traditional spatial upscaling approaches usually require high-quality anti-aliased source images. Therefore, games that do not use anti-aliasing must implement anti-aliasing, which requires more time to integrate traditional approaches. Furthermore, traditional spatial upscalers often produce poor quality upscaled output when anti-aliasing is poorly implemented. Also, upscaling quality depends on the input source resolution. When the source resolution is very low, there is insufficient information to draw from to reproduce thin details, which leads to new artifacts such as shimmering and poor edge reconstruction.

[0002] The present disclosure may be better understood, and its numerous features and advantages made apparent to those skilled in the art by reference to the following drawings, in which: The use of the same reference numbers in different drawings indicates similar or identical items. [Brief description of the drawings]

[0003] [Figure 1] FIG. 2 is a block diagram of an example device for implementing the super-resolution upscaling techniques described herein, in accordance with some embodiments. [Diagram 2] 2 is a more detailed block diagram of the example device of FIG. 1 illustrating further details associated with executing processing tasks in one or more accelerated processing devices, in accordance with some embodiments. [Diagram 3] FIG. 3 is a block diagram of the graphics pipeline of FIG. 2 according to some embodiments. [Figure 4] FIG. 2 is a synthesis diagram of a graphics pipeline for a super-resolution upscaler according to some embodiments. [Diagram 5] 5 is a block diagram illustrating a more detailed diagram of the super-resolution upscaler of FIG. 4 according to some embodiments. [Figure 6] 6 is a block diagram illustrating various stages / paths of the super-resolution upscaler of FIG. 5 according to some embodiments. [Figure 7] 6 illustrates a downsampling technique of the auto-exposure component of the super-resolution upscaler of FIG. 5 according to some embodiments. [Figure 8] 6 illustrates an exemplary enhanced depth value and motion vector construction process performed by the reconstruction enhancement component of the super-resolution upscaler of FIG. 5 according to some embodiments. [Figure 9] 6 is a block diagram illustrating various sub-stages of the reprojection accumulate stage of the super-resolution upscaler of FIG. 5 according to some embodiments. [Figure 10] 6 illustrates an exemplary resampling process performed by the upsampling component of the super-resolution upscaler of FIG. 5 according to some embodiments. [Figure 11] 6 illustrates an exemplary color bounding box construction process performed by the upsampling component of the super-resolution upscaler of FIG. 5 in accordance with some embodiments. [Figure 12] 6 illustrates an exemplary color data reprojection process performed by the reprojection component of the super-resolution upscaler of FIG. 5 in accordance with some embodiments. [Figure 13] 6 illustrates an exemplary color correction process performed by the color corrector component of the super-resolution upscaler of FIG. 5 in accordance with some embodiments. [Figure 14] 6 illustrates an example filter performed by the image sharpening component of the super-resolution upscaler of FIG. 5 in accordance with some embodiments. [Figure 15]FIG. 1 is a flow diagram illustrating an overall method for spatially upscaling rendered frames of a video stream by reconstructing a high-resolution image representative of the rendered frames using temporal feedback, according to some embodiments. [Figure 16] FIG. 1 is a flow diagram illustrating another general method for spatially upscaling rendered frames of a video stream by reconstructing a high-resolution image representative of the rendered frames using temporal feedback, according to some embodiments. [Figure 17] FIG. 1 is a flow diagram illustrating another general method for spatially upscaling rendered frames of a video stream by reconstructing a high-resolution image representative of the rendered frames using temporal feedback, according to some embodiments. [Figure 18] FIG. 18 is a flow diagram illustrating a more detailed method of the reprojection accumulation process shown in block 1616 of FIG. 17, according to some embodiments. [Figure 19] FIG. 18 is a flow diagram illustrating a more detailed method of the reprojection accumulation process shown in block 1616 of FIG. 17, according to some embodiments. [Figure 20] FIG. 2 is a flow diagram illustrating an overall method of a pixel locking process shown in accordance with some embodiments. [Figure 21] FIG. 18 is a flow diagram illustrating a more detailed method of the pixel locking process shown in block 1614 of FIG. 17, according to some embodiments. [Figure 22] 18. A flow diagram illustrating a more detailed method for the reprojection shown in block 1806 of FIG. 18 and the lock update process shown in block 1808 of FIG. 18, according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0004] As described below, an exemplary approach can improve upscaling by using temporal feedback to reconstruct high-resolution images while preserving or improving image quality compared to native rendering, allowing for "practical implementations" of costly rendering operations such as hardware ray tracing.

[0005] FIG. 1 illustrates an example device 100 capable of implementing one or more of the functions described herein, such as super resolution upscaler 332 (FIG. 3). Device 100 may include, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, a tablet computer, a wearable computing device, etc. In one or more embodiments, device 100 includes one or more processors 102, memory 104, a storage device 106, an input device 108, and an output device 110. In at least some embodiments, device 100 includes one or more input drivers 112 and one or more output drivers 114. It should be understood that device 100 may include additional components not shown in FIG. 1.

[0006] In one or more embodiments, the processor 102 includes a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU located on the same die, or one or more processor cores, each of which may be a CPU or a GPU. In at least some embodiments, the CPU includes one or more single-core or multi-core CPUs. In various alternatives, the memory 104 may be located on the same die as the processor 102 or may be located separately from the processor 102. The memory 104 may include, for example, volatile or non-volatile memory, such as random access memory (RAM), dynamic RAM, or cache.

[0007] Storage 106 may include, for example, fixed or removable storage such as a hard disk drive, solid state drive, optical disk, or flash drive. Input devices 108 may include, for example, a keyboard, a keypad, a touch screen, a touch pad, a detector, a microphone, an accelerometer, a gyroscope, a biometric scanner, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals). Output devices 110 may include, for example, a display, a speaker, a printer, a haptic feedback device, one or more lights, an antenna, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals).

[0008] The input driver 112 communicates with the processor 102 and the input device 108, allowing the processor 102 to receive input from the input device 108. The output driver 114 communicates with the processor 102 and the output device 110, allowing the processor 102 to send output to the output device 110. It should be noted that the input driver 112 and the output driver 114 are optional components, and the device 100 would operate similarly without the input driver 112 and the output driver 114. The device 100 also includes one or more accelerated processing devices (APDs) 116. In at least some embodiments, the APD 116 is coupled to, for example, one or more display devices 118. The APD 116 may be part of the output device 110 or may be separate from the output device. The APD 116 accepts compute commands and graphics rendering commands from the processor 102, processes the compute commands and graphics rendering commands, and provides pixel output to the display device 118 for display.

[0009] As described in more detail below, the APD 116 includes one or more parallel processing units for performing calculations according to the single instruction, multiple data (SIMD) paradigm. In one or more embodiments, the APD 116 is used to implement a GPU, and in these embodiments, the parallel processing units are referred to as shader cores or streaming multiprocessors (SMX). Each parallel processing unit includes one or more processing elements, such as a scalar and / or vector floating point unit, an arithmetic logic unit (ALU), etc. In various embodiments, the parallel processing units also include special purpose processing units (not shown), such as an inverse square root unit and a sine / cosine unit.

[0010] Although various functions are described herein as being performed by or in conjunction with APD 116, in various alternatives, the functions described as being performed by APD 116 are not driven by a host processor (e.g., processor 102) but are additionally or alternatively performed by other computing devices having similar functionality for providing graphical output to display device 118. For example, any processing system that performs processing tasks according to the SIMD paradigm is capable of performing the functions described herein. Alternatively, computing systems that do not perform processing tasks according to the SIMD paradigm are similarly capable of performing the functions described herein.

[0011] FIG. 2 is a block diagram of device 100 illustrating further details related to the execution of processing tasks on APD 116. In one or more embodiments, processor 102 maintains one or more control logic modules in system memory 104 for execution by processor 102. The control logic modules include, for example, operating system 202, kernel mode driver 204, and applications 206. These control logic modules control various features of the operation of processor 102 and APD 116. For example, operating system 202 communicates directly with hardware and provides an interface to the hardware for other software executing on processor 102. Kernel mode driver 204 controls the operation of APD 116, for example, by providing application programming interfaces (APIs) to software executing on processor 102 (e.g., applications 206) to access various features of APD 116. In one or more embodiments, kernel mode driver 204 also includes a just-in-time compiler that compiles programs for execution by processing components (e.g., SIMD units 210) of APD 116.

[0012] In one or more embodiments, APD 116 includes any cooperating collection of hardware and / or software that performs functions and calculations related to accelerating graphics processing tasks, data parallel tasks, nested data parallel tasks in an accelerated manner for resources such as traditional CPUs, traditional GPUs, and combinations thereof. GPUs, accelerated processing units (APUs), and general purpose graphics processing units (GPGPUs) are terms commonly used in the art, although the term "accelerated processing device (APD)" as used herein is a broader term.

[0013] The APD 116 executes commands and programs for selected functions, such as graphics operations and non-graphics operations that may be suitable for parallel processing. In one or more embodiments, the APD 116 is used to perform graphics pipeline operations, such as pixel operations, geometry operations, and image rendering to the display device 118, based on commands received from the processor 102. The APD 116 also performs operations not directly related to graphics operations, such as operations related to video, physics simulations, computational fluid dynamics, and the like, based on commands received from the processor 102. For example, such commands include special instructions not typically defined in the instruction set architecture (ISA) of the APD 116. In at least some embodiments, the APD 116 receives image geometry representing a graphics image along with one or more commands or instructions for rendering and displaying the image. In various embodiments, the image geometry corresponds to a representation of a two-dimensional (2D) or three-dimensional (3D) computer graphics image.

[0014] APD 116 includes compute units 208 (shown as compute units 208-1 through 208-3), such as a processing core that includes one or more SIMD units 210 (shown as SIMD units 210-1 through 210-6) that perform operations in parallel according to the SIMD paradigm as required by processor 102. The SIMD paradigm allows multiple processing elements to share a single program control flow unit and program counter to execute the same program, but execute the program on different data. In one embodiment, each SIMD unit 210 includes a number of lanes, each lane executing the same instructions simultaneously with other lanes in SIMD unit 210, but may execute the instructions on different data. Lanes may be switched off with prediction when not all lanes are required to execute a given instruction. Prediction may also be used to execute programs with branching control flows. More specifically, for programs with conditional branch instructions or other instructions where control flow is based on calculations performed by individual lanes, prediction of lanes corresponding to currently not executed control flow paths and serial execution of different control flow paths allows arbitrary control flow.

[0015] The basic unit of execution in the compute unit 208 is the work item. Each work item represents a single instantiation of a program that executes in parallel on a particular lane. Work items can execute simultaneously on a single SIMD unit 210 as a "wavefront." One or more wavefronts are included in a "workgroup," which is a collection of work items designated to execute the same program. A workgroup can be executed by executing each of the wavefronts that make up the workgroup. Alternatively, a wavefront may execute sequentially on a single SIMD unit 210, or may execute partially or fully in parallel on different SIMD units 210. A wavefront is the largest collection of work items that can execute simultaneously on a single SIMD unit 210. Thus, if commands received from the processor 102 indicate that a particular program is parallelized to an extent that it cannot be executed simultaneously on a single SIMD unit 210, the program is split into wavefronts that are parallelized on two or more SIMD units 210, or serialized on the same SIMD unit 210 (or both parallelization and serialization are performed, as appropriate). The scheduler 212 performs operations related to scheduling various wavefronts on the different compute units 208 and SIMD units 210.

[0016] The parallelism provided by one or more compute units 208 is well suited for graphics-related operations such as pixel value calculations, vertex transformations, tessellation, geometry shading operations, and other graphics operations. Thus, in one or more embodiments, the graphics processing pipeline 214 receives graphics processing commands from the processor 102 and provides computational tasks to the compute units 208 for parallel execution. Some graphics pipeline operations, such as pixel processing and other parallel computational operations, require the same command stream or computation kernel to be executed on a stream or collection of input data elements. Each instantiation of the same compute kernel executes simultaneously on multiple SIMD 210 units in one or more compute units 208 to process such data elements in parallel. As referred to herein, for example, a compute kernel is a function that includes instructions that are declared in a program and executed on the APD compute units 208. The function may also be referred to as a kernel, a shader, a shader program, or a program.

[0017] The graphics processing pipeline 214 includes multiple stages (not shown in FIG. 2 for simplicity) configured to process various primitives simultaneously in response to a draw call. In one or more embodiments, the stages of the graphics processing pipeline 214 simultaneously process various primitives generated by an application 206, such as a video game. When geometry data is sent to the graphics processing pipeline 214, a hardware state setting is selected to define the state of the graphics processing pipeline 128. Examples of states include rasterizer state, blend state, depth stencil state, primitive topology type of the submitted geometry, as well as shaders (e.g., vertex shader, domain shader, geometry shader, hull shader, pixel shader, etc.) used in rendering the scene. Shaders implemented in the graphics pipeline state are represented by corresponding bytecodes. In some cases, information representing the graphics pipeline state is hashed or compressed to provide a more efficient representation of the graphics pipeline state.

[0018] As described in more detail below, the APD 116 is configured to implement features of the present disclosure by performing a number of functions. For example, the APD 116 is configured to implement a super resolution upscaler 332 that receives a low resolution rendered frame 502 of a video stream. The super resolution upscaler 332 spatially upscales the low resolution rendered frame 502 using temporal feedback (e.g., previously upscaled frame(s) of the video stream) to reconstruct a high resolution frame 508 that represents the rendered frame.

[0019] Figure 3 illustrates further details of the graphics pipeline 214 shown in Figure 2, according to one or more embodiments. As described below, the graphics pipeline 214 can perform super-resolution upscaling using temporal feedback (e.g., data associated with one or more previously upscaled frames) to reconstruct a high-resolution image while maintaining or improving image quality compared to native rendering. In at least some embodiments, the graphics pipeline 214 is implemented in the APD 116 shown in Figures 1 and 2. It should be understood that the super-resolution upscaling techniques described herein can also be applied to other pipeline configurations.

[0020] In one or more embodiments, the graphics processing pipeline 214 is configured to render graphics as images that depict a scene having three-dimensional geometry in a virtual space (also referred to herein as "world space"), but possibly two-dimensional geometry. Typically, the graphics processing pipeline 214 receives a representation of a three-dimensional scene, processes the representation, and outputs a two-dimensional raster image. These stages of the graphics processing pipeline 214 first process data that is properties of the endpoints (or vertices) of geometric primitives, which provide information about the object being rendered. Typical primitives in three-dimensional graphics include triangles and lines, and the vertices of these geometric primitives provide information such as xyz coordinates, textures, and reflectance.

[0021] The graphics pipeline 214 has access to storage resources 334 (also referred to herein as "storage components"), such as one or more memories or hierarchies of caches used to implement buffers, store vertex data, texture data, and the like. The storage resources 334 are implemented, for example, using some embodiments of the system memory 104 shown in FIGS. 1 and 2. Some embodiments of the storage resources 334 include (or have access to) one or more caches 336, one or more random access memory units 338, video random access memory unit(s) (not shown), one or more processor registers (not shown), and the like, depending on the nature of the data at a particular stage of rendering. It will thus be understood that the storage resources 334 refers to processor-accessible memory utilized by the graphics processing pipeline 214.

[0022] Graphics processing pipeline 214 includes stages, each of which performs a specific function. The stages represent a subdivision of the functionality of graphics processing pipeline 214. Each stage may be implemented partially or wholly as a shader program running on a programmable processing unit, such as SIMD unit 210 of FIG. 2, or partially or wholly as fixed-function, non-programmable hardware external to the programmable processing unit. Stages 301 and 303 represent the front-end geometry processing portion of graphics processing pipeline 214 before rasterization. Stages 305 through 311 represent the back-end pixel processing portion of graphics processing pipeline 214.

[0023] In the input assembler stage 301 of the graphics pipeline 214, the input assembler 302 is configured to access information from storage resources 334 that is used to define objects that represent portions of a model of a scene. For example, in various embodiments, the input assembler stage 220 reads primitive data (e.g., points, lines, and / or triangles) from a user-populated buffer (e.g., a buffer populated at the request of software executed by the processor 102, such as an application 206) and assembles the data into primitives that are used by other pipeline stages of the graphics processing pipeline 200. As used herein, the term “user” refers to an application 206 or other entity that provides shader code and three-dimensional objects for rendering to the graphics processing pipeline 214. The term “user” is used to distinguish the overall activity performed by the APD 116. The input assembler 302 assembles vertices into a number of different primitive types (such as line lists, triangle strips, or primitives with adjacency) based on primitive data contained in user-supplied buffers, and formats the assembled primitives for use by the rest of the graphics processing pipeline 214.

[0024] In one or more embodiments, the graphics processing pipeline 214 operates on one or more virtual objects that are defined by a set of vertices established in world space and have geometry defined relative to coordinates in a scene. For example, input data utilized by the graphics processing pipeline 214 may include a polygon mesh model of the scene geometry whose vertices correspond to primitives processed in the rendering pipeline in accordance with aspects of the present disclosure, and this initial vertex geometry is established within the storage resources 334 during the application stage executed by the CPU.

[0025] In the vertex processing stage 303 of the graphics pipeline 214, one or more vertex shaders 304 process the vertices of the primitives via the input assembler 302. For example, a vertex shader 304 receives as input a single vertex of a primitive and outputs a single vertex. The vertex shaders 304 perform various per-vertex operations such as transformation, skinning, morphing, and per-vertex lighting. Transformation operations include various operations for transforming the coordinates (e.g., XY coordinates and Z depth values) of the vertices. These operations include one or more of a modeling transformation, a view transformation, a projection transformation, a perspective division, and a viewport transformation. As used herein, such transformations are considered to modify the coordinates or "position" of the vertex at which the transformation is performed. Other operations of the vertex shader 304 change attributes other than coordinates.

[0026] In one or more embodiments, the vertex shader(s) 304 are implemented partially or wholly as a vertex shader program that executes on one or more compute units 208. The vertex shader program is based on a program provided by the processor 102 and pre-written by a computer programmer. The kernel mode driver 204 compiles such a computer program to generate a vertex shader program having a format suitable for execution within the compute unit 208. Some embodiments of shaders, such as the vertex shader 304, provide large-scale single instruction multiple data (SIMD) processing such that multiple vertices are processed simultaneously. In at least some embodiments, the graphics pipeline 214 implements a unified shader model such that all shaders included in the graphics pipeline 214 have the same execution platform on a shared large-scale SIMD compute unit 210. In such embodiments, the shaders, including the vertex shader(s) 304, are implemented using a common set of resources, referred to herein as a unified shader pool 306.

[0027] In one or more embodiments, the vertex processing stage 303 performs additional vertex processing calculations that subdivide primitives and generate new vertices and new geometry in world space. In at least some embodiments, these additional vertex processing calculations are performed by one or more of the hull shader 308, the tessellator 310, the domain shader 312, and the geometry shader 314. The hull shader 308 operates on input high-order patches or control points that are used to define input patches. The hull shader 308 outputs tessellation factors and other patch data. In one or more embodiments, the primitives generated by the hull shader 308 are provided to the tessellator 310. The tessellator 310 receives an object (e.g., a patch) from the hull shader 308 and generates information identifying a primitive that corresponds to the input object, for example, by tessellating the input object based on the tessellation factors provided to the tessellator 310 by the hull shader 308. Tessellation divides input high-order primitives, such as patches, into a set of lower-order output primitives that represent finer-grained detail, e.g., as indicated by a tessellation factor that specifies the granularity of the primitives produced by the tessellation process. Thus, a model of a scene is represented by a smaller number of high-order primitives (to save memory or bandwidth), and additional detail is added by tessellating the high-order primitives.

[0028] The domain shader 312 inputs the domain location and (optionally) other patch data. The domain shader 312 operates on the information provided and generates a single vertex for output based on the input domain location and other information. The geometry shader 314 receives an input primitive and outputs up to four primitives that the geometry shader 314 generates based on the input primitive. In some embodiments, the geometry shader 314 retrieves vertex data from the storage resource 334 and generates new graphics primitives, such as lines and triangles, from the vertex data in the storage resource 334. Specifically, the shader 314 retrieves the vertex data of an entire primitive and generates zero or more primitives. For example, the geometry shader 314 can operate on a triangle primitive with three vertices. The geometry shader 314 can perform a variety of different types of operations, including operations such as point sprint expansion, dynamic particle system operations, far fin generation, shadow volume generation, single pass render to cube map, per-primitive material swapping, and per-primitive material setup. In at least some embodiments, one or more of the hull shader 308, domain shader 312, or geometry shader 314 are implemented as shader programs executing on a programmable processing unit, such as the SIMD unit 210, while the tessellator 310 is implemented by fixed-function hardware.

[0029] Once front-end processing is complete, the scene is defined by a set of vertices, each with a set of vertex parameter values ​​stored in storage resources 334. In one particular embodiment, the vertex parameter values ​​output from the vertex processing stage 303 include positions defined in different homogeneous coordinates for different zones.

[0030] As noted above, stages 305 through 311 represent the back-end processing of the graphics processing pipeline 214. The rasterizer stage 305 includes a rasterizer 316 that accepts and rasterizes simple primitives and upstream generated images. The rasterizer 316 performs shading operations as well as operations such as clipping, perspective division, scissoring, and viewport selection. The rasterizer 316 generates a set of pixels that are subsequently processed in the pixel processing / shader stage 307 of the graphics pipeline 214. In some embodiments, the set of pixels includes one or more tiles. In one or more embodiments, the rasterizer 316 is implemented by fixed function hardware.

[0031] The pixel processing stage 307 includes one or more pixel shaders 318 that input a pixel flow (e.g., including a set of pixels generated by the rasterizer 316) and output zero or more pixel flows in response to the input pixel flow. The pixel shaders 318 calculate output values ​​for screen pixels based on the primitives generated upstream and the results of rasterization. In one or more embodiments, the pixel shaders 318 apply textures from a texture memory, which may be implemented as part of the storage resources 334. Each of the output(s) generated by the one or more pixel shaders 318, such as color values, depth values, and stencil values, is stored in one or more corresponding buffers, such as a color buffer 320, a depth buffer 322, and a stencil buffer 324. The combination of the color buffer 320, the depth buffer 322, and optionally the stencil buffer 324, is referred to as a frame buffer 326. In at least some embodiments, the graphics pipeline 214 implements multiple frame buffers 326, including intermediate buffers, such as a front buffer, a back buffer, a render target, a frame buffer object, and the like. The operations of the pixel shader 318 are performed by shader programs executing on the programmable processing unit 210.

[0032] In one or more embodiments, the pixel shader 318 or another shader accesses shader data, such as texture data stored in the storage resources 334. The texture data defines textures, which are bitmap images used in various parts of the graphics processing pipeline 214. For example, in some cases, the pixel shader 318 applies textures to pixels to improve the apparent rendering complexity (e.g., provide a more "photorealistic" appearance) without increasing the number of vertices rendered. In other cases, the vertex shader 304 uses the texture data to modify primitives, increasing the complexity, for example, by generating or modifying vertices to improve aesthetics. For example, the vertex shader stage 304 uses a height map stored in the storage resources 334 to change the displacement of vertices. This type of technique can be used, for example, to generate more realistic looking water by changing the position and number of vertices used in rendering water, as compared to using only textures in the pixel processing stage 307. In some cases, the geometry shader 314 accesses the texture data from the storage resources 334.

[0033] The output merger stage 309 includes an output merger 328 that accepts outputs from the pixel processing stage 307, merges those outputs, and performs operations such as Z-test, alpha blending, stenciling, etc., on each pixel received from the pixel shader 318 to determine the final color of the screen pixel. For example, the output merger 328 combines various types of output data (e.g., pixel shader values, depth and stencil information, etc.) with the contents of the color buffer 320, depth buffer 322, and optionally stencil buffer 324, and stores the combined output back in the frame buffer 326. The output 309 of the merger stage can be referred to as rendered pixels that collectively form the rendered frame. In one or more embodiments, the output merger 328 is implemented by fixed function hardware.

[0034] A post-processing stage 311 is performed after the output merger stage 309. In the post-processing stage 311, one or more post-processors 330 operate on the rendered frame stored in the frame buffer 326 (or for each pixel) to apply one or more post-processing effects, such as ambient occlusion or tone mapping, before the frame is output to a display. In at least some embodiments, the post-processors 330 are implemented using one or more vertex shaders 304, one or more pixel shaders 318, etc. The post-processed frame is written to a frame buffer 326, such as a back buffer for display or an intermediate buffer for further post-processing. In at least some embodiments, the graphics pipeline 214 can include other shaders or components, such as a compute shader 340, a ray tracer 342, a mesh shader 344, etc., and communicate with one or more of the other components of the graphics pipeline 214. The vertex processing stage 303, the rasterizer stage 305, the pixel processing stage 307, and at least a portion of the post-processing stage 311 are collectively referred to herein as the “renderer 313” or “rendering stage 313” of the graphics pipeline 214.

[0035] In many cases, the amount of processing resources required to render a full high-resolution image can make it difficult to render frames while meeting current frame rates, such as at least 60 frames per second (fps). Thus, in one or more embodiments, the application 206 (e.g., a video game) uses the renderer 313 to generate rendered images at a lower render resolution or size (e.g., 1920×1080 pixels) than one or more final output / presentation resolutions or sizes (e.g., 3840×2160 pixels) to meet timing requirements and reduce processing resource requirements. In at least some embodiments, the super-resolution upscaling stage 315, including the super-resolution upscaler 332 (referred to herein as “upscaler 332” for brevity), processes the low-resolution rendered images to generate upscaled images that represent the content of the low-resolution rendered images at a resolution equal to (or at least close to) the target presentation resolution. In one or more embodiments, the upscaling stage 315 is performed as part of the post-processing stage 311. However, in other embodiments, upscaler 332 is part of a different processing stage of graphics pipeline 214. In one or more embodiments, upscaler 332 is implemented using one or more vertex shaders 304, pixel shaders 318, compute shaders 340, etc., or a combination thereof.

[0036] The upscaler 332 improves the rendering performance of an application by implementing a temporal upscaling algorithm that operates on multiple inputs. Thus, in one or more embodiments, the upscaler 332 is placed in the graphics pipeline 214 at a position that ensures a balance between the best visual quality and performance. Placing other image space processes before the upscaler has the advantage that these other image space processes run at a lower resolution, providing performance benefits to the application 206. However, this placement may not be appropriate for some types of image space processes / techniques. For example, some image space processes may introduce noise or grain into the final image (e.g., to simulate a physical camera). Placing such image space processes before the upscaler may result in the upscaler amplifying the noise and introducing undesirable artifacts into the final upscaled image. Thus, in one or more embodiments, the upscaler 332 is typically placed between the post-processor 330 that operates on the current frame at the render resolution and the post-processor 330 that operates on the upscaled frame generated by the upscaler 332 at the presentation resolution. However, in at least some embodiments, techniques such as noise and grain can be calculated at a lower resolution and then applied / composited to the upscaled frame, for example using bilinear sampling.

[0037] FIG. 4 shows a graphics pipeline integration diagram 400 of the upscaler 332. In this embodiment, the first stage is the rendering portion 313 of the graphics pipeline 214, which renders the current frame at a given render resolution in a render color space. The pre-upscale post-processing stage 311-1 receives the rendered frame as input and performs one or more post-processing operations on the rendered frame at the render resolution and in the render color space. The pre-upscale post-processing stage 311-1 typically includes a post-processor 330 that performs post-processing operations that access the rendered frame's depth buffer 322. Examples of such post-processing operations include screen space reflection, screen space ambient occlusion, noise removal (e.g., shadows or reflections), exposure, etc. The upscaling stage 315 receives the output of the pre-upscale post-processing stage 311-1 (or rendering stage 313) as input and performs at least one of an upscaling or anti-aliasing operation on the pre-scaled post-processed frame or the rendered frame (if no post-processing has been performed) in a linear color space. The output of the upscaling stage 315 is at least one of an upscaled and / or anti-aliased frame having a target presentation resolution in the presentation color space. The post-upscale post-processing stage 311-2 receives the upscaled (and anti-aliased) frame as input and performs one or more post-processing operations on the upscaled frame at the presentation resolution and in the presentation color space. The post-upscale post-processing stage 311-2 typically uses the anti-aliased frame or includes a post-processor 330 that performs operations to add effects to the frame that would otherwise result in undesirable artifacts when the frame is upscaled. Examples of such post-processing operations include film grain, chromatic aberration, vignetting, tone mapping, blooming, depth of field, motion blur, etc.The final stage 417 in the embodiment shown in FIG. 4 is the presentation of a user interface or heads-up display, which the graphics pipeline 214 renders at a presentation resolution and within a presentation color space.

[0038] 5 is a block diagram illustrating a more detailed view of the upscaler 332. As described below, the upscaler 332 is temporal and takes as input data associated with a currently rendered (aliased) frame 502 (also referred to herein as the "current frame 502" or the "input frame 502") and data related to a previously presented (upscaled) frame. In one or more embodiments, the rendered frame 502 and the previously upscaled frame are frames of a video stream. The previously presented frame is the most recent frame that was processed by the upscaler 332 and presented to the user at the target presentation resolution. In one or more embodiments, the input data associated with the previously presented frame includes an output buffer 506-1 that holds data, such as a color buffer (e.g., output color), depth buffer, etc., for the previously presented frame. The resolution / size of the output buffer 506-1 corresponds to the presentation resolution.

[0039] In one or more embodiments, the input data associated with the current frame 502 includes various buffers provided by the application 206, such as the color buffer 320-1, the depth buffer 322-1, and the motion vector buffer 504. The input color buffer 320-1 is a render resolution color buffer for the current frame 502 provided by the application 206. In one or more embodiments, the input color buffer 320-1 includes color data, such as pixel color values, that are generated based on sub-pixel jittering performed during rendering of the current frame 502. Sub-pixel jittering can be performed for each frame of a sequence of frames during the rendering stage. During sub-pixel jittering, the center point of color determination for a pixel of a frame is shifted slightly to another point within the pixel using a sub-pixel offset, and color data is determined for this offset location. In one or more embodiments, the output merger 328 of the graphics pipeline 214 combines the multi-sampled color data determined for each sub-pixel sample to determine the final color data for the corresponding pixel. The output merger 328 then stores the final color data in the color buffer 320-1 for frame 502. In one or more embodiments, the jitter positions are determined randomly or according to a set pattern or sequence, such as a Halton sequence (e.g., Halton[2,3]). The Halton sequence provides spatially separated points that cover the available pixel space. In one or more embodiments, jittering is applied to the rendering of multiple object types, including opaque objects, alpha transparent objects, and ray traced objects. For rasterized objects, a sub-pixel jittering value can be applied to a camera projection matrix, which is then used to perform a transformation during vertex shading. For ray traced rendering, the sub-pixel jittering is applied to the origin of the ray, which is often the camera position.

[0040] In some configurations, the input depth buffer 322-1 includes the stencil buffer 324. However, in other configurations, the input depth buffer 322-1 is stored separately from the stencil buffer 324. The input depth buffer 322-1 is a render resolution depth buffer for the current frame 502 provided by the application 206. In one or more embodiments, the resolution / size of the input color buffer 320-1 and the input depth buffer 322-1 is equal to the render resolution. The motion vector buffer 504 contains two-dimensional motion vectors that encode the motion from a pixel in the current frame 502 to the location of the same pixel in the previous (upscaled) frame. In one or more embodiments, the motion vectors are provided by the application 206, for example, in the range [(<-width,-height>...<-width,height>)]. For example, a value for a pixel in the top left corner of the screen is<width,height> The motion vectors represent movement across the full width and height of the input surface, originating from the bottom right corner of the screen. In at least some configurations, one or more of opaque, alpha tested, or alpha blended objects write their motion vectors for every pixel they cover. If vertex shader effects such as UV scrolling are applied, these calculations are factored into the motion calculations to produce enhanced results. In some configurations, the resolution / size of the motion vector buffer 504 is equal to the render resolution. However, in other configurations, the resolution / size of the motion vector buffer 504 is equal to the presentation resolution. In one or more embodiments, the color buffer 320-1, depth buffer 322-1, and motion vector buffer 504 are texture data types, although other data types are equally applicable.

[0041] In at least some embodiments, the upscaler 332 processes additional external input resources, such as a reactive mask 602 (FIG. 6) and an exposure 604 (FIG. 6) of the rendered frame 502, when upscaling the rendered frame 502. In the context of the upscaler 332, "reactive" refers to how much influence the samples rendered for the current frame 502 have on the generation of the final upscaled image. In one or more embodiments, the samples rendered for the current frame 502 contribute a relatively small amount to the results calculated by the upscaler 332, with exceptions. For example, to produce the best results for fast moving alpha blended objects, one or more stages of the upscaler 332 (e.g., the reprojection accumulation stage 611 (FIG. 6)) can be configured to be more reactive to such pixels. Because it can be difficult to determine which pixels have been rendered using alpha blending, either from color, depth, or motion vectors, in one or more embodiments, the application 206 provides a reactive mask 602 as an input to the upscaler 332. The reactive mask 602 provides a mechanism for the application 206 to identify areas of the rendered frame 502 that do not leave a footprint in the input depth buffer 322-1 or that contain motion vectors 504. In other words, the reactive mask 602 guides the upscaler 332 as to where it should rely less on past information when compositing the current pixel, and instead allow samples of the current frame 502 to contribute more to the final result. The reactive mask 602 allows the application 206 to provide a value, for example, from [0...1], where a value of 0 indicates that the pixel is not reactive at all (e.g., the upscaler 332 uses its default compositing strategy) and a value of 1 indicates that the pixel is fully reactive. In one or more embodiments, the alpha value used in compositing an alpha-blended object (eg, a particle) into a scene is implemented as a proxy for reactivity.In such an embodiment, the application 206 writes the alpha value of each covered pixel for the alpha-blended object to a corresponding pixel of the reactive mask 602. In one or more embodiments, the resolution / size of the reactive mask 602 is equal to the render resolution. Further, in one or more embodiments, the reactive mask 602 is a texture data type, although other data types are equally applicable.

[0042] Exposure 604 informs the upscaler 332 of the exposure value calculated by the application 206 for the rendered frame 502. In one or more embodiments, the exposure value matches the exposure that the application 206 will use during the next tone mapping pass. In at least some embodiments, the upscaler 332 calculates its own exposure values ​​for internal use at various stages. Furthermore, in one or more embodiments, the output generated by the upscaler 332 has this internal tone mapping inverted before the final output is written. That is, the upscaler 332 returns results in the same domain (or close to it) as the original input signal. In one or more embodiments, the exposure 604 is a texture data type, although other data types are equally applicable.

[0043] The upscaler 332 uses the current frame input and the previous frame input to generate a super-resolution upscaled (and anti-aliased) frame 508 (also referred to herein for brevity as the “upscaled frame 508”) at the target presentation resolution that corresponds to the rendered frame 502. In one or more embodiments, the upscaler 332 stores data representing the upscaled frame 508 in an output buffer 506-2. The upscaler 332 maintains an output buffer history 510 for accessing output buffers 506-1 generated for previously presented frames. The upscaler 332 implements one or more components for performing the upscaling and anti-aliasing operations described herein. For example, the upscaler 332 implements an auto exposure component 512, an input color adjustment component 514, a reconstruction expansion component 516, a depth clip component 518, a locking component 520, a reprojection accumulation component 522, and a sharpening component 524. 6, each of these components is implemented in a corresponding stage of upscaler 332. Although components 512-524 of upscaler 332 are shown implemented separately from one another, it should be understood that two or more of these components may be combined.

[0044] FIG. 6 illustrates various stages / paths of the upscaler 332, such as an auto exposure stage 601, an input color adjustment stage 603, a reconstruction expansion stage 605, a depth clip stage 607, a locking stage 609, a reprojection accumulation stage 611, and a sharpening stage 613. FIG. 6 also illustrates the inputs processed by each stage and the outputs generated by each stage. In the example illustrated in FIG. 6, the hatched boxes represent stages of the upscaler 332, the dashed boxes represent input / output buffers, and the solid boxes represent intermediate / processing buffers. In one or more embodiments, the application 206 interacts with the upscaler 332 through one of several different application programming interfaces (APIs). For example, the application 206 can instantiate the upscaler 332, issue one or more calls to the upscaler 332, pass one or more data structures and inputs to the upscaler 332, and so on, through the API. In one or more embodiments, when the upscaler 332 is instantiated (e.g., invoked by the application 206), storage resources 334, such as GPU local memory, are allocated for consumption by the upscaler 332. The upscaler 332 uses these storage resources 334 to store intermediate textures calculated by the upscaler 332 and also to store persistent textures across many frames of the application 206.

[0045] In the auto exposure stage 601, the application 206 provides the color buffer 320-1 of the currently rendered frame 502 as input to the auto exposure component 512. If the content of the color buffer 320-1 is high dynamic range (HDR), the application 206 can indicate this to the auto exposure component 512, for example, by setting a flag in a data structure provided by the application 206 to the upscaler 332. The auto exposure component 512 processes the color buffer 320-1 to generate up to two intermediate storage resources, depending on the configuration provided by the application 206. The first intermediate storage resource is a current luma texture 606, which is a low-resolution representation of the luma of the input colors. For example, the current luma texture 606 is a texture that is 50% (or some other percentage) of the render resolution texture (e.g., color buffer 320-1) that contains the luma values ​​of the currently rendered frame 502. The current luma texture 606 is used by the shading change detection process of the reprojection accumulation stage 611, as described below. A second intermediate storage resource is an exposure texture 604 that contains exposure values ​​calculated for the rendered frame 502. The exposure texture 604 is optionally used by subsequent stages of the upscaler 332 depending on the configuration of the upscaler 332. For example, the application 206 can indicate whether the exposure 604 should be used during the upscaling process, for example, by setting a flag in a data structure provided to the upscaler 332. In other embodiments, the exposure 604 is provided as an input by the application 206 to the auto exposure component 512. Optionally, the exposure 604 is used by the exposure calculations of the input color adjustment stage 603 to apply tone mapping, and further by the reprojection accumulation stage 611 to invert the local tone mapping before generating the output (e.g., the upscaled frame 508) by the upscaler 332. In one or more embodiments, the resolution / size of the exposure 604 is 1×1 pixel, although other sizes are equally applicable.

[0046] In at least some embodiments, the auto-exposure component 512 implements at least one downsampling mechanism, such as the AMD FidelityFX® Single Pass Downsampler, and uses shader dispatch, such as Single Compute Shader Dispatch, to generate a mipmap chain. A mipmap is a collection of successively lower resolution bitmap images of a texture. That is, a mipmap contains multiple versions of the same texture, each version at a different resolution. These different versions are sometimes referred to as "mipmap levels," "levels," or "mips." In one or more embodiments, unlike traditional pyramidal approaches, the downsampling mechanism implemented by the auto-exposure component 512 not only generates a specific set of mipmap levels for any input texture (e.g., input color buffer 320-1), but also performs some calculations on the data as it is stored to a target location in memory.

[0047] In one or more embodiments, the downsampling mechanism of the auto exposure component 512 is configured to write only to the second (e.g., half resolution) mipmap level 702 and the last (1x1) mipmap level 704, as illustrated in Figure 7. Furthermore, different calculations are applied at each of these levels 702 and 704 to calculate the amount to be used in the next stage of the upscaler 332. Thus, the remaining levels of the mipmap chain do not need to be backed by GPU local memory (or other types of memory). The second mipmap level 702 contains the current luminance 606, and the last mipmap level 704 contains the exposure 604, and these values ​​are calculated during the downsampling of the color buffer 320-1. In one embodiment, the current luminance 606 contains a texture at 1 / 32 render resolution such that each pixel contains the average luminance of a 32x32 pixel area in the source texture. The texture is sampled using the coordinates of the current pixel, and then bilinear sampling is used to obtain the average luminance of a 32x32 pixel region centered on the input pixel.

[0048] 6, in one or more embodiments, the various stages of the upscaler 332 operate in a color space, such as YCoCg, that is different from the rendered frame 502 (e.g., RGBa). Thus, to avoid repeatedly computing conversions from the color space used by the application 206, the upscaler 332 implements an input color adjustment stage 603 to apply all adjustments to the colors in one go. For example, in the input color adjustment stage 603, the input color adjustment component 514 takes as inputs the input color buffer 320-1 provided by the application 206 and an exposure 604 provided by the application 206 or generated by the auto exposure component 512 of the auto exposure stage 601. The input color adjustment component 514 processes these inputs to perform various adjustment operations on the input colors of the rendered frame 502. One example of such an adjustment operation includes dividing the input colors (e.g., RGB) by pre-exposure values ​​provided by the application 206 as part of the current frame 502 to obtain the original input colors. In many cases, different frames use different pre-exposure values ​​(to aid in calculation accuracy from a gaming perspective), and the original input colors of the frames are further divided by the pre-exposure values ​​to bring all images into a comparable range space. Another example of an adjustment operation includes aligning the input image to a mid-gray level by multiplying the input colors by the exposure value 604. A further example of an adjustment operation includes converting the exposed colors to the YCoCg color space. The YCoCg color space includes a luma value (Y) and two chrominance values, chrominance green (Cg) and chrominance orange (Co). The luma value (Y) represents the luminance in the rendered frame 502 (the achromatic portion of the frame), and the chroma components (Co and Cg) represent the color information of the rendered frame 502. The results of the input color adjustment stage 603 are then cached in a adjusted color buffer (texture) 608 that can be read by subsequent stages of the upscaler 332. In one or more embodiments, the resolution / size of the adjusted color buffer 608 is equal to the render resolution.

[0049] As part of the adjustment process, the input color adjustment component 514 generates a luma history buffer (texture) 610 with a resolution / size equal to the render resolution. The luma history buffer 610 contains the luma values ​​Y of the currently prepared input color of the rendered frame 502. In one or more embodiments, the luma history buffer 610 is persistent (i.e., not available for aliasing or cleared every frame). Thus, multiple frames (e.g., 4 frames) of luma history are kept in the luma history buffer 610 and are accessible during the input color adjustment stage 603 for any one frame. However, at the end of the input color adjustment stage 603, the luma history values ​​are shifted down. That is, in at least some embodiments, the next stage of the upscaler 332 has access to a smaller number of luma frames, e.g., the three most recent frames of luma (the current frame and the two previous frames). Thus, in this example, for a current frame of n, the values ​​stored in the luma history buffer 610 are as follows:

[0050] [Table 1]

[0051] In one or more embodiments, the input color adjustment component 514 encodes a stability factor into the alpha channel of the luma history buffer 610. The stability factor is a measure of the stability of luma over the current frame 502 and a predetermined number of frames (e.g., three frames) that precede the current frame 502.

[0052] In addition to performing the input color adjustment operations described above, the input color adjustment component 514 clears the reprojected / previous depth buffer 616 to a known value, as described below. This process prepares the previous depth buffer 616 for the reconstruction expansion stage 605 of the next rendered frame of the application 206. The input color adjustment component 514 selects the clear value based on, for example, the configuration of the previous depth buffer 616. For example, the input color adjustment component 514 clears the previous depth buffer 616 to a maximum z-far value, which is typically 0 for inverted depths. In at least some configurations, the previous depth buffer 616 is cleared as a result of the reconstruction expansion stage 605 using an atomic operation to populate the previous depth buffer 616.

[0053] In the reconstruction extension stage 605, the reconstruction extension component 516 receives as input the input depth buffer 322-1 and the motion vector buffer 504 provided by the application 206. The reconstruction extension component 516 processes this input to generate an extended depth buffer 612 for the previous frame, a buffer 614 containing an extended set of motion vectors in UV space, and a reprojected / previous depth buffer 616. In one or more embodiments, the reconstruction extension component 516 applies motion vector scaling to convert non-screen space motion vectors to screen space motion vectors before processing the motion vectors. The extended depth buffer 612 is a texture containing the extended depth values ​​determined from the input depth buffer 322-1. The extended motion vector buffer 614 is a texture containing the extended two-dimensional motion vectors determined from the input motion vector buffer 504. In one or more embodiments, a first color channel, such as a red channel, and a second color channel, such as a green channel, of the expanded motion vector buffer 614 contain two-dimensional motion vectors in normalized device coordinate (NDC) space. The previous depth buffer 616 is a texture that contains reconstructed depth values ​​of the previous frame. Each of these buffers has a resolution / size equal to the render resolution of the current frame 502.

[0054] More specifically, the reconstruction and extension component 516 calculates the extended depth values ​​and motion vectors from the input depth values ​​and motion vectors contained in the input depth buffer 322-1 and the input motion vector buffer 504 of the current frame 502, respectively. The extended depth values ​​and motion vectors accentuate edges of the geometry rendered in the input depth buffer 322-1. For example, edges of geometry often introduce discontinuities in a continuous series of depth values. Thus, when the depth values ​​and motion vectors are extended, they naturally follow the contours of the geometric edges present in the input depth buffer 322-1. In one or more embodiments, the reconstruction and extension component 516 calculates the extended depth values ​​and motion vectors by considering the depth values ​​of a 3×3 (or other kernel size) neighborhood of each pixel of the current frame 502. The reconstruction and extension component 516 then selects the depth value and corresponding motion vector of the neighborhood whose depth value is closest to the camera. This process is illustrated in Figure 8, which shows one embodiment of a geometry 802 of a current frame 502 in which a center pixel 804 of a 3x3 kernel 806 is updated with the depth value and motion vector from a pixel 808 with the maximum depth value. The reconstruction dilation component 516 stores the determined dilated depth value in a dilated depth buffer 612 and stores the determined dilated motion vector in a dilated motion vector buffer 614.

[0055] The reconstruction augmentation component 516 estimates the location of each pixel from the current frame's depth buffer 322-1 in the previous frame using the augmented motion vectors in the augmented motion vector buffer 614. For example, the reconstruction augmentation component 516 applies the augmented motion vector calculated for the pixel to the value in the augmented depth buffer 612 to determine the location of the pixel in the previous frame. In other words, each depth sample from the augmented depth buffer 612 is reprojected to its location in the previous frame using the augmented motion vector of the sample. The reprojected depth samples are distributed among the affected depth samples, for example, using backward or inverse reprojection. Since many pixels may be reprojected to the same pixel in the previous frame, the reconstruction augmentation component 516 uses an atomic operation to resolve the value of the closest depth value for each pixel. For example, in one or more embodiments, the reconstruction augmentation component 516 uses an atomic operation such as InterlockedMax or InterlockedMin provided by a high-level shader language (HLSL) or equivalent. In some embodiments, the reconstruction extension component 516 performs a different atomic operation (e.g., InterlockedMax or InterlockedMin) depending on whether the input depth buffer 322-1 is inverted or non-inverted. The reconstruction extension component 516 stores the reconstructed / estimated depth values ​​in the reconstructed / previous depth buffer 616.

[0056] 6, in the depth clip stage 607, the depth clip component 518 receives as input the expanded depth buffer 612, the expanded motion vector buffer 614, and the previous depth buffer 616. The depth clip component 518 processes this input to generate a disocclusion mask / map 618 that indicates the disoccluded areas of the current rendered frame 502. For example, as the camera moves from an initial position (previous frame) to a new position (current frame), pixels that were initially occluded from the perspective of the camera's previous position may become visible (disoccluded) from the perspective of the camera's current position. In one or more embodiments, the disocclusion mask 618 is a texture that contains values ​​that indicate how much of a corresponding pixel in the current frame 502 has been disoccluded. In one implementation, a value of 0 indicates that the entire pixel was occluded in the previous frame and is now disoccluded, and a value of 1 indicates that the pixel was fully visible in the previous frame and is also fully visible in the current frame 502. A value between 0 and 1 indicates that the pixel was visible in the previous frame to an extent proportional to the value. In one or more embodiments, the disocclusion mask 618 has the same resolution / size as the render resolution.

[0057] The depth clip component 518 generates the disocclusion mask 618, for example, by calculating per-pixel depth values ​​from the previous and new camera positions. The depth clip component 518 compares the delta between the depth values ​​to a separation value, such as the Akerley separation constant ksep., which provides the minimum distance between two objects represented in the floating-point depth buffer, and which the depth clip component 518 uses to determine with a high degree of certainty that the pixels were originally distinct from one another. In one or more embodiments, if the depth clip component 518 determines that the delta between the depth values ​​is greater than the separation value calculated for the application's configuration of the depth buffer 612, the depth clip component 518 determines that the pixels represent distinct objects. However, if the depth clip component 518 determines that the delta between the depth values ​​does not exceed the separation value, the depth clip component 518 cannot determine with certainty that the pixels represent distinct objects. The depth clip component 518 stores values ​​in the pixel disocclusion mask 618, for example, in the range [0...1], where a value of 1 maps to a delta greater than or equal to the separation value.

[0058] In the locking stage 609 of the upscaler 332, the locking component 520 generates new pixel locks based on the pixels of the current frame 502. The pixel locks are consumed in the reprojection accumulation stage 611 of the upscaler 332, as described below. The locking component 520 receives as inputs input lock state information 620, such as an input lock state texture / buffer 620-1 (also referred to herein as "lock state 620-1" for brevity) and an adjusted color buffer 608. Based on this input, the locking component 520 generates locks for one or more pixels of the current frame 502, if applicable. In one or more embodiments, the generated locks are stored in an output lock state texture / buffer 620-2 (also referred to herein as "lock state 620-2" for brevity). In one or more embodiments, the upscaler 332 maintains the lock state 620 as an array of two textures: an input lock state 620-1 and an output lock state 620-2. The input locked state 620-1 is a locked texture generated for a previously upscaled frame that is used as an input by the upscaler 332 when processing the current frame 502. The output locked state 620-2 is a locked texture generated for the current frame 502 that is used as an input for the next frame to be processed by the upscaler 332. The input locked state 620-1 and the output locked state 620-2 each have a resolution / size equal to the target (upscaled) presentation resolution and maintain the locked state 620. The two texture array configurations of the locked state 620 allow neighboring pixels of the pixel processed by the locking stage 609 to be read while avoiding read-modify-write conflicts. It should be understood that the locking stage 609 is not limited to image upscaling. For example, the locking stage 609 can be applied to temporal anti-aliasing (TAA) and other techniques.

[0059] In at least some embodiments, the input lock state 620-1 and the output lock state 620-2 each include lock state information for each pixel of the associated frame that indicates whether the pixel is associated with a lock or is unlocked. In one or more embodiments, the lock acts as a classifier that indicates that the locked pixel contains high frequency information (e.g., luma change between opposing neighborhoods of the pixel). In at least some embodiments, the lock acts as a mask that indicates that color correction in the reprojection accumulation stage 611 will not be performed on the pixel associated with the lock. The net effect of a locked pixel is that more color data from previous frames is used when calculating the final super-resolved pixel color during the reprojection accumulation stage 611. For example, if shading changes between frames, the next color correction process may limit the influence of the history color so that it is closer to the color of the current frame 502 sample surroundings, which causes small / faint features that are not visible in all jittered renders to be clamped and disappear from the blended image. Locking increases the contribution of samples containing small / thin features, preventing these features from being clamped and disappearing from the blended image.

[0060] In one or more embodiments, the lock state 620 consists of two values: a red channel value and a green channel value. In another embodiment, the lock state 620 consists of three values: a red channel, a green channel, and a blue channel. The red channel of the lock state 620 contains the remaining lifetime of the pixel lock and is initialized based on the jitter sequence length. For example, for 2x upscaling, the sequence length is 32 and the lock may be active for 32 frames. The length of 32 frames is stored in the red channel as the initial lifetime of the lock and decreases by approximately 1 / 32 times (in the example of 2x upscaling) each frame. In one or more other embodiments, once a pixel is locked, the locking component 520 populates the red channel of the output lock state 620-2 with a remaining lifetime value that indicates that the lock is new (i.e., created for the current frame) and whether the lock is the first / initial lock of the pixel or a subsequent lock (e.g., second lock). For example, the locking component 520 puts a remaining lifetime value of -1 in the red channel of the output lock state 620-2 to indicate that the lock is the first lock of the pixel, where the negative sign indicates that the lock is a new lock. If the pixel was locked in a previous frame, the new lock is considered the next pixel lock (e.g., the second lock), and the locking component 520 puts a remaining lifetime value of -2 in the red channel of the output lock state 620-2 to indicate that the lock is new (i.e., generated for the current frame) and is a subsequent lock of the pixel. In one or more embodiments, if the initial value of the lock's remaining lifetime is -1, the remaining lifetime decrements between 0 and 1, and if the initial value of the lock's remaining lifetime is -2, the remaining lifetime decrements between 0 and 2. As described in more detail below, in these embodiments, rather than decrementing the remaining lifetime by a factor of 1 / 32 (in the 2x upscaling example), the remaining lifetime is dynamically decremented depending on the contribution of the locked pixel to the final color of the current frame. Additionally, when the final lock of a pixel is stored in the reprojection accumulation stage 611, the absolute value of the remaining lifetime is stored.Thus, in at least some embodiments, the red channel of the final output locked state 620-2 includes a non-negative value for the remaining lifetime, and thus, when the final output locked state 620-2 becomes the next frame input locked state 620-1, the remaining lifetime stored in the red channel is also a non-negative value.

[0061] The green channel of the locked state 620 contains the current luminance of the current frame 502 at the time the pixel is locked (e.g., the average luma of the scene in the area around the pixel) and is obtained by the locked state 620 from the luma texture 606. In at least some embodiments, the locking component 520 populates the green channel of the locked state 620 during the reprojection stage of the reprojection accumulation stage 611. The luminance information stored in the currently generated green channel of the locked state 620 is used by the reprojection accumulation stage 611 when processing the next frame. The luminance value stored in the green channel of the locked state 620 is used by the reprojection accumulation stage 611 as part of shading change detection for the current frame 502, which allows the upscaler 332 to unlock the pixel if there is a discontinuous change in the pixel's appearance (e.g., an abrupt change to the shading of the pixel). In at least some embodiments, the blue channel of the locked state 620 is populated with a confidence / trust factor that indicates how reliable the associated lock is. It should be appreciated that each of the remaining lock life information, current brightness information and confidence / reliability factors may be stored interchangeably in any of the red, green and blue channels of the lock state 620 .

[0062] In one or more embodiments, the locking component 520 determines whether to generate a lock for a pixel of the current frame 502 by comparing a neighborhood (e.g., 3×3 pixels) of the luminance value associated with the pixel against a luminance difference threshold. In other words, the locking component 520 determines whether the relative luminance difference between the pixel and the neighborhood is less than the luminance difference threshold. The luminance difference can be characterized as max(A,B) / min(A,B), where A and B are luminance / luma values. One example of the luminance difference threshold is a predefined percentage of similarity to the median luminance of the pixel's neighborhood (e.g., 3×3). The locking component 520 obtains the luminance value from the adjusted color buffer 608. Through the use of the neighborhood, the locking component 520 can detect thin features (e.g., wires or chain-linked fences) in the current frame 502 that should be locked to preserve details in the final super-resolution frame / image. If the shading change difference (e.g., luminance difference) between the pixel and its luminance value neighbors meets the luminance change threshold, the locking component 520 locks the pixel by generating an entry in the output lock state 620-2. For example, if the luminance change is less than (or equal to) the luminance change threshold, the locking component 520 locks the pixel. Otherwise, the locking component 520 retains the pixel's current locked state (e.g., unlocked state), which in at least some embodiments is also reflected in the output lock state 620-2. More specifically, the locking component 520 compares the luminance value of the center pixel of a 3×3 region with the luminance values ​​of adjacent pixels in the 3×3 (or other) neighborhood. If the relative luminance difference between the center pixel and the neighboring pixels is less than the luminance difference threshold, the locking component 520 sets a bit in a mask. The locking component 520 considers the center pixel as a candidate for locking if its luminance is outside the luminance range of all neighboring pixels. After the locking component 520 compares the luminance of the central pixel with all of the neighboring pixels, the locking component 520 determines whether the final mask contains noise, i.e., whether it does not have 2x2 (or other) regions where all bits are set.If the mask is noisy and the intensity of the central pixel is outside the intensity range of all neighboring pixels, the locking component 520 generates a lock for the central pixel and stores the lock in the output lock state 620-2.

[0063] Besides generating a new lock, the locking component 520 (or another component of the upscaler 332) also updates the lock in the input lock state 620-1 generated for the previously upscaled frame. For example, the locking component 520 decreases the red channel value of the pixel lock in the input lock state 620-1 generated for the previously upscaled frame by the initial pixel lock length divided by the total length of the (color buffer / camera) jitter sequence. For example, if the initial length of the lock is initially set to 1, subtracting 1 / jitter sequence length from the remaining lifetime of the lock will result in a remaining lifetime of the lock of 0 after one iteration of the jitter sequence. In another embodiment, the locking component 520 (or another component of the upscaler 332) dynamically decreases / decrements the red channel value of the pixel lock in the input lock state 620-1 depending on the contribution of the locked pixel to the final color of the current frame. When the lock reaches zero (or another threshold), the locking component 520 considers the lock expired and releases the lock. In one or more embodiments, the initial length of the lock is modified by a confidence / certainty factor that indicates how confident the locking component 520 is that the associated pixel should be locked, which affects the length for which the pixel is locked. The locking component 520 updates one or more locks in the input lock state 620-1 by releasing the lock(s), as described below.

[0064] In the reprojection accumulation stage 611, the reprojection accumulation component 522 receives as inputs the disocclusion mask 618, the dilated motion vector buffer 614, the reactive mask 602, the output buffer 506-1 of the previous frame, the current luma texture 606, the luma history 610, the adjusted color buffer 608, and the lock state 620. The reprojection accumulation component 522 processes this input to generate an output buffer (texture) 506-2 of the current frame 502 at the target presentation resolution / size, and also generates a reprojected pixel lock (texture) 622 from the previous frame that can be mapped onto the current frame 502. As described below, the reprojection accumulation component 522 accumulates the reprojected color data from the previous frame with the upsampled color data from the current frame 502 and stores the accumulated color data in the output buffer 506-2. In one or more embodiments, output buffer 506-2 is a texture used internally by upscaler 332 and is distinct from presentation buffer 626 generated by sharpening stage 613. Additionally, in one or more embodiments, output buffer 506-2 is part of output buffer 506, which is represented as an array of two additional textures (e.g., output buffer 506-1 and output buffer 506-2) consumed by reprojection accumulation stage 611. For odd frames, output buffer 506-1 is read and output buffer 506-2 is written (or vice versa). For even frames, output buffer 506-2 is read and output buffer 506-1 is written (or vice versa).

[0065] As seen in FIG. 9, the reprojection accumulation stage 611 is comprised of multiple substages. In one or more embodiments, the substages include a shading change detection substage 901, an upsampling substage 903, a reprojection substage 905, a lock update substage 907, a color correction substage 909, one or more tone mapping substages 911 (illustrated as tone mapping substage 911-1 and tone mapping substage 911-2), an accumulation substage 913, and an inverse tone mapping substage 915. In at least some embodiments, the reprojection accumulation component 522 includes one or more subcomponents 902-914 that perform operations associated with one or more of the reprojection accumulation substages 901-913. In one or more embodiments, the reprojection accumulation stage 611 is performed on a frame-by-frame basis, and the data flow illustrated in FIG. 9 is performed once per output / presentation resolution pixel for each pixel in parallel, and further, the data flow illustrated in FIG. 9 is performed in parallel for each pixel of the upscaled frame 508. 9, a "2x2 Bilinear" reference indicates that bilinear sampling is performed at the corresponding stage, a "5x5 Lanczos" reference indicates that Lanczos resampling is performed at the corresponding stage, and a "1" reference indicates that point sampling is performed, but it should be understood that other types of sampling are equally applicable.

[0066] As described below, the reprojection accumulation component 522, via one or more of the subcomponents 902-914, performs a number of operations, including: For the current frame 502, performing an upsampling of the color buffer 320-1 and reprojecting the (historical) color data 922 and the pixel lock provided by the output buffer 506-1 of the previous frame as if seen from the current camera's point of view; cleaning the reprojected color data 922; and accumulating the final historical color data / values ​​928 (FIG. 9) and the upsampled color data 918 for the current frame 502; Optionally, inverse tone mapping and sharpening the super-resolved color values ​​930 produced by the reprojection-accumulation stage 611.

[0067] At the start of the reprojection accumulation stage 611, the shading change detection sub-stage 901 receives the luminance history 610, the current luminance texture 606, and the input lock state 620-1. The shading change detection component 902 processes the inputs to evaluate each pixel of the current frame 502 to detect changes in the shading of the pixel. In one or more embodiments, the shading change detection component 902 also performs texture filtering, such as bilinear filtering, on these inputs. In one implementation, the shading change detection component 902 uses the lock state 620-1 of the current pixel in the previously upscaled frame to determine whether the pixel is locked or unlocked. If the pixel is locked, the shading change detection component 902 determines whether the shading of the pixel has changed by comparing the luminance (Y) value of the pixel at the time the lock was generated to a shading change threshold. In one or more embodiments, the shading change detection component 902 compares the luminance of the pixel to the average luminance of the pixel's neighborhood (e.g., 32×32) in the current frame 502. If the difference between the luminance of the pixel and the average luminance of the neighboring pixels meets (e.g., is greater than or equal to) a shading change threshold, the shading change detection component 902 determines that the shading of the pixel has changed. In one or more embodiments, the shading change detection component 902 obtains the luminance of the pixel from the green channel of the pixel's input lock state 620-1. If the pixel is unlocked, in one or more embodiments, the shading change detection component 902 determines whether the shading of the pixel has changed based on the luminance value of the pixel in the current frame and the historical luminance value(s) of the pixel or a neighbor of the pixel in the past. For example, if the difference between the luminance value of the pixel in the current frame and the historical luminance value(s) of the pixel (or a neighbor of the pixel in the past) meets a shading change threshold, the shading change detection component 902 determines that the shading of the pixel has changed.In one or more embodiments, the shading change detection component 902 obtains the current luminance value of the pixel from the current luminance texture 606 and the historical luminance values ​​of the pixel from the luminance history 610. The shading change detection component 902 generates shading change data 924 that includes, for example, a bit / flag indicating whether the shading of the pixel has changed. The shading change data 924 is received as input by the upsampling sub-stage 903, the lock update sub-stage 907, and the color correction sub-stage 909.

[0068] The upsampling component 904 of the upsampling sub-stage 903 receives as input the shading change data 924 from the shading change detection component 902 and the adjusted color buffer 608 and upsamples the adjusted color buffer 608. In one or more embodiments, the upsampling component 904 uses the shading change data to reshape the filter kernel, which may result in lower or higher sample weights being used in the accumulation sub-stage 913. Upsampling the adjusted color buffer 608 involves interpolating between existing pixels in the adjusted color buffer 608 to obtain estimates of their values ​​at new pixel locations. For example, if the current frame 502 has been upsampled from 1920×1080 pixels to 3840×2160 pixels, the upsampling component 904 interpolates between the original 1920×1080 pixels to estimate color values ​​for the new pixels upsampled to the higher 3840×2160 pixel resolution. In at least some embodiments, the upsampling component 904 implements Lanczos resampling to upscale the pixels of the adjusted color buffer 608, although one or more different scaling techniques or algorithms, such as Sinc resampling, nearest neighbor interpolation, bilinear algorithms, bicubic algorithms, box sampling, etc., or combinations thereof, may be applied as well.

[0069] In general, Lanczos resampling is an interpolation method used to calculate new values ​​for digitally sampled data. When used in digital image resizing, the Lanczos function indicates which pixels in the rendered (original) image make up each pixel and which part of the upsampled image it comprises. For example, FIG. 10 shows a number of pixels 1002 rendered from a current frame 502. In each rendered pixel 1002, point P represents a set of low-resolution samples 1004 that are resampled and for which Lanczos weights are calculated. Each point S in the pixel 1002 represents a proposed resolution target pixel 1006 (also referred to herein as an upsampled pixel 1006) for calculation based on resampling from the set of low-resolution samples 1004. The location of the upsampled pixel 1002-1 serves as the center of a Lanczos resampling kernel 1008, such as the 5×5 Lanczos resampling kernel represented in FIG. 10 by the dashed grid. The Lanczos function 1010 (illustrated as Lanczos function 1010-1 and Lanczos function 1010-2) is centered on the upsampled pixel 1006-1 being calculated. Thus, the color value of the upsampled pixel 1006-1 is determined by applying a Lanczos resampling kernel 1008 (e.g., Lanczos(x,2)) to the low-resolution samples 1004-1 in a 5×5 neighborhood (grid) of pixels surrounding the upsampled pixel 1006-1. In other words, the weight of each low-resolution sample 1004 in the 5×5 pixel neighborhood is determined using the Lanczos resampling and the distance from the low-resolution sample 1004 to the presentation resolution target pixel 1006-1. The final color value of the upsampled pixel 1006-1 is based on the sum of the weights of each low-resolution sample 1004 in the 5×5 pixel neighborhood. In some cases, the Lanczos resampling may result in ringing artifacts.Thus, in one or more embodiments, the final color value of the upsampled pixel 1006-1 is clamped, for example, using a central 2x2 kernel range, to reduce ringing artifacts.

[0070] While the above examples implement a 5×5 pixel neighborhood during the Lanczos resampling process, in some cases a 4×4 neighborhood is sampled due to the zero weighted contribution of pixels on the periphery of the 5×5 neighborhood. Furthermore, in one or more aspects, the implementation of the Lanczos kernel varies depending on the GPU being implemented. For example, in one or more GPU embodiments, a look-up table (LUT) can be used to encode the sinc(x) function of the Lanczos kernel. The LUT can be used to balance arithmetic logic (ALU) operations in the reprojection accumulation stage 611 with memory usage. For example, for a given jitter, the sample Lanczos values ​​used are the same for all pixels of a frame. Thus, in one or more embodiments, these Lanczos values ​​are pre-computed and recorded in a LUT. The LUT is passed to a shader that implements the upsampling component 904 as a texture. Then, rather than repeatedly calculating the Lanczos values, the pre-computed Lanczos values, which may be resident in a cache, are repeatedly read in the shader by the upsampling component 904. In one or more embodiments, to reduce the ALU usage of the shader core, the LUT is read through a fixed function sampling block. If the shader stage is not bandwidth limited, implementing the LUT is faster than using ALU cycles.

[0071] In at least some embodiments, the upsampling component 904 also calculates a YCoCg bounding box 1102 (FIG. 11) for each pixel to be upsampled. For example, FIG. 11 shows a YCoCg bounding box 1102 constructed from a 3×3 neighborhood of pixels 1104 centered on the current pixel (e.g., pixel 1002-1 in FIG. 10). It should be understood that the bounding box 1102 also has a third dimension for Cg (not shown for illustrative purposes). The bounding box 1102 is used during the color correction sub-stage 909 described below. In one or more embodiments, the upsampling component 904 generates the bounding box 1102 for each pixel in the current frame 502 by calculating the minimum and maximum of each channel of all pixels in a neighborhood (e.g., 3×3) surrounding the current pixel. The different shading / patterns of the squares 1106 in the 3×3 pixel neighborhood 1104 represent the Y value of the pixel. The location of each square within the bounding box 1102 represents where in the bounding box 1102 the sample is located. In at least some embodiments, the output of the upsampling sub-stage 903 includes upsampled (pixel) color data / values ​​918 for the upsampled pixels (i.e., pixels upsampled to the presentation resolution) generated for the current frame 502, and includes a YCoCg bounding box 1102 generated for each upsampled pixel.

[0072] The reprojection component 906 of the reprojection sub-stage 905 receives as input the expanded motion vector buffer 614, the input lock state 620-1, and the output buffer 506-1 of a previously upscaled frame (i.e., a frame previously processed by the upscaler 332). The reprojection component 906 processes this input and reprojects the output color buffer from the output buffer 506-1 and the pixel lock information from the input lock state 620-1 of the previous frame onto the new locked pixels generated for the current frame 502 as if seen from the camera's point of view of the current frame 502. For example, FIG. 12 shows frame 1202, which represents a previously upscaled frame. The reprojection component 906 samples the expanded motion vectors in the expanded motion vector buffer 614. The reprojection component 906 then applies the sampled motion vector 1204 (illustrated as motion vector 1204-1 and motion vector 1204-2) to the output buffer 506-1 of the previously upscaled frame (e.g., color data). FIG. 12 shows the sampled two-dimensional motion vector 1204 being applied to the location of the current pixel 1206 from the output buffer 506-1 of the previously upscaled frame. That is, the reprojection component 906 subtracts the sampled motion vector 1204 from the location coordinate of the current pixel to obtain the historical coordinate. The reprojection component 906 then uses the historical coordinate (the pixel location after transformation) to sample the input lock state 620-1 and the historical color information (from the output buffer 506-1).

[0073] In one or more embodiments, the reprojection component 906 performs Lanczos resampling to sample a neighborhood of pixels surrounding the transformed pixel location. For example, FIG. 12 shows a Lanczos(x,2) resampling kernel 1208 applied to a 5×5 grid 1210 of pixels surrounding the transformed pixel location 1206-1. In one or more embodiments, the reprojection component 906 reprojects the lock using bilinear interpolation at the transformed pixel location. Thus, the output of the reprojection sub-stage 905 is a presentation resolution image that includes all data of the previous frame that can be mapped to the current frame 502. In other words, the output of the reprojection sub-stage 905 is the reprojected color data 922 and the reprojected pixel lock 622 of the pixel in the previously upscaled frame that can be mapped to the current frame 502. For example, the remaining lifetime from the red channel, the current luminance of the current frame 502 from the green channel, and the confidence / reliability factor from the blue channel (if implemented) in the input lock state 620-1 of a pixel for which no new lock has been generated for the current frame 502 are reprojected / written as the reprojection lock 622. In one or more embodiments, the reprojection lock 622 refers to the reprojected lock information written to the pixel's output lock state 620-2 or the reprojected lock information stored in the intermediate lock texture. In at least some embodiments, the reprojection component 906 determines whether a new lock has been generated for a pixel based on the pixel's output lock state 620-2 that includes a negative remaining lifetime value. In one or more embodiments, if the pixel's input lock state 620-1 indicates that the pixel has only been locked once (e.g., the remaining lifetime value is "1"), the lock is not active and the reprojection component 906 does not perform a reprojection operation on the pixel.In at least some embodiments, the reprojected pixel lock 622 is reprojected to the output lock state 620-2 and used as a lock for pixels in the current frame 502 that have not been updated with new sample information, while the pixel locks generated during the locking stage 609 are used for pixels that have updated samples in their area. It should be understood that the locked reprojection process of the reprojection sub-stage 905 is not limited to image upscaling. For example, the locked reprojection process can be applied to TAA and other techniques.

[0074] The lock update component 908 of the lock update sub-stage 907 receives as input the de-occlusion mask 618, the reprojected lock 622, and the shading change data 924. The lock update component 908 processes this input to generate an updated set of pixel locks 926, where updating the pixel locks includes identifying a set of pixel locks from the reprojected locks 622 to be passed to the color correction sub-stage 909. For example, if the lock update component 908 determines that the de-occlusion mask 618 indicates that the pixel associated with the reprojected pixel lock is occluded, the lock update component 908 releases the lock. As another example, if the lock update component 908 determines that the shading change data 924 indicates that the shading of the pixel associated with the reprojected pixel lock has changed by more than a shading change threshold, the lock update component 908 releases the lock. In one or more embodiments, the lock update component 908 updates the reprojected lock to reflect the decayed and released lock.

[0075] The lock update component 908 then determines which of the remaining reprojected locks are reliable for the current frame 502. In one or more embodiments, the lock update component 908 determines the reliability of the pixel lock by comparing the luminance value associated with the pixel lock (i.e., the luminance value of the pixel at the time the pixel lock was generated) to a luminance value of the current luminance texture 606 that represents the average luminance of the scene in a neighborhood of pixels centered around the pixel associated with the current pixel lock. If the luminance separation between the compared luminance values ​​exceeds (or is equal to) the luminance separation threshold, the lock update component 908 determines that the pixel lock is unreliable and releases the lock. However, if the luminance separation between the compared luminance values ​​is below (or equal to) the luminance separation threshold, the lock update component 908 determines that the pixel lock is reliable and maintains the pixel lock.

[0076] In one or more embodiments, the lock update component 908 processes the input to identify new pixel locks generated during the locking stage 609 of the current non-reprojected frame 502. As described above, the lock update component 908 identifies a new pixel lock based on, in at least some embodiments, the red channel of the output lock state 620-2 having a negative value or the remaining lock lifetime being equal to the initial lock lifetime. If a new pixel lock is identified, the lock update component 908 does not perform the unlock operation described above for the new pixel lock. However, in one or more embodiments, the lock update component 908 updates the green channel of the output lock state 620-2 with current luminance information from the current luminance texture 606 for new pixel locks that are active (e.g., pixel locks that correspond to at least the second pixel lock for a pixel). In one or more embodiments, the lock update component 908 updates the confidence / reliability factors of the pixel locks stored in the blue channel of the lock state 620 of the new active pixel lock and the reprojected pixel locks (if implemented). For example, in at least some embodiments, the lock update component 908 increases the confidence / confidence factors of the new pixel lock and the reprojected lock based on the current upsample weight contributions of the associated pixels. The increase value is in the normalized range [0,1] and is determined by dividing the upsampling weights of the current frame (e.g., the Lanczos weights calculated in the upsampling substage 903) by a constant value, e.g., the maximum cumulative weight. This increase value is added to the current confidence / confidence factors stored in the blue channel of the lock state 620. In at least some embodiments, the confidence / confidence factors are clamped to the range [0,1].

[0077] In one or more embodiments, the lock update component 908 decreases the lock lifetime of the new pixel lock and the reprojected lock. For example, the lifetime of the reprojected or new pixel lock is decreased by approximately 1 / 32 times (in the example of 2x upscaling). As another example, the lifetime of the reprojected or new pixel lock is dynamically decremented depending on the contribution of the locked pixel to the final color of the current frame, i.e., the current upsample weight contribution. In at least some embodiments, the block update component 908 calculates the decrease value by dividing the upsample weight of the current frame by a factor based on the jitter sequence length and the average upsample kernel weight. In one or more embodiments, this decrease value is in the normalized range [0,1]. The lock update component 908 then updates the remaining lifetime of the pixel lock stored in the red channel of the locked state 620 by subtracting the decrease value from the current remaining lifetime of the pixel lock. In one or more embodiments, the resulting value is clamped to the range [0,1] to avoid negative lifetime values.

[0078] The output of the lock update sub-stage 907 is a set of updated locks 926 (i.e., lock states 620) including the reprojected pixel locks 622 determined to be reliable, or includes a set of reprojected locks 622 with an indicator associated with each lock that identifies whether the lock is reliable or unreliable. In one or more embodiments, the updated set of pixel locks 926 includes the newly generated active pixel locks for the current frame 502. In at least some embodiments, each reprojected pixel lock in the updated set of locks 926 includes an updated remaining lifetime, intensity value, and confidence / confidence factor, and each newly generated active lock includes an updated intensity value and confidence / confidence factor, as described above. The updated output lock states 620-2 for the new pixel locks are then stored in the output buffer 506-2 for use during upsampling of the next frame. It should be understood that the lock update sub-stage 907 is not limited to image upscaling. For example, the lock update sub-stage 907 can be applied to TAA and other techniques.

[0079] The set of updated locks 926 is then passed as input to the color correction sub-stage 909. The color corrector component 910 of the color correction sub-stage 909 receives as inputs, for example, upsampled color data 918 for upsampled pixels of the current frame 502, color data 922 for reprojected pixels of a previously upscaled frame, the set of updated locks 926, the deocclusion mask 618, luma stability factors (e.g., the w channel of the luma history 610), shading change data 924, and a lock contribution factor for the newly locked or reprojected locked pixel. The lock contribution factor is calculated as a function of the upsampled color data 918, any transparency and composition masks, and any reactive mask 602 weights. The color corrector component 910 processes the inputs to determine final historical color data / values ​​928 for each upsampled pixel of the current frame 502 and its contribution to the final upscaled pixel. In one or more embodiments, when determining the final history color value 928 for each upsampled pixel, the color corrector component 910 reduces the influence of the history samples based on various factors including, among others, lock state or disocclusion. In other words, the color corrector component 910 modifies the history color values ​​in some cases to match the shading of the current frame 502. For example, if a pixel location is (partially) disoccluded, it means that the pixel location was occluded in the previously upscaled frame. Thus, the history color data for the pixel location is no longer valid to improve the quality of the current pixel. Therefore, the color corrector component 910 reduces the influence of the history samples of the disoccluded pixels to avoid artifacts such as ghosting. In this example, the color corrector component 910 reduces the influence of the history samples by modulating the color values ​​of the history samples by the disocclusion mask 618.

[0080] FIG. 13 illustrates how the correction component 910 adjusts the distance (d) of the upsampled pixel from the color bounding box 1102 by the history sample (S h 9 illustrates another example of reducing the influence of the historical sample 1302 relative to the upsampled color data 918. In this example, the greater the distance (d), the more the influence of the historical sample 1302 is reduced by the color corrector component 910. In one or more embodiments, if the color corrector component 910 determines that the historical sample is locked based on the updated set of locks 926 and further determines that the reactive mask 602 or the deocclusion mask 618 does not override the lock, the color corrector component 910 increases the influence of the historical color data of the historical sample relative to the upsampled color data 918.

[0081] In one or more embodiments, at least one of the upsampled color data 918 or the final historical color data 928 is passed to a tone mapping sub-stage 911 (illustrated as tone mapping sub-stage 911-1 and tone mapping sub-stage 911-2). For example, if the upsampled color data 918 or the final historical color data 928 includes high dynamic range data, the upsampled color data 918 or the final historical color data 928 is passed to the tone mapping sub-stage 911, respectively. A tone mapping component 912 (illustrated as tone mapping component 912-1 and tone mapping component 912-2) then applies one or more local tone mapping operations to the upsampled color data 918 or the final historical color data 928 to reduce artifacts, such as firefly artifacts. In one or more embodiments, tone mapping ensures that color values ​​are in the range 0 to 1 by dividing the color by the maximum channel value, e.g., Color.rgb=color.rgb / (1+max(max(color.r,color.g),color.b)), where max is the maximum channel value, color.r is the red channel value, color.g is the green channel value, and color.b is the blue channel value.

[0082] An accumulation component 914 of accumulation sub-stage 913 receives as input upsampled color data 918, which may or may not be tone mapped, and final historical color data 928 for each pixel of the current frame 502. For each upsampled pixel of the current frame 502, accumulation component 914 accumulates (e.g., blends) the upsampled color data 918 with the upsampled pixel's final historical color data 928 to collectively form final accumulated super-resolved color values ​​930 for the upsampled version of the current frame 502. An implicit result of the temporal accumulation of multiple frames in accumulation sub-stage 913 is that the resulting final accumulated super-resolved color values ​​930 are anti-aliased. In one or more embodiments, the accumulation component 914 blends the upsampled color data 918 of the current frame 502 with the final historical color data 928 with a relatively low linear interpolation factor such that the super-resolution color value 930 of the upscaled frame 508 contains a small amount of color data from the current frame 502. In other words, the final historical color data 928 determined for each upsampled pixel contributes more to the final accumulated super-resolution color value 930 than the upsampled color data 918 of the current frame 502. The blending of the two color values ​​may be characterized as follows: Blended Value=A * (1-W)+B *W, where A is the final historical color data 928, B is the color data from the current frame 502, and W is the contribution weight. In one or more embodiments, the contribution weight of the upsampled color data 918 can be increased based on the reactivity mask 602, which the accumulation component 914 also receives as an input. The reactivity mask 602 instructs the accumulation component 914 where it should rely less on historical information when synthesizing the current upsampled pixel, and instead make the upsampled pixel's color data 918 contribute more to the final accumulated super-resolution color value 930 (also referred to herein as "super-resolution color value 930" for brevity). For example, for a given pixel of the current frame 502, the accumulation component 914 multiplies the reactivity value from the reactivity mask 602 associated with the pixel with a contribution weight W to increase the contribution weight of the upsampled color data 918. Thus, the blending of two color values ​​taking into account the reactivity mask 602 can be characterized as follows: Blended Value=A * (1-(R*W)+B * (R * W) where A is the final historical color data 928, B is the color data from the current frame, and R is the reactivity value from the reactivity mask 602.

[0083] The super-resolved color values ​​930 represent an upscaled version of the current frame 502 at the presentation resolution and are stored in the output buffer 506-2. In other words, the output buffer 506-2 contains the upscaled frame 508. In one or more embodiments, the output buffer 506-2 is used as the output buffer 506-1 that the upscaler 332 takes as input when processing the next frame. In one or more embodiments, the super-resolved color values ​​930 are passed to an inverse tone mapping sub-stage 915 before being stored in the output buffer 506-2. In these embodiments, the inverse tone mapping component 916 inverse tone maps the super-resolved color values ​​930. The inverse tone mapping process reverses the color tone mapping process described above such that the color range of the super-resolved color values ​​930 matches the color range of the input frame 502. In one or more embodiments, the inverse tone mapping is characterized by Color.rgb=color.rgb / (1.f-max(max(color.r,color.g),color.b)), where f denotes a floating point number, max is the maximum channel value, color.r is the red channel value, color.g is the green channel value, and color.b is the blue channel value.

[0084] Referring to FIG. 6, the sharpening component 524 of the sharpening stage 613 receives as input the output buffer 506-2 containing the upscaled frame 508. The sharpening component 524 performs one or more sharpening operations on the pixels of the upscaled frame 508 represented by the output buffer 506-2. The sharpening component 524 stores the sharpened upscaled frame 508 in the presentation buffer 626. In one or more embodiments, the sharpening component 524 performs Robust Contrast Adaptive Sharpening (RCAS). However, other image sharpening techniques are equally applicable. RCAS generates additional clarity and sharpness in the final upscaled image, and resolves to maximize local sharpness as much as possible before clipping when translating local contrast into a variable amount of sharpness. RCAS also incorporates a process to limit the sharpening of potential noise in the image. More specifically, as shown in Figure 14, the RCAS operates on sampled data 1402 (Figure 14) from the output buffer 506-2 using a 5-tap filter 1404 (Figure 14) configured in a cross pattern. During the RCAS process, the sharpening component 524 calculates weights w based on the following equation:

[0085]

number

[0086] The sharpening component 524 selects either w0 or w1 based on the weight w that does not result in clipping, and further limits the weight w. The sharpening component 524 multiplies by a "sharpness" amount, which is a sharpness value provided by the application 206 that modifies the perceptual sharpness of the output image. The RCAS process performs high-pass filtering and shaping after normalization for local contrast. This process is used as a noise detection filter to reduce the effect of RCAS on grain and focus on actual edges. In addition, the RCAS process also supports pass-through alpha. Once the sharpening process is complete, the sharpening component 524 stores the data of the upscaled frame 508 in the presentation buffer 626. In at least some embodiments, the post-upscaling post-processing stage 311-2 performs one or more post-processing operations on the upscaled frame 508 stored in the presentation buffer 626. The post-processed upscaled frame 508 is then output by the graphics processing pipeline 214 for presentation to the display device 118. Alternatively, the upscaled frame 508 is output by the graphics processing pipeline 214 after sharpening for presentation on the display device 118 without performing any post-processing operations after upscaling.

[0087] FIG. 15 illustrates, in the form of a flow chart, an overview of one exemplary method 1500 for performing spatial upscaling of a rendered frame 502 of a video stream by using temporal feedback to reconstruct a high resolution image representative of the rendered frame 502. It should be understood that the process described below with respect to the method 1600 has been described in more detail above with reference to FIGS. 6 through 14. At block 1502, the upscaler 332 obtains a first frame of the video stream rendered at a first resolution. The first frame is defined by a first plurality of pixels associated with a first set of color data. At block 1504, the upscaler obtains a second frame of the video stream upscaled to a second resolution higher than the first resolution. The second frame is defined by a second plurality of pixels associated with a second set of color data. At block 1506, the upscaler locks one or more pixels of the first plurality of pixels. At block 1508, the upscaler 332 upsamples the first plurality of pixels to the second resolution. The upsampling generates upsampled color data for the upsampled first plurality of pixels based on the first set of color data. At block 1510, the upscaler 332 accumulates the upsampled color data with the second set of color data to generate final color data for the upsampled first plurality of pixels. The upsampled color data associated with the one or more locked pixels is maintained during the accumulation of upsampled color data for at least one of the rendered frame or the next frame. Furthermore, color data of the second set of color data associated with the pixel lock contributes more to the final color data than corresponding color data of the upsampled color data. At block 1512, the upscaler 332 stores the upsampled first plurality of pixels together with the final color data as an upscaled frame representing the first frame at the second resolution.

[0088] 16 and 17 outline, in flow chart form, another exemplary method 1600 for performing spatial upscaling of a rendered frame 502 of a video stream by using temporal feedback to reconstruct a high resolution image representative of the rendered frame 502. It should be understood that the process described below with respect to the method 1600 has been described in more detail above with reference to FIGS. 6 through 14. At block 1602, the rendering stage 313 of the graphics pipeline 214 renders the frame 502 at a first resolution. At block 1604, the rendered frame 502 is passed to the super-resolution upscaler 332 of the graphics pipeline 214. At block 1606, the auto-exposure stage 601 of the upscaler 332 processes the color buffer 320-1 of the rendered frame 502 to generate an exposure texture 604 that includes exposure value(s) determined for the current frame, and a current luminance texture 606 that includes current luminance value(s) for the rendered frame 502. In block 1608, the input color adjustment stage 603 of the upscaler 332 processes the exposure texture 604 and the color buffer 320-1 of the rendered frame 502 to generate an adjusted color buffer 608 and a luma history texture 610 that contains historical luma information / values ​​for the rendered frame 502.

[0089] At block 1610, the reconstruction enhancement stage 605 of the upscaler 332 processes the depth buffer 322-1 and the motion vector buffer 504 of the rendered frame 502 to generate a previous depth buffer 616, an enhanced depth buffer 612, and an enhanced motion vector buffer 614 for the rendered frame 502. At block 1612, the depth clip stage 607 of the upscaler 332 processes the previous depth buffer 616, the enhanced depth buffer 612, and the enhanced motion vector buffer 614 to generate a disocclusion mask 618 for the rendered frame 502. At block 1614, the pixel locking stage 609 of the upscaler 332 processes the adjusted color buffer 608 to lock one or more pixels of the rendered frame 502 and generate a locked texture 620.

[0090] At block 1616, the reprojection accumulation stage 611 of the upscaler 332 processes the disocclusion mask 618, the dilated motion vector buffer 614, the reactive mask 602 associated with the pixels of the rendered frame 502, the output buffer 506-1 for the previously upscaled frame, the current luma texture 606, the luma history texture 610, the adjusted color buffer 608, and the locked state texture 620 to generate a super-resolution (and anti-aliased) upscaled frame 508 representing the rendered frame. The upscaled frame 508 has a second resolution that is greater than the first resolution of the rendered frame 502. At block 1618, the sharpening stage 613 of the upscaler 332 sharpens the upscaled frame 508. At block 1620, the sharpened upscaled frame 508 is stored in a presentation buffer 626 for presentation to the display device 118. In one or more embodiments, further post-processing is performed on the upscaled frame 508 before presenting the upscaled frame 508 to the display device 118. Flow returns to block 1602, where the rendered frame is then processed by the upscaler 332.

[0091] 18 and 19 together illustrate in flow chart form a more detailed method 1800 of the reprojection accumulation process shown in block 1616 of FIG. 17. It should be understood that the process described below with respect to method 1800 has been described in more detail above with reference to FIGS. 9 through 14. At block 1802, shading change detection sub-stage 901 processes the locked state texture 620, the current intensity texture 606, and the intensity history texture 610 to generate shading change data 924 for each pixel of the rendered frame 502. At block 1804, upsampling sub-stage 903 processes the adjusted color buffer 608 to upsample pixels of the rendered frame 502 at the second resolution (i.e., the target presentation resolution). Upsampling sub-stage 903 generates a color bounding box 1102 for each upsampled pixel. In block 1806, the reprojection sub-stage 905 processes the expanded motion vector buffer 614, the locked texture 620, and the previously upscaled frame output buffer 506-1 to reproject pixel data (e.g., color data 922) and pixel locks 622 from the previously upscaled frame that can be mapped onto the rendered frame 502.

[0092] At block 1808, the lock update stage 907 processes the deocclusion mask 618, the reprojected pixel locks 622, and the shading change data 924 to generate an updated set of pixel locks 926, including the reliable locks. Locks determined to be unreliable are released. At block 1810, the color correction sub-stage 909 processes the upsampled pixel data (e.g., color data 918), the upsampled pixel color bounding box 1102, the updated set of pixel locks 926, the deocclusion mask 618, the reprojected pixel color data 922, and the shading change data 924 to determine final historical color data / values ​​928 for each upsampled pixel.

[0093] At block 1812, one or more tone mapping sub-stages 911 tone map the final historical color data 928 and the color data 918 of the upsampled pixel. However, in at least some embodiments, the tone mapping operation is optional. At block 1814, the accumulation sub-stage 913 processes the final historical color data 928 and the potentially tone mapped color data 918 to generate a final accumulated super-resolved color value 930 of the upsampled pixel. At block 1816, the inverse tone mapping sub-stage 915 inverse tone maps the super-resolved color value 930. However, in at least some embodiments, the inverse tone mapping operation is optional. At block 1818, the potentially inverse tone mapped super-resolved color value 930 is stored in the output buffer 506-2, representing an upscaled frame 508 having a second resolution and corresponding to the rendered frame 502.

[0094] FIG. 20 illustrates in flow chart form an overall process for generating a pixel lock for one or more pixels. It should be understood that the process described below with respect to method 2000 has been described above in more detail with reference to FIG. 9. At block 2002, the locking stage 609 obtains a first frame of the video stream. The first frame is defined by a first plurality of pixels associated with a second of the color data. At block 2004, the locking stage 609 determines that a pixel of the first plurality of pixels includes high frequency information. At block 2006, in response to determining that the pixel includes high frequency information, the locking stage 609 generates a first pixel lock for the pixel such that the color data associated with the pixel is maintained during the color accumulation process of the next frame. At block 2008, the locking stage 609 obtains pixel lock information for a second frame of the video stream. The second frame is defined by a second plurality of pixels. The pixel lock information identifies at least a second pixel lock associated with at least one pixel of the second plurality of pixels. In block 2010, the locking stage 609 reduces the remaining lifetime of at least the second pixel lock.

[0095] FIG. 21 illustrates in flow chart form a more detailed method 2100 of the pixel locking process shown at block 1614 of FIG. 17. It should be understood that the process described below with respect to method 2100 has been described in more detail above with reference to FIG. 9. At block 2102, the locking stage 609 of the upscaler 332 receives as input the adjusted color buffer 608 and the locked state texture 621 for a previously upscaled frame. At block 2104, for each pixel of the rendered frame 502, the locking stage 609 compares a defined neighborhood of intensity values ​​obtained from the adjusted color buffer 608 against an intensity threshold. At block 2106, the locking stage 609 determines whether the intensity threshold has been met. At block 2108, if the intensity threshold has not been met (e.g., if the intensity change between the neighboring pixel and the current pixel exceeds the intensity threshold), the locking stage 609 does not generate a lock for the current pixel. The process proceeds to block 2106. At block 2110, if the luminance threshold is not met (e.g., the luminance change between the neighboring pixel and the current pixel is below the luminance threshold), the locking stage 609 generates a lock for the current pixel. As part of creating the lock at block 2112, the locking stage 609 stores the remaining lifetime of the lock in the red channel of the current pixel's lock state texture 620. When the lock is first generated, the pixel's full lifetime remains. At block 2114, the locking stage 609 stores the pixel's luminance value at the time the lock was generated in the green channel of the current pixel's lock state texture 620. At block 2116, for each locked pixel in the previously upscaled frame, the locking stage 609 updates the remaining lifetime of the lock in the red channel of that pixel's lock state texture 620.

[0096] FIG. 22 illustrates in flow chart form a more detailed method 2200 of the reprojection shown at block 1806 of FIG. 18 and the lock update process shown at block 1808 of FIG. 18. It should be understood that the process described below with respect to the method 2200 has been described in more detail above with reference to FIG. 9 and FIG. 12. At block 2202, the reprojection sub-stage 905 samples the expanded motion vectors in the expanded motion vector buffer 614. The reprojection sub-stage 905 applies the samples to the output buffer 506-1 of the previously upscaled frame to generate reprojected color data 922 for the pixels of the previously upscaled frame. At block 2204, the reprojection sub-stage 905 samples the expanded motion vectors and applies the samples to the lock state texture 620 of the previously upscaled frame to generate reprojected lock 622 for the pixels of the previously upscaled frame. As described above, the expanded motion vectors are not applied to the lock state texture 620 of the previously upscaled frame for pixels where the lock state texture 620 in the rendered (current) frame indicates that a new pixel lock has been generated in the rendered frame. At block 2206, the reprojection sub-stage 905 passes the reprojected color data 922 as an input to a color correction sub-stage 909. At block 2208, the reprojection sub-stage 905 passes the reprojected lock 622 as an input to a lock update sub-stage 907.

[0097] For each reprojected pixel lock, the lock update sub-stage 907 determines whether the associated pixel is occluded based on the deocclusion mask 618, at block 2210. If the pixel is occluded, the lock update sub-stage 907 releases the reprojected lock and then processes the reprojected pixel lock, at block 2212. If the pixel is not occluded, at block 2214, the lock update sub-stage 907 maintains the reprojected pixel lock. Alternatively, or in addition to block 2212, the lock update sub-stage 907 determines whether the shading of the pixel has changed by more than a shading change threshold, at block 2216. If the shading of the pixel has changed by more than a shading change threshold, at block 2218, the lock update sub-stage 907 releases the reprojected lock and then processes the reprojected pixel lock. If the shading of the pixel has not changed by more than a shading change threshold, at block 2220, the lock update sub-stage 907 maintains the reprojected pixel lock. At block 2222, the lock update sub-stage 907 determines that the lock is reliable. At block 2224, the lock update sub-stage 907 updates any active pixel locks generated for the rendered frame, for example, by updating the luminance stored in the green channel and updating the confidence / belief factors stored in the blue channel. Once all reprojected and active pixel locks for the rendered frame have been processed, the lock update sub-stage 907 passes a set of updated pixel locks 926 to the color correction sub-stage 909.

[0098] In some embodiments, the above apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips). In at least some embodiments, electronic design automation (EDA) and computer aided design (CAD) software tools can be used to design and manufacture these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs include code executable by a computer system to operate the computer system to operate on code representing circuits of one or more IC devices to perform at least a portion of a process for designing or adapting a manufacturing system to manufacture the circuits. This code may include instructions, data, or a combination of instructions and data. The software instructions representing the design tool or manufacturing tool are typically stored in a computer readable storage medium accessible to the computing system. Similarly, code representing one or more stages of the design or manufacture of the IC device is stored in and accessed from the same computer readable storage medium or a different computer readable storage medium.

[0099] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tape, magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or micro-electromechanical systems (MEMS) based storage media. The computer-readable storage medium (e.g., system RAM or ROM) may be internal to the computing system, the computer-readable storage medium (e.g., a magnetic hard drive) may be permanently attached to the computing system, the computer-readable storage medium (e.g., an optical disk or Universal Serial Bus (USB)-based flash memory) may be removably attached to the computing system, or the computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to the computer system via a wired or wireless network.

[0100] In some embodiments, certain aspects of the techniques described above are implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored or otherwise tangibly embodied in a non-transitory computer-readable storage medium. The software may include instructions and specific data that, when executed by the one or more processors, operate the one or more processors to perform one or more aspects of the techniques described above. The non-transitory computer-readable storage medium may include, for example, a magnetic or optical disk storage device, a solid-state storage device such as a flash memory, a cache, a random access memory (RAM), or other non-volatile memory device(s), etc. The executable instructions stored in the non-transitory computer-readable storage medium may be implemented as source code, assembly language code, object code, or other form of instructions that can be interpreted or otherwise executed by one or more processors.

[0101] In addition to the above, it should be noted that not all activities or elements described in the summary description are required, some of the specific activities or devices may not be required, one or more additional activities may be performed, and one or more additional elements may be included. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will appreciate that various changes and modifications can be made without departing from the scope of the invention as set forth in the claims. Thus, the specification and drawings should be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the invention.

[0102] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, the benefits, advantages, solutions to problems, and features by which any benefit, advantage, or solution may occur or be manifested are not to be construed as critical, essential, or essential features of any or all claims. Moreover, the specific embodiments described above are illustrative only, as the disclosed invention may be modified and practiced in different but similar manners apparent to those skilled in the art having the benefit of the teachings herein. No limitations are intended to the details of construction or design herein shown, other than as set forth in the appended claims. It is therefore apparent that the specific embodiments described above may be altered or modified, and all such variations are considered to be within the scope of the disclosed invention. Accordingly, the protection sought herein is set forth in the appended claims.

Claims

1. 1. A method in an accelerated processing device, comprising: obtaining a first frame of a video stream, the first frame being defined by a plurality of pixels associated with a first set of color data; generating a pixel lock, in response to a first set of pixels of the plurality of pixels including high frequency spatial information, for each pixel of the first set of pixels such that color data associated with the pixel is maintained during a color accumulation process performed by the accelerated processing device; method.

2. The color accumulation process comprises: upscaling at least one of the first frame or the second frame of the video stream, the second frame following the first frame; or Temporal anti-aliasing at least one of the first frame or the second frame; Executed during one or more of 10. The method of claim 1.

3. For each pixel of the first set of pixels: comparing the luminance values ​​of a neighborhood of a pixel with the luminance value of said pixel; maintaining a current locked state of the pixel in response to a luminance difference between a luminance value of a neighbor of the pixel and the luminance value of the pixel exceeding a luminance threshold; generating the pixel lock in response to the luminance difference not exceeding the luminance threshold.

10. The method of claim 1.

4. generating the pixel lock storing the initial lifetime of the pixel lock in a lock state texture of the pixel; storing a luminance value of the pixel at the time the pixel lock was generated in the locked state texture.

10. The method of claim 1.

5. storing the initial lifetime of the pixel lock includes storing a value indicating whether the pixel is an initial pixel lock of the pixel or a subsequent pixel lock of the pixel. The method of claim 4.

6. generating the pixel lock further includes storing a confidence factor associated with the pixel lock in the lock state texture, the confidence factor affecting an amount of time the pixel lock is maintained for the pixel; The method of claim 4.

7. The method of claim 6, wherein storing the initial lifetime of the pixel lock comprises storing the initial lifetime of the pixel lock in a first color channel of the lock state texture; storing the luminance value of the pixel includes storing the luminance value in a second color channel of the locked state texture; storing the reliability factor includes storing the reliability factor in a third color channel of the locked state texture. The method of claim 6.

8. obtaining pixel lock information for a second frame of the video stream, the second frame being defined by the plurality of pixels associated with a second set of color data and being a previous frame of the video stream processed by the accelerated processing device, the pixel lock information identifying, for each pixel of a second set of pixels of the plurality of pixels, a pixel lock generated for the second frame; 10. The method of claim 1.

9. and, in response to obtaining the pixel lock information, decreasing a remaining lifetime of each pixel lock generated for each of the second set of pixels.

9. The method of claim 8.

10. the first frame is rendered at a first resolution and the second frame is upscaled to a second resolution greater than the first resolution; 9. The method of claim 8.

11. reprojecting the pixel lock information of the second frame onto the pixel lock information of the first frame to generate a reprojected pixel lock.

9. The method of claim 8.

12. Reprojecting the pixel lock information comprises: in response to one or more pixels of the second set of pixels not corresponding to any pixel of the first set of pixels, reprojecting at least a remaining lifetime of each pixel lock associated with the one or more pixels, and reprojecting a luminance value of each pixel of the one or more pixels when the pixel lock was generated for the pixel. The method of claim 11.

13. releasing the reprojected pixel locks in response to an indication that any of the reprojected pixel locks is unreliable; and maintaining the reprojected pixel lock in response to an indication that the reprojected pixel lock is reliable. The method of claim 11.

14. Indicating that the reprojected pixel lock is unreliable in response to a shading change of the plurality of pixels associated with the reprojected pixel lock exceeding a shading change threshold; and and indicating that the reprojected pixel lock is reliable in response to the shading change of the pixel not exceeding the shading change threshold.

14. The method of claim 13.

15. The method of claim 14, further comprising: indicating that the reprojected pixel lock is unreliable in response to any one of the plurality of pixels associated with the reprojected pixel lock being occluded; and and indicating that the reprojected pixel lock is reliable in response to the pixel being unoccluded.

14. The method of claim 13.

16. 1. A system comprising: an acceleration processing device; The acceleration processing device obtaining a first frame of a video stream, the first frame being defined by a plurality of pixels associated with a first set of color data; generating, for each pixel of the first set of pixels, a pixel lock in response to the first set of pixels of the plurality of pixels including high frequency spatial information, such that color data associated with the pixel is maintained during a color accumulation process performed by the accelerated processing device; configured to: system.

17. The accelerating processing device may further include: upscaling at least one of the first frame or the second frame of the video stream, the second frame following the first frame; or Temporal anti-aliasing at least one of the first frame or the second frame; configured to execute during one or more of 17. The system of claim 16.

18. The acceleration processing device, comparing the luminance values ​​of a neighborhood of a pixel with the luminance value of said pixel; maintaining a current locked state of the pixel in response to a luminance difference between a luminance value of a neighbor of the pixel and the luminance value of the pixel exceeding a luminance threshold; generating the pixel lock in response to the luminance difference not exceeding the luminance threshold; configured to:

17. The system of claim 16.

19. The acceleration processing device storing the initial lifetime of the pixel lock in a lock state texture of the pixel; storing the luminance value of the pixel when the pixel lock was created in the locked state texture; and generating the pixel lock by 17. The system of claim 16.

20. the acceleration processing device is configured to store an initial lifetime of the pixel lock by storing a value indicating whether the pixel is an initial pixel lock of the pixel or a subsequent pixel lock of the pixel; 20. The system of claim 19.

21. the accelerated processing device is configured to generate the pixel lock by storing a confidence factor associated with the pixel lock in the lock state texture, the confidence factor affecting the length of time the pixel lock is maintained for the pixel; 20. The system of claim 19.

22. The acceleration processing device: storing an initial lifetime of the pixel lock by storing the initial lifetime of the pixel lock in a first color channel of the lock state texture; storing a luminance value of the pixel by storing the luminance value in a second color channel of the locked state texture; configured to store the reliability factor by storing the reliability factor in a third color channel of the locked state texture.

22. The system of claim 21.

23. The acceleration processing device and configured to obtain pixel lock information for a second frame of the video stream, the second frame being defined by the plurality of pixels associated with a second set of color data and being a previous frame of the video stream processed by the accelerated processing device, the pixel lock information identifying, for each pixel of the second set of pixels of the plurality of pixels, a pixel lock generated for the second frame.

17. The system of claim 16.

24. The acceleration processing device and configured to decrease a remaining lifetime of each pixel lock generated for each of the second set of pixels in response to the pixel lock information being obtained.

24. The system of claim 23.

25. the first frame is rendered at a first resolution and the second frame is upscaled to a second resolution greater than the first resolution; 24. The system of claim 23.

26. The acceleration processing device configured to reproject the pixel lock information of the second frame onto the pixel lock information of the first frame to generate a reprojected pixel lock.

24. The system of claim 23.

27. The acceleration processing device, configured, in response to one or more pixels of the second set of pixels not corresponding to any pixel of the first set of pixels, to reproject the pixel lock information by reprojecting at least a remaining lifetime of each pixel lock associated with the one or more pixels and reprojecting a luminance value of each pixel of the one or more pixels when the pixel lock was generated for the pixel.

27. The system of claim 26.

28. The acceleration processing device, releasing the reprojected pixel locks in response to an indication that any of the reprojected pixel locks is unreliable; and maintaining the reprojected pixel lock in response to an indication that the reprojected pixel lock is reliable; and configured to:

27. The system of claim 26.

29. The acceleration processing device, indicating the reprojected pixel lock as unreliable in response to a shading change of the plurality of pixels associated with the reprojected pixel lock exceeding a shading change threshold; indicating that the reprojected pixel lock is reliable in response to the shading change of the pixel not exceeding the shading change threshold; configured to:

29. The system of claim 28.

30. The acceleration processing device, indicating the reprojected pixel lock as unreliable in response to any pixel of the plurality of pixels associated with the reprojected pixel lock being occluded; Indicating that the reprojected pixel lock is reliable in response to the pixel being unoccluded; configured to:

29. The system of claim 28.

31. A computer-readable storage medium embodying a set of executable instructions for operating at least one processor, comprising: The set of instructions obtaining a first frame of a video stream, the first frame being defined by a plurality of pixels associated with a set of color data; generating, for each pixel of the set of pixels of the plurality of pixels in response to the pixel of the set of pixels including high frequency spatial information, a pixel lock such that color data associated with the pixel is maintained during a color accumulation process performed by the accelerated processing device; causing the at least one processor to perform A computer-readable storage medium.