Sampling of Partially Resident Textures
The method addresses inefficiencies in texture mapping by using a residency map descriptor to sample partially resident texture data, reducing memory access latency and improving performance in computer graphics applications.
Patent Information
- Application Number
- JP2022554757
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-29
- Filing Date
- 2021-03-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-03-08
AI Technical Summary
Existing texture mapping techniques in computer graphics face inefficiencies when dealing with partially resident textures (PRTs), as they require multiple memory accesses and can lead to delays due to cache misses.
A method for sampling partially resident texture data using a residency map descriptor, which allows for efficient retrieval of texture data from a mipmap stored in memory by specifying the residency map and associated registers, reducing the number of memory accesses required.
This approach improves performance by reducing memory access latency and increasing efficiency in texture sampling operations, especially when dealing with PRTs, by allowing texture data to be obtained in a single instruction rather than multiple separate operations.
Smart Images

Figure 0007692431000001 
Figure 0007692431000002 
Figure 0007692431000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 994,441, filed Mar. 25, 2020, and U.S. Patent Application No. 17 / 084,520, filed Oct. 29, 2020, and is incorporated herein by reference in its entirety as if fully set forth herein.
Background Art
[0002] Texture mapping is a method for adding detail, surface texture, color, or other attributes to a computer - generated graphic or a three - dimensional (3D) model. When rendering a computer - generated graphic, one or more textures can be applied (or mapped) to each of the geometric primitives of the graphic. These textures include, for example, color, luminance, and / or other data that is mapped to each of the geometric primitives.
[0003] Mipmaps are a pre - calculated and optimized collection of texture images at different resolutions. Typically, before rendering a graphics scene, all of the textures of the scene and the related mipmaps need to be resident or partially resident in video memory. Textures that are partially resident in video memory can be called partially resident textures (PRTs).
[0004] A more detailed understanding can be obtained from the following description given by way of example with the accompanying drawings.
Brief Description of the Drawings
[0005]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
DETAILED DESCRIPTION OF THE INVENTION
[0006] Some embodiments provide a method for sampling partially resident texture data. An instruction including a residency map descriptor is received. The instruction is executed to obtain partially resident texture data from a mipmap stored in memory based on the residency map descriptor.
[0007] In some embodiments, the residency map descriptor includes a residency map. In some embodiments, the residency map descriptor includes information indicating how the residency map is stored in memory. In some embodiments, the residency map descriptor includes a specification of the register that stores the residency map descriptor. In some embodiments, the residency map descriptor includes a specification of the register that stores the residency map. In some embodiments, the instruction includes an index of an image sample and an index of a mipmap descriptor, and further includes executing an instruction to obtain partially resident texture data from a mipmap stored in memory based on the residency map descriptor, the index of the image sample, and the index of the mipmap descriptor. In some embodiments, the image sample includes coordinates of an image. In some embodiments, the index of the image sample includes a specification of the register that stores the image sample.
[0008] In some embodiments, the mipmap descriptor includes information indicating how the mipmap is stored in memory. In some embodiments, the index of the mipmap descriptor includes a specification of the register that stores the mipmap descriptor.
[0009] Some embodiments provide a processor configured to sample partially resident texture data. The processor includes circuitry configured to receive an instruction that includes a residency map descriptor. The processor also includes circuitry configured to execute an instruction to obtain partially resident texture data from a mipmap stored in memory based on the residency map descriptor.
[0010] In some embodiments, the residency map descriptor includes a residency map. In some embodiments, the residency map descriptor includes information indicating how the residency map is stored in memory. In some embodiments, the residency map descriptor includes the specifications of the register that stores the residency map descriptor. In some embodiments, the residency map descriptor includes the specifications of the register that stores the residency map. In some embodiments, the instruction includes an index of an image sample and an index of a mipmap descriptor, and the circuit is configured to execute an instruction to obtain texture data that is partially resident from a mipmap stored in memory based on the residency map descriptor, the index of the image sample, and the index of the mipmap descriptor. In some embodiments, the image sample includes the coordinates of an image. In some embodiments, the index of the image sample includes the specifications of the register that stores the image sample. In some embodiments, the mipmap descriptor includes information indicating how the mipmap is stored in memory. In some embodiments, the index of the mipmap descriptor includes the specifications of the register that stores the mipmap descriptor.
[0011] FIG. 1 is a block diagram of an exemplary device 100 that can implement one or more features of the present disclosure. Device 100 can include, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a cellular phone, or a tablet computer. Device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Additionally, device 100 can optionally include an input driver 112 and an output driver 114. It should be understood that device 100 can include additional components not shown in FIG. 1.
[0012] In various alternative examples, the processor 102 includes a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU located on the same die, or one or more processor cores, each of which may be a CPU or a GPU. In various alternative examples, the memory 104 may be located on the same die as the processor 102 or may be located separately from the processor 102. The memory 104 includes volatile or non-volatile memory (e.g., random access memory (RAM), dynamic RAM, cache).
[0013] The storage device 106 includes a fixed or removable storage device (e.g., hard disk drive, solid state drive, optical disk, flash drive). The input device 108 includes, but is not limited to, a keyboard, keypad, touch screen, touch pad, detector, microphone, accelerometer, gyroscope, biometric scanner, or network connection (e.g., wireless local area network card for transmitting and / or receiving wireless IEEE802 signals). The output device 110 includes, but is not limited to, a display, speaker, printer, tactile feedback device, one or more lights, antenna, or network connection (e.g., wireless local area network card for transmitting and / or receiving wireless IEEE802 signals).
[0014] The input driver 112 communicates with the processor 102 and the input device 108, enabling the processor 102 to receive inputs from the input device 108. The output driver 114 communicates with the processor 102 and the output device 110, enabling the processor 102 to send outputs to the output device 110. Note that the input driver 112 and the output driver 114 are optional components, and the device 100 will operate similarly in the absence of the input driver 112 and the output driver 114. The output driver 116 includes an acceleration processing device (APD) 116 coupled to the display device 118. The APD receives compute commands and graphics rendering commands from the processor 102, processes those compute commands and graphics rendering commands, and provides pixel outputs for display to the display device 118. As will be described in more detail below, the APD 116 includes one or more parallel processing units for performing computations according to the single-instruction-multiple-data (SIMD) paradigm. Thus, while various functions are described herein as being performed by or in conjunction with the APD 116, in various alternatives, the functions described as being performed by the APD 116 are not driven by the host processor (e.g., the processor 102) and are additionally or alternatively performed by other computing devices having similar capabilities to provide graphic output to the display device 118. For example, any processing system that executes processing tasks according to the SIMD paradigm is considered capable of performing the functions described herein. Alternatively, computing systems that do not perform processing tasks according to the SIMD paradigm are contemplated to perform the functions described herein.
[0015] FIG. 2 is a block diagram of device 100 showing additional details regarding the execution of processing tasks on APD 116. Processor 102 maintains one or more control logic modules for execution by processor 102 within system memory 104. The control logic modules include operating system 120, kernel mode driver 122, and application 126. These control logic modules control various features of the operation of processor 102 and APD 116. For example, operating system 120 communicates directly with the hardware and provides an interface to the hardware for other software executed on processor 102. Kernel mode driver 122 controls the operation of APD 116, for example, by providing an application programming interface (API) to software (e.g., application 126) executed on processor 102 to access various functions of APD 116. Also, kernel mode driver 122 includes a just-in-time compiler that compiles programs for execution by processing components of APD 116 (such as SIMD unit 138, described in more detail below).
[0016] APD 116 executes commands and programs for selected functions such as graphic operations and non-graphic operations suitable for parallel processing. APD 116 can be used to execute graphics pipeline operations such as pixel operations, geometric calculations, and rendering of images to display device 118 based on commands received from processor 102. Also, APD 116 executes computational processing operations not directly related to graphic operations, such as operations related to video, physical simulation, computational fluid dynamics, or other tasks, based on commands received from processor 102.
[0017] APD116 includes a computing unit 132 that includes one or more SIMD units 138 that perform operations in parallel according to the SIMD paradigm at the request of the processor 102. The SIMD paradigm is one in which multiple processing elements share a single program control flow unit and a program counter and thus execute the same program, but can execute that program with different data. In one example, each SIMD unit 138 includes 16 lanes, and each lane can execute the same instruction simultaneously with other lanes within the SIMD unit 138, but can execute that instruction with different data. The lanes can be switched off predictively when not all lanes need to execute a given instruction. Also, prediction can be used to execute programs having branch control flow. More specifically, for programs having conditional branches or other instructions where the control flow is based on calculations performed by individual lanes, prediction of lanes corresponding to currently unexecuted control flow paths and serial execution of different control flow paths enables any control flow.
[0018] The basic unit of execution within the computing unit 132 is the work item. Each work item represents a single instantiation of a program that is executed in parallel within a particular lane. Work items can be executed simultaneously as a "wavefront" on a single SIMD processing unit 138. One or more wavefronts are included in a "work group", which includes a collection of work items that are specified to execute the same program. A work group can be executed by executing each of the wavefronts that make up the work group. In an alternative example, wavefronts are executed sequentially on a single SIMD unit 138, or partially or fully in parallel on different SIMD units 138. A wavefront can be considered the maximum collection of work items that can be executed simultaneously on a single SIMD unit 138. Thus, if a command received from the processor 102 indicates that a particular program is parallelized to an extent that it cannot be executed simultaneously on a single SIMD unit 138, that program is split into wavefronts that are parallelized across two or more SIMD units 138 or serialized on the same SIMD unit 138 (or both parallelized and serialized as required). The scheduler 136 performs operations related to scheduling the various wavefronts on different computing units 132 and SIMD units 138.
[0019] The parallel processing provided by the computing unit 132 is suitable for graphics-related operations such as pixel value calculation, vertex transformation, and other graphics operations. Thus, in some cases, the graphics pipeline 134 that receives graphics processing commands from the processor 102 provides compute tasks to the computing unit 132 for parallel execution.
[0020] Further, the calculation unit 132 is used to perform calculation tasks that are not related to graphics or not part of the "normal" operation of the graphics pipeline 134 (e.g., custom operations performed to supplement the processing done on the operation of the graphics pipeline 134). An application 126 or other software executed on the processor 102 transmits a program that defines such a calculation task to the APD 116 for execution.
[0021] FIG. 3 is a block diagram showing further details of the graphics processing pipeline 134 shown in FIG. 2. The graphics processing pipeline 134 includes stages (phases) each performing a specific function. The stages represent sub - divisions of the functionality of the graphics processing pipeline 134. Each stage is implemented partially or fully as a shader program executed within the programmable processing unit 202, or partially or fully as fixed - function non - programmable hardware external to the programmable processing unit 202.
[0022] The input assembler stage 302 reads a buffer filled by the user (e.g., a buffer filled by a software request executed by the processor 102 such as the application 126) and assembles (assembles) that data into primitives used by the rest of the pipeline. The input assembler stage 302 can generate different types of primitives based on the primitive data contained in the buffer filled by the user. The input assembler stage 302 formats the assembled primitives for use by the rest of the pipeline.
[0023] The vertex shader stage 304 processes the vertices of the primitives assembled by the input assembler stage 302. The vertex shader stage 304 performs various per-vertex operations such as transformation, skinning, morphing, and per-vertex lighting for each vertex. The transformation operations include various operations for transforming the coordinates of the vertices. These operations include one or more of modeling transformation, view transformation, projection transformation, perspective division, and viewport transformation. In this specification, such a transformation is considered to change the coordinates or "position" of the vertex at which the transformation is performed. Other operations of the vertex shader stage 304 change attributes other than the coordinates.
[0024] The vertex shader stage 304 is implemented partially or fully as a vertex shader program executed on one or more compute units 132. The vertex shader program is provided by the processor 102 and is based on a program pre-written by a computer programmer. The driver 122 compiles such a computer program to generate a vertex shader program having a form suitable for execution within the compute unit 132.
[0025] The hull shader stage 306, the tessellator (mosaicizer) stage 308, and the domain shader stage 310 operate together to implement tessellation, which transforms simple primitives into more complex primitives by subdividing the primitives. The hull shader stage 306 generates patches for tessellation based on the input primitives. The tessellator stage 308 generates a set of samples for the patches. The domain shader stage 310 calculates the vertex positions of the vertices corresponding to the samples of the patches. The hull shader stage 306 and the domain shader stage 310 can be implemented as shader programs executed on the programmable processing unit 202.
[0026] The geometry shader stage 312 performs vertex operations based on primitives. Various different types of operations can be performed by the geometry shader stage 312, including operations such as point sprint expansion, dynamic particle system operations, fur-fin generation, shadow volume generation, single pass render-to-cubemap, per-primitive material swapping, and per-primitive material setup. In some cases, the shader program executed on the programmable processing unit 202 performs the operations of the geometry shader stage 312.
[0027] The rasterizer stage 314 accepts and rasterizes simple primitives and is generated upstream. Rasterization involves determining which screen pixels (or sub-pixel samples) are covered by a particular primitive. Rasterization is performed by fixed-function hardware.
[0028] The pixel shader stage 316 calculates the output values of screen pixels based on the primitives generated upstream and the results of rasterization. The pixel shader stage 316 can apply textures from texture memory. The operations of the pixel shader stage 316 are performed by the shader program executed on the programmable processing unit 202.
[0029] The output merge stage 318 accepts the outputs from the pixel shader stage 316, merges those outputs, and performs operations such as z-test and alpha blending to determine the final color of the screen pixels.
[0030] The texture data that defines a texture is stored and / or accessed by a texture unit 320. A texture is a bitmap image that is used at various points within a graphics processing pipeline 134. For example, in some cases, a pixel shader stage 316 applies a texture to pixels to improve the apparent rendering complexity (e.g., to provide a more “realistic” appearance) without increasing the number of vertices being rendered.
[0031] In some cases, a vertex shader stage 304 uses texture data from the texture unit 320 to modify primitives to increase complexity, e.g., by generating or modifying vertices for improved aesthetics. In one example, the vertex shader stage 304 uses a height map stored in the texture unit 320 to modify the displacement of vertices. This type of technique can be used to generate a more realistic appearance of water, for example, by modifying the position and number of vertices used to render water, as compared to a texture that is only used at the pixel shader stage 316. In some cases, a geometry shader stage 312 accesses texture data from the texture unit 320.
[0032] In the field of computer graphics, a texture map or texture is a set of data that can be applied to the surface of a primitive. In some cases, a texture is an image such as a bitmap image. Textures are typically two-dimensional (2D), but a texture can be one-dimensional (1D), three-dimensional (3D), or in some cases of other degrees of dimensionality. In some cases, a texture is a mathematical function rather than an image.
[0033] Textures are typically applied to 3D graphics primitives by mapping points on the texture to points on the primitive. Applying a texture to a 3D graphics object composed of such primitives is analogous to wrapping a real-world object with patterned paper or skin.
[0034] In some cases, the complexity of the scene does not require the full resolution of the texture. For example, as a primitive moves away from the viewer, it appears smaller and the complexity of the object visible to the viewer decreases. In some such cases, a lower-resolution version of the texture is loaded into memory instead of the full-resolution texture, saving memory space. The complexity of the 3D model representation from the user's perspective is called its "level of detail" (LOD). In some cases, the LOD is adjusted for other reasons. Less important objects, for example, may have a relatively low LOD in some cases.
[0035] A mipmap is a set of bitmap images of the texture, which typically includes the full-resolution texture and progressively smaller scaled-down replicas of the full-resolution texture at lower resolutions. When the complexity of the 3D graphics is sufficient, the full-resolution (i.e., highest LOD) texture is sampled from the mipmap and applied to the primitive. When the complexity of the 3D graphics is not sufficient (e.g., as the primitive moves away from the viewer, the primitive appears smaller), a lower LOD version of the texture is sampled from the mipmap and applied to the primitive. In some cases, different LOD versions of the texture from the mipmap are processed, such as by interpolation, to provide an intermediate resolution.
[0036] In some applications, all textures and mipmaps of a scene need to be resident in video memory before rendering the scene. If the video memory is not large enough for all textures and related mipmaps that should be fully resident in the video memory, it may be necessary to reduce the size, resolution, and / or LOD of the textures in order to fully resident the textures and mipmaps.
[0037] In some applications, only some of the textures and related mipmaps need to be resident in video memory. In some cases, by specifying a maximum or minimum LOD, only some of the textures and their related mipmaps need to be resident in video memory. In such cases, the texture is called a partially resident texture (PRT).
[0038] In some cases, various permutations of fully and partially resident textures are possible. For example, in some cases, all levels of the mipmap are resident in video memory and the complete texture is resident at each level.
[0039] In some cases, only some levels of the mipmap are resident in video memory, but the complete texture is resident at each resident level. In such cases, the application restricts the shader view (e.g., via a shader resource descriptor (SRD)) to a subset of the levels of the mipmap. The levels of the mipmap that are resident in memory are adjacent in such cases (i.e., no skipped levels - "no hole in the middle of the chain"), but the most detailed level, or the least detailed level (i.e., "end") may be skipped by the application in some cases. In some such cases, the hardware (e.g., texture processing unit, SIMD processor, etc.) clamps the LOD to the current mipmap level.
[0040] In some cases, only some levels of mipmaps are resident in video memory, and only a portion of the texture resides in each resident level. Similar to the case where only some levels of mipmaps are resident in video memory, a complete texture resides in each resident level, except that the memory resident levels of the mipmaps do not have to be completely resident. In such cases, in addition to clamping the LOD, the hardware assumes that the values are zero for non-resident tiles. This metric is available in some applications and hardware that support PRT (e.g., texture processing units, SIMD processors, etc.).
[0041] In some cases, some portions of the texture (e.g., tiles) reside in different levels of mipmaps than other portions. In some such cases, all or only some of the levels of the mipmaps are resident in memory. In some such cases, the hardware fetches each tile from the memory resident level of the mipmap that is closest to the computed LOD. For example, if the hardware computes LOD = 0 for a tile and the most detailed resident tile is at LOD = 1, the hardware fetches the tile at LOD = 1. On the other hand, if the hardware computes LOD = 3 for a tile and the most detailed resident tile is at LOD = 1, the hardware fetches the tile at LOD = 3. This technique is sometimes called minmip mapping or PRT++.
[0042] FIG. 4 is a block diagram showing an exemplary PRT mapping scheme 400. The PRT mapping scheme 400 includes an exemplary mipmap chain including three mipmaps 410 0 ~410 2 and virtual memory 420 and physical memory 430 (e.g., implemented in the local graphics memory of memory 104 or APD 116 as illustrated and described with respect to FIGS. 1 and 2, the memory of texture unit 320 as illustrated and described with respect to FIG. 3, or any other suitable physical memory). The mipmaps 410 0represents the LOD of the texture. In this example, the mipmap 410 0 is a version of the texture that includes the maximum level of detail (e.g., the highest resolution image) within the mipmap chain. The mipmap 410 1 is the mipmap 410 0 is a filtered or downsampled version of the mipmap 410, representing a lower detail or lower resolution LOD for the texture than the mipmap 410 0 . Similarly, the mipmap 410 2 is the mipmap 410 1 is a filtered or downsampled version of the mipmap 410, representing a lower detail or lower resolution LOD for the texture than the LOD of the mipmap 410 1 . Mipmaps and the LODs associated with each mipmap are well-known to those skilled in the art.
[0043] In the example of FIG. 4, the mipmaps 410 0 ~410 2 are divided into memory tiles. Each of the mipmaps 410 0 ~410 2 is divided into memory tiles (of the same size in this example). In this example, the mipmap 410 2 is represented by a single memory tile and represents the least detailed LOD within the mipmap chain 410 0 ~410 2 . Each of the memory tiles contains the texel (i.e., texture pixel, the basic unit of the texture map) information of its respective mipmap. Although only three mipmaps are shown in FIG. 4, those skilled in the art will understand that the mipmap chain 410 0 ~410 2is illustrative, and in some embodiments, it will be understood that the mipmap chain may include more mipmaps (providing a larger range of LOD) or fewer mipmaps (providing a smaller LOD range) than the three mipmaps shown in the texture mapping scheme 400 of FIG. 4. Further, based on the description herein, one of ordinary skill in the art will recognize that the description herein is equally applicable to 1D textures and 3D textures, as well as arrays of various textures. In some cases, if the mipmap 410 0 ~410 2 is smaller (in size) than a single memory tile, one or more of the mipmaps 410 0 ~410 2 can be loaded into a single memory tile to conserve memory resources.
[0044] In the example of FIG. 4, each of the memory tiles is associated with each address space within the virtual memory 420. For example, the memory tile (labeled "RGB") associated with the upper left corner of the mipmap 410 0 is associated with the address space 420 0 within the page table of the virtual memory 420 (shown by the dashed line with an arrow). The memory tile immediately to the right of the memory tile at the upper left corner of the mipmap 410 0 (labeled "X") is associated with the address space 420 1 within the page table of the virtual memory 420.
[0045] Note that if the size of two or more consecutive page table entries is relatively small (e.g., one or more of the mipmaps 410 0 ~410 2 can be loaded into a single memory tile), in some embodiments, these page table entries can be mapped to a single address space within the page table of the virtual memory 420.
[0046] In this example, in a raster-like fashion (e.g., a horizontal traversal from left to right, a vertical traversal downwards, and then a horizontal traversal from left to right), the mipmap 410 0 ~410 2 each of the memory tiles from is associated with each address space 420 0 ~420 20 in the page table of the virtual memory 420. The illustrated association of the split memory tiles of the mipmap 410 0 ~410 2 to the page table of the virtual memory 420 in FIG. 4 is exemplary. Those skilled in the art will recognize that other schemes are available for associating the memory tiles of the mipmap 410 0 ~310 2 with the virtual memory 420.
[0047] After the mipmap 410 0 ~410 2 is split into memory tiles, the first subset of the memory tiles is mapped to each address space within the physical memory 430 of FIG. 4. The dashed lines with arrows in FIG. 4 indicate, in this example, which of the memory tiles from the mipmap 410 0 ~410 2 are mapped to each address space within the physical memory 430.
[0048] The memory tiles within the first subset that are mapped to the physical memory 430 are selected in any suitable manner for any suitable use. The physical memory 430 is implemented by any suitable hardware, such as that illustrated and described with respect to FIGS. 1 - 3. The selection of the memory tiles is based on any suitable measure or criterion.
[0049] For example, a memory tile can be selectable for a first subset based on one or more of the following factors. Whether an application anticipates that a memory tile will be needed in a texture mapping operation, and a given LOD of a mipmap and one or more memory tiles associated with that mipmap. In other words, to reduce physical memory consumption, an application may, in some cases, resident (e.g., by issuing appropriate instructions) a particular memory tile associated with a mipmap chain.
[0050] For example, an application can instruct the processor to resident a memory tile associated with mipmap 410 2 (having the lowest, or least detailed, LOD within the mipmap chain), to resident most of the memory tiles associated with mipmap 410 1 (the next higher mipmap being the LOD for mipmap 410 2 ), and to resident a few of the memory tiles associated with mipmap 410 0 (having the most detailed LOD within the mipmap chain).
[0051] In some cases, optimization of the physical memory space allocated to the texture mapping process is realized as a result of this mapping scheme, enabling more physical memory space for memory tiles with more detailed LODs to be resident in the video memory.
[0052] An exemplary selection of the first subset of memory tiles mapped to physical memory 430 is shown by the memory tiles labeled "RGB" within mipmap chain 410 0 ~410 2 . These selected memory tiles are each address space within virtual memory 420 (e.g., address spaces 410 0 、410 5 、4106 and 410 13 and 410 16~18 and 420 20 are associated with. In the first subset of memory tiles, the mipmaps 410 0 ~410 2 from the memory tiles (e.g., the address space 410 from the virtual memory 420 1~4 , 410 7~12 , 410 14 , 410 15 , and 410 19 ) are marked as invalid (the memory tiles labeled with "X" in FIG. 4) and are not mapped to the physical memory 430.
[0053] The first subset of memory tiles mapped to the physical memory 430 may, in some cases, share the same texture information. In particular, the address spaces within the page table of the virtual memory 420 may, in some cases, be mapped to the same address spaces within the physical memory 430. For example, the address spaces 410 5 and 410 6 in the page table of the virtual memory 420 are each mapped to the address space 432 within the physical memory 430. Similarly, the address spaces 420 0 and 410 18 in the page table of the virtual memory 420 are each mapped to the address space 434 within the physical memory 430.
[0054] After the first subset of memory tiles is mapped to the physical memory 430 (i.e., resident in the physical memory 430), this information is accessible to the application (e.g., via the API and driver modules such as the kernel mode driver 122 illustrated and described with respect to FIG. 2) during the texture mapping operation (e.g., of the application 126 as illustrated and described with respect to FIG. 2).
[0055] The stored portions of the texture and associated mipmaps are called PRTs because only a portion of the texture and associated mipmaps need to be resident in video memory for texture mapping operations. This is in contrast to texture mapping operations that require textures to be fully resident.
[0056] To apply PRTs to primitives, the application issues texture sampling instructions to the processor, which may, in some cases, return a status code in response. For example, the status code indicates the success or failure of fetching the texture information requested by the application. In some cases, the failure to fetch texture information is the result of an attempt to access information from a mipmap that has a higher level of detail than the mipmaps currently resident in video memory.
[0057] In some cases, a residency map is used to track partially resident textures and their associated mipmaps. The residency map is sometimes also called a "mini mipmap" or LOD index table.
[0058] FIG. 5 is a diagram of an exemplary residency map 510 showing the available (i.e., memory resident) tiles of mipmaps 410 0 ~410 2 as illustrated and described in connection with FIG. 4. The residency map 510 indicates the memory resident tiles corresponding to the mipmap tiles of interest. For example, with respect to the residency map 510, the memory tile located at the upper left corner (labeled "0") corresponds to the memory resident tile located at the upper left corner of mipmap 410 0 (labeled "RGB"). Here, the label "0" is at the level of mipmap 410 0 and at lower detail levels 410 1 and 410 2However, it indicates that a texture is available for that tile.
[0059] In this example, the PRT and its associated mipmaps are tracked using a single resident map (e.g., only using resident map 510). In such a case, the value 0 in the memory tile located at the upper left corner indicates that the texture is available at the corresponding upper left corner at all levels of the mipmap chain, and the value 1 in the memory tile located at the upper right corner indicates that the texture is available at the corresponding upper right corner at levels 1 and 2. In other words, the resident map represents the availability of the coarser (i.e., less detailed) mipmaps at the indicated levels and in the mipmap chain.
[0060] For example, for a non-resident memory tile such as the memory tile located at the upper right corner of the mipmap 410 (labeled "X") 0 memory tile information from a coarser (e.g., less detailed) LOD mipmap 410 1 can be used. This is shown in the resident map 510, and the memory tile located at the upper right corner (labeled "1") indicates that the texture is available at the corresponding upper right corner at level 1 (i.e., at the upper right corner of the mipmap 410 0 labeled "RGB") and level 2.
[0061] In some embodiments, the application and / or the processor tracks resident texture information via the resident map 510. In some embodiments, sampling is processed using two dependent operations to sample the resident map-corresponding texture.
[0062] In some embodiments, the first operation is a sampling instruction that includes coordinates read from a texture resource and a resource descriptor of a residency map. Also, the second operation is a sampling operation that includes coordinates read from a texture, a sampled LOD value, and a resource descriptor of a texture resource. In some embodiments, the resource descriptor includes information describing a texture resource (i.e., a mipmap) that includes its dimensions, size, format, whether it has a residency map, etc. In some embodiments, the residency map descriptor includes information describing a residency map that includes its dimensions, size, format, etc.
[0063] Accordingly, in some embodiments, a first texture instruction is executed to access a residency map 510 and calculate a specific LOD for fetching within a page table of virtual memory 420. Here, the first operation maps image coordinates onto the residency map based on the residency map resource descriptor, samples an LOD value from the residency map, and returns the sampled LOD value.
[0064] In some embodiments, the first texture instruction includes a descriptor of the residency map and a sample of the image to which the texture is applied, and in response to the first texture instruction, the processor returns an LOD corresponding to the sample from the residency map.
[0065] Entries in the residency map 510 provide information about textures resident in video memory and related mipmaps, so that an application can determine the LODs available for texture sampling instructions.
[0066] Thus, in some embodiments, the application issues a second texture instruction to the processor to access the mipmap based on the LOD returned by the first texture instruction (e.g., clamping the LOD to a specific maximum LOD), the sample of the image to which the texture is applied, and the descriptor of the texture. In response to the second texture instruction, the processor returns texture information corresponding to the image sample. Here, the second operation maps the image coordinates onto the mipmap based on the mipmap resource descriptor, samples the texture from the mipmap, and returns the sampled texture information.
[0067] FIG. 6 is a block diagram showing an exemplary memory access for the first texture sampling instruction described above. The memory access is described with respect to the SIMD register file 600. The SIMD register file 600 can be implemented as part of the SIMD unit 138 as illustrated and described with respect to FIG. 2, or in any other suitable hardware or another way as shown in FIG. 2. Also, FIG. 6 shows a memory 602 that also reflects any relevant cache entries. The memory 602 can be implemented in or by the local graphics memory of the memory 104 or the APD 116, the memory of the texture unit 320 as illustrated and described with respect to FIG. 3, or any other appropriate physical memory, as illustrated and described with respect to FIGS. 1 and 2.
[0068] In this example, the residency map descriptor 604 and the texture descriptor 606 are stored in the SIMD register 600. The processor obtains the LOD 612 from the residency map 614 stored in the memory 602 and executes the first texture sampling instruction 608 based on the residency map descriptor 604 and the sample 610 of the image (e.g., as described above) to store the LOD 612 in the SIMD register 600.
[0069] FIG. 7 is a block diagram showing an exemplary memory access for the second texture sampling instruction described above. After LOD612 becomes available in SIMD register 600, the processor executes a second texture sampling instruction 700 based on LOD612, texture descriptor 606, and image sample 610 to obtain texture data 702 from mipmap 616. In this example, the obtained texture data 702 is stored in SMID register 600.
[0070] The texture sampling described above with respect to the exemplary first and second texture sampling instructions described above requires two separate memory accesses to fetch texture data 702, one for resolving the LOD and one for fetching the texture. This dependent fetch operation has the potential to cause delays from up to two separate cache misses. Thus, in some embodiments, performance may be improved by reducing the number of video memory accesses to 1 in some cases.
[0071] FIG. 8 is a block diagram showing an exemplary memory access for an exemplary PRT texture fetch instruction. In this example, the residency map descriptor 804 includes the residency map 614. The processor executes a PRT texture fetch instruction 800 based on the residency map descriptor 804, the sample 610 of the image to be texture processed, and the texture descriptor 606 to obtain texture data 702 from mipmap 616. In this example, the obtained texture data 702 is stored in SMID register 600.
[0072] The resident map descriptor 804 includes, in this example, the resident map 614. Thus, the resident map descriptor 804 is of a larger size than the resident map descriptor 604 illustrated and described with respect to FIGS. 6 and 7. However, since the resident map descriptor 804 itself includes a resident map, the determination of the LOD (corresponding to the LOD 612 described with respect to FIGS. 6 and 7) does not require an initial instruction to obtain the LOD from the resident map 614 stored in the video memory 602, but can be performed as part of the execution of the PRT texture fetch instruction 800.
[0073] In some cases, obtaining the texture data 702 using only one instruction (e.g., the PRT texture fetch instruction 800) instead of two has the advantage of improving performance by reducing the processing time due to memory access latency. In some embodiments, this also applies when the size of the resident map descriptor (e.g., the resident map descriptor 804) is larger than in the case of two instructions, for example, to include the resident map as an immediate value (e.g., the resident map 614). In some embodiments, the resident map descriptor is not actually larger in size than in the case of two instructions. For example, the resident map descriptor 604 includes fields (e.g., a pointer to memory) that do not require the resident map descriptor 804 (e.g., because it includes an immediate value and does not specify memory). Thus, in some cases, the resident map descriptor for the case of one instruction (immediate value) is the same size as or smaller than the resident map descriptor for the case of two instructions (pointer).
[0074] In some embodiments, the application can control the size of the residency map descriptor, including the actual residency map (i.e., as an immediate value in the case of one instruction), by controlling the number of tiles into which the texture is divided. For example, in some embodiments, a residency map descriptor with an embedded immediate residency map can be smaller than, the same size as, or larger than a residency map descriptor without an embedded immediate residency map (e.g., specifying the residency map in memory) by adjusting the number of tiles, and the more tiles there are, the larger the embedded residency map and thus the larger the descriptor.
[0075] It should be understood that many variations are possible based on the disclosure herein. Although features and elements are described above in specific combinations, each feature or element can be used alone without other features and elements or in various combinations with or without other features and elements.
[0076] The various functional units shown in the figures and / or described in this specification (including, but not limited to, processor 102, input driver 112, input device 108, output driver 114, output device 110, acceleration processing device 116, scheduler 136, graphics processing pipeline 134, computing unit 132, SIMD unit 138, texture processing unit 320, etc.) may be implemented as a general-purpose computer, processor or processor core, or as a program, software or firmware stored in a non-transitory computer-readable storage medium or another medium executable by a general-purpose computer, processor or processor core. The provided method can be implemented in a general-purpose computer, processor or processor core. Suitable processors include, by way of example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application specific integrated circuits (ASICs), field programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors can be manufactured by configuring the manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data including netlists (such instructions can be stored on a computer-readable medium). The result of such processing may be a mask work, which is then used in a subsequent semiconductor manufacturing process to manufacture a processor implementing aspects of the embodiments.
[0077] The methods or flow diagrams provided herein may be implemented in a computer program, software, or firmware incorporated in a non-transitory computer-readable storage medium for execution by a general-purpose computer or processor. Examples of non-transitory computer-readable storage media include read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media (e.g., CD-ROM disks and digital versatile disks (DVDs)).
Claims
1. A method for sampling texture data that is partially resident, comprising: receiving an instruction including a residency map descriptor, wherein the residency map descriptor includes a residency map; and executing the instruction to obtain partially resident texture data from a mipmap stored in memory based on the residency map descriptor. A method.
2. The method of claim 1, wherein the residency map descriptor includes information indicating how the residency map is stored in the memory. The method of claim 1.
3. The method of claim 1, wherein the residency map descriptor includes specifications of a register for storing the residency map descriptor. The method of claim 1.
4. The method of claim 1, wherein the residency map descriptor includes specifications of a register for storing the residency map. The method of claim 1.
5. The instruction includes an index of an image sample and an index of a mipmap descriptor, and further comprises executing the instruction to obtain partially resident texture data from a mipmap stored in the memory based on the residency map descriptor, the index of the image sample, and the index of the mipmap descriptor. The method of claim 1.
6. The method of claim 5, wherein the image sample includes coordinates of an image. The method of claim 5.
7. The method of claim 5, wherein the index of the image sample includes specifications of a register for storing the image sample. The method of claim 5.
8. The method of claim 5, wherein the mipmap descriptor includes information indicating how the mipmap is stored in the memory. The method of claim 5.
9. The method of claim 5, wherein the index of the mipmap descriptor includes specifications of a register for storing the mipmap descriptor. The method of claim 5.
10. A processor configured to sample partially resident texture data, comprising: a circuit configured to receive an instruction including a residency map descriptor, wherein the residency map descriptor includes a residency map; and a circuit configured to execute the instruction to obtain partially resident texture data from a mipmap stored in memory based on the residency map descriptor. A processor.
11. The processor of claim 10, wherein the residency map descriptor includes information indicating how the residency map is stored in the memory. The processor of claim 10.
12. The resident map descriptor includes specifications of registers that store the resident map descriptor. The processor of claim 10.
13. The resident map descriptor includes specifications of registers that store the resident map. The processor of claim 10.
14. The instruction includes an index of an image sample and an index of a mipmap descriptor, and the circuit is configured to execute the instruction to obtain partially resident texture data from the mipmap stored in the memory based on the resident map descriptor, the index of the image sample, and the index of the mipmap descriptor. The processor of claim 10.
15. The image sample includes coordinates of an image. The processor of claim 14.
16. The index of the image sample includes specifications of registers that store the image sample. The processor of claim 14.
17. The mipmap descriptor includes information indicating how the mipmap is stored in the memory. The processor of claim 14.
18. The index of the mipmap descriptor includes specifications of registers that store the mipmap descriptor. The processor of claim 14.
Citation Information
Patent Citations
Residency map descriptors
US10540802B1
Partially Resident Textures
US20120147028A1
Methods of and apparatus for using textures in graphics processing systems
US20140152684A1