Input / output filter unit, control method, and computer-readable storage medium
Patent Information
- Application Number
- CN202110662530.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-17
- Filing Date
- 2021-06-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-06-15
Smart Images

Figure CN113888389B_ABST
Abstract
Description
Technical Field
[0001] This application relates to a method and unit for performing filtering in a graphics processing unit, specifically, performing texture filtering. Background Technology
[0002] A graphics processing unit (GPU) can be used to process geometric data (e.g., vertices defining primitives or patches) generated by an application in order to generate image data. For example, a GPU can determine the pixel values (e.g., color values) of an image to be stored in a frame buffer, which can then be output to a display.
[0003] The GPU can process received geometry data in two phases—the geometry processing phase and the rasterization phase. During the geometry processing phase, vertex shaders are applied to the geometry data (e.g., the vertices of bounded primitives or patches) received from the application (e.g., a game application) to transform the geometry data into rendering space (e.g., screen space). Other functions such as clipping and culling can also be performed in the geometry processing phase to remove geometry data (e.g., primitives or patches) that fall outside the view frustum, and / or lighting / attribute processing.
[0004] During the rasterization stage, transformed primitives are mapped to pixels and a color is assigned to each pixel. This can include rasterizing the transformed geometry data (e.g., by performing a scan transformation) to generate primitive fragments. The term "fragment" is used herein to refer to a sample of primitives at a sampling point that will be processed to render pixels of the image. In some examples, there may be a one-to-one mapping from pixels to fragments. However, in other examples, there may be more fragments than pixels, and this oversampling can allow for higher quality rendering of pixel values.
[0005] Hidden primitive fragments (e.g., those hidden by other fragments) can then be removed using a process called hidden surface removal. Textures and / or shading can then be applied to the unhidden primitive fragments to determine the pixel values of the rendered image. For example, in some cases, the color of a fragment can be identified by applying a texture to it. As those skilled in the art know, a texture, also known as a texture map, is an image used to represent pre-computed colors, lighting, shadows, etc. A texture map is formed by multiple texels (i.e., color values), which can also be referred to as texture elements or texture pixels. Applying texture to a fragment typically involves mapping the fragment's position in rendering space to a position or location in the texture, and using the color of that position in the texture as the texture color of the fragment. As described below, the texture color can then be used to determine the final color of the fragment. A fragment whose color is determined based on a texture can be called a texture-mapped fragment.
[0006] Since fragment locations rarely map directly to specific texels, the texture color of a fragment is typically identified through a process called texture filtering. In its simplest case, which can be called point sampling or point filtering, a fragment is mapped to a single texel (e.g., the texel closest to the location of interest), and the texel value (i.e., the color) can be used as the texture color of the fragment. However, in most cases, more sophisticated filtering techniques are used to determine the texture color of a fragment, combining multiple texels that are close to relevant locations in the texture. For example, filtering techniques, such as, but not limited to, bilinear, trilinear, or anisotropic filtering, can be used to combine multiple texels that are close to relevant locations in the texture to determine the texture color of a fragment.
[0007] The texture color output through texture filtering can then be used as input to a fragment shader. As those skilled in the art know, a fragment shader (which may alternatively be called a pixel shader) is a program (e.g., set instructions) that operates on individual fragments to determine their color, brightness, contrast, etc. A fragment shader may receive a fragment (e.g., its position) and one or more other input parameters (e.g., texture coordinates) as input and output a color value according to a specific shader program. In some cases, the output of the pixel shader may be further processed. For example, when there are more samples than pixels, anti-aliasing techniques, such as multi-sample anti-aliasing (MSAA), can be used to generate the color of a particular pixel from multiple samples (which may be referred to as subsamples). Anti-aliasing techniques apply filters, such as, but not limited to, box filters, to multiple samples to generate a single color value for the pixel.
[0008] A GPU that performs hidden surface removal before texturing and / or shading is said to be implementing "deferred" rendering. In other examples, the GPU may not implement deferred rendering, in which case texturing and shading can be applied to fragments before performing hidden surface removal on them. In either case, the rendered pixel values can be stored in memory (e.g., a frame buffer).
[0009] Because texture filtering and pixel filtering (e.g., MSAA filtering) are complex operations, GPUs can have dedicated hardware to perform texture filtering and pixel filtering, rather than programming one or more ALUs (Arithmetic Logic Units) to perform the filtering. For example, now refer to Figure 1The example GPU 100 is illustrated. The example GPU 100 includes multiple ALU clusters 102 (which may be referred to as unified shading clusters), each ALU cluster including multiple ALUs that can be configured to execute various types of shaders generated by one of multiple data master components 104, 106, 108 (e.g., vertex shaders running during the geometry processing phase, fragment / pixel shaders running during the rasterization phase, and compute shaders). For example, in Figure 1 In the GPU 100, there are a vertex data master unit 104 for initiating or generating vertex shader tasks, a pixel data master unit 106 for initiating or generating pixel or fragment shader tasks, and a computation data master unit 108 for initiating or generating computation shader tasks.
[0010] exist Figure 1 In the example, microcontroller 110 receives vertices, pixels, and computation tasks from a host (e.g., a central processing unit (CPU)) and causes the corresponding data master components 104, 106, and 108 to generate or initiate tasks. For example, when microcontroller 110 receives a vertex task, it can be configured to cause vertex data master component 104 to generate the task. In response to receiving a task request from microcontroller 110, data master components 104, 106, and 108 generate the task and send it to scheduler 112 (which may also be referred to as a coarse-grained scheduler), where it is added to a task queue. Scheduler 112 is configured to allocate resources to tasks in the queue, then schedule the tasks and publish them to ALU clusters 102 (e.g., to fine-grained schedulers (FGS) within the ALU cluster). Each ALU cluster 102 then schedules and executes the tasks received from scheduler 112 (e.g., via FGS).
[0011] As mentioned above, in some cases, during the rasterization stage, texture filtering identifies the texture color of one or more fragments. During the texture filtering process, the position or location of a specific fragment within the texture from which it is drawn (which may be referred to as correlated texture coordinates or mapped texture coordinates) is identified, one or more texels located near the identified position (which may be referred to as correlated texels) are read from the texture, and the texture color of the fragment is determined by applying one or more filters to the correlated texels. To perform texture filtering efficiently, Figure 1The GPU 100 has a dedicated unit for performing texture filtering, referred to as texture unit 114. Exemplary texture filtering methods or techniques that can be implemented by texture unit 114 include, but are not limited to: bilinear filtering, wherein four texels closest to the identified texture location are read and combined by a distance-weighted average to produce a texture color for a fragment; trilinear filtering, comprising performing a texture lookup and bilinear filtering at two closest mipmap levels (one with higher detail and one with lower detail), and then linearly interpolating the results to produce a texture color for the fragment; anisotropic filtering, wherein several texels around the identified texture location are read, but on a sample pattern mapped according to the projected shape of the texture at that fragment; and percentage asymptotic filtering (PCF), which uses depth comparison to determine the texture color of the fragment. Thus, texture unit 114 is configured to acquire one or more samples (i.e., texels) from a texture stored in memory (not shown), perform a filtering operation on the acquired samples (i.e., texels) according to a texture filtering method, and provide the result of the filtering operation to an ALU cluster as input, for example, for a fragment / pixel shader task. Specifically, as described above, the texture color of the fragment generated by texture unit 114 can be provided to the ALU cluster as input to fragment / pixel shader tasks (e.g., tasks generated by pixel data master component 106). In some cases, memory (not shown) can be accessed via one or more interfaces 118 and / or system-level cache 120.
[0012] As mentioned above, in some cases, the output of the pixel shader can be further processed before it is output. Specifically, one or more filters can be applied to the output of the pixel shader (referred to herein as pixel filtering) to implement one or more post-processing techniques. For example, when there are more samples than pixels, a box filter or another filter can be applied to the output of multiple samples to implement anti-aliasing techniques, such as, but not limited to, MSAA, to generate the color of a specific pixel. To perform this pixel filtering efficiently, Figure 1The GPU 100 has a dedicated unit called the pixel backend 116, which is configured to receive the output of a fragment / pixel shader task from the ALU cluster 102, determine the individual pixel colors from it, and output the pixel colors to memory. In some cases, this may include, for example, applying a box filter to the data received from the ALU cluster 102 to implement MSAA, or downsampling the data received from the ALU cluster 102 and writing the filtered output to memory. However, in other cases, this may simply include outputting the received pixels. The pixel backend 116 may also be able to perform format conversions. For example, the pixel backend 116 may receive color values in one format (e.g., 16-bit floating-point format (FP16)) and output color values in another format (e.g., 8-bit fixed-point or integer format (8INT)).
[0013] The embodiments described below are provided by way of example only and do not limit the implementation methods that address any or all of the drawbacks of known methods and hardware for performing texture filtering and pixel filtering. Summary of the Invention
[0014] This summary is provided to introduce some concepts that are further described in the following detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0015] This document describes an input / output filter unit for a graphics processing unit. The input / output filter unit includes: a first buffer configured to store data received from and output to a first component of the graphics processing unit; a second buffer configured to store data received from and output to a second component of the graphics processing unit; a weight buffer configured to store filter weights; a filter group configured to perform any of a plurality of types of filtering on an input data set, the plurality of types of filtering including one or more texture filtering types and one or more pixel filtering types; and control logic configured to cause the filter group to: (i) perform one of the plurality of types of filtering on a data set stored in one of the first and second buffers using a set of weights stored in the weight buffers, and (ii) store the result of the filtering in one of the first and second buffers.
[0016] A first aspect provides an input / output filter unit for a graphics processing unit, the input / output filter unit comprising: a first buffer configured to store data received from and output to a first component of the graphics processing unit; a second buffer configured to store data received from and output to a second component of the graphics processing unit; a weight buffer configured to store filter weights; a filter group configured to perform any of a plurality of types of filtering on an input data set, the plurality of types of filtering including one or more types of texture filtering and one or more types of pixel filtering; and control logic configured to cause the filter group to: (i) perform one of the plurality of types of filtering on a data set stored in one of the first and second buffers using a set of weights stored in the weight buffers, and (ii) store the filtering result in one of the first and second buffers.
[0017] A filter group may include one or more filter blocks, each filter block including multiple arithmetic components, which may be selectively enabled to enable the filter group to perform one of a variety of filtering types.
[0018] The plurality of arithmetic components can be configured to form a pipeline.
[0019] The plurality of arithmetic components may include a set of arithmetic components forming an n-input x n-weighted filter, where n is an integer.
[0020] The set of arithmetic components may include n multiplier components, each of which is configured to multiply an input value and a weight, and a plurality of adder components forming an adder tree, the adder tree being configured to produce the sum of the outputs of the n multipliers.
[0021] The plurality of arithmetic components may further include n comparators, each of which is configured to compare input values and provide the result of the comparison as input to an n-input x n-weighted filter.
[0022] The plurality of arithmetic components may further include a scaling component configured to receive the output of the n-input x n-weighted filter and generate a scaled version of the output.
[0023] The filter group may include multiple filter blocks.
[0024] The control logic can be configured to cause the filter group to perform one of the multiple types of filtering on a data set stored in one of the first and second buffers by having one of the filter blocks perform a first part of the filtering of the type in a first round of the filter block and a second part of the filtering of the type in a second round of the filter block.
[0025] Temporary data may be generated during at least one of the first and second rounds, and the temporary data is stored in one of the first and second buffers.
[0026] The texture filtering of one or more types may include one or more of bilinear filtering, trilinear filtering, anisotropic filtering, and percentage asymptotic filtering.
[0027] The one or more types of pixel filtering may include one or more of downsampling, upsampling, and multisampling anti-aliasing box filtering.
[0028] The filter group can be further configured to perform texture blending.
[0029] The filter group can be further configured to perform a set of convolutional operations as part of the convolutional layers processing the neural network.
[0030] The input / output filter unit may further include a texture address generator configured to generate addresses of one or more relevant texels for performing a type of texture filtering on a fragment or pixel.
[0031] The input / output filter unit may further include a weight generator configured to generate the weight set for performing one or more types of filtering and store the generated weights in the weight buffer.
[0032] The first component may be a cluster of arithmetic logic units configured to perform coloring tasks, and the second component is a memory.
[0033] The control logic can be configured to cause the filter group to perform a filtering task among multiple filtering tasks, including a texture filtering task and a pixel filtering task. Causing the filter group to perform a pixel filtering task may include: causing the filter group to use a set of weights stored in the weight buffer via an arithmetic logic unit cluster to perform one or more types of pixel filtering on a data set stored in a first buffer, and storing the result of the pixel filtering in a second buffer for output to the memory. Causing the filter group to perform a texture filtering task may include: causing the filter group to use a set of weights stored in the weight buffer to perform one or more types of texture filtering on a data set from memory stored in the second buffer, and storing the result of the texture filtering in the first buffer for output to the arithmetic logic unit cluster.
[0034] The control logic can be configured to store the filtered results in another of the first and second buffers.
[0035] The input / output filter unit can be contained in hardware on an integrated circuit.
[0036] A second aspect provides a method for controlling an input / output filter unit, the input / output filter unit including a first buffer, a second buffer, a weight buffer, and a configurable filter group, the method comprising: receiving information identifying a filtering task, the information identifying the filtering task including an identification of a data set stored in one of the first and second buffers, a weight set stored in the weight buffer, and information on one type of filtering among multiple types of filtering, wherein the multiple types of filtering include one or more types of texture filtering and one or more types of pixel filtering; causing the configurable filter group to: perform the identified type of filtering on the identified data set using the identified weight set; and storing the filtering result in one of the first and second buffers.
[0037] The third aspect provides a graphics processing unit that includes the input / output filter unit of the first aspect.
[0038] The input / output filter units and graphics processing units described herein can be included in hardware on an integrated circuit. A method for manufacturing the input / output filter units or graphics processing units as described herein in an integrated circuit manufacturing system can be provided. An integrated circuit definition dataset can be provided, which, when processed in an integrated circuit manufacturing system, configures the system to manufacture the input / output filter unit or graphics processing unit. A non-transitory computer-readable storage medium can be provided, on which a computer-readable description of the input / output filter unit or graphics processing unit is stored, which, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture an integrated circuit including the input / output filter unit or graphics processing unit.
[0039] An integrated circuit manufacturing system may be provided, comprising: a non-transitory computer-readable storage medium storing a computer-readable description thereon of an input / output filter unit or a graphics processing unit described herein; a layout processing system configured to process the computer-readable description to generate a circuit layout description of an integrated circuit including the input / output filter unit or the graphics processing unit; and an integrated circuit generation system configured to manufacture the input / output filter unit or the graphics processing unit according to the circuit layout description.
[0040] Computer program code for performing the methods described herein may be provided. A non-transitory computer-readable storage medium having computer-readable instructions stored thereon may be provided, which, when executed at a computer system, cause the computer system to perform the methods described herein.
[0041] As will be apparent to those skilled in the art, the above features can be appropriately combined, and can be combined with any aspect of the examples described herein. Attached Figure Description
[0042] The example will now be described in detail with reference to the accompanying drawings, in which: Figure 1 This is a block diagram of a first exemplary graphics processing unit; Figure 2 This is a block diagram of a second exemplary graphics processing unit, including an input / output filter unit; Figure 3 It includes one or more filter blocks. Figure 2 A block diagram of an exemplary implementation of an input / output filter unit; Figure 4 This is a schematic diagram illustrating bilinear filtering; Figure 5 yes Figure 3 A block diagram illustrating an exemplary implementation of a filter block; Figure 6 It is control Figure 3 A flowchart illustrating an exemplary method for an input / output filter unit; Figure 7 This is a block diagram of an exemplary computer system in which the input / output filter unit and / or graphics processing unit described herein can be implemented; and Figure 8 This is a block diagram of an exemplary integrated circuit manufacturing system for generating integrated circuits that include the input / output filter units and / or graphics processing units described herein.
[0043] The accompanying drawings illustrate various examples. Those skilled in the art will understand that the element boundaries (e.g., boxes, groups of boxes, or other shapes) shown in the drawings represent one example of a boundary. In some examples, it may be that one element can be designed as multiple elements, or multiple elements can be designed as one element. Where appropriate, common reference numerals are used throughout the drawings to indicate similar features. Detailed Implementation
[0044] The following description is given by way of example to enable those skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be readily apparent to those skilled in the art. Embodiments are described by way of example only.
[0045] The inventor has discovered that, Figure 1 The texture unit 114 and pixel backend 116 perform similar filtering or analogous filtering operations on similar data, but they are separate hardware units. This leads to inefficiency because it duplicates hardware, unnecessarily increasing the cost and complexity of the GPU. This can be solved by replacing the texture unit 114 and pixel backend 116 with a single input / output (I / O) filter unit, which can be dynamically configured to perform texture filtering, pixel filtering, or both. This not only avoids duplication of filter logic but also increases GPU throughput and improves load balancing when texture filtering tasks outnumber pixel filtering tasks or vice versa. For example, when there is no pixel filtering to perform, it allows all I / O filter unit resources to be used for texture filtering instead of leaving the pixel backend 116 idle. Similarly, when there is no texture filtering to perform, it allows all I / O filter unit resources to be used for pixel filtering instead of leaving the texture unit 114 idle or unused. Furthermore, the connections (e.g., wires) between ALU cluster 102 and other components are expensive, and therefore replacing the two units connected to ALU cluster 102 with one allows for a reduction in the number of connections (e.g., wires) from ALU cluster 102.
[0046] Therefore, this document describes an input / output filter unit for a graphics processing unit (GPU) that can perform both texture filtering and pixel filtering. Texture filtering includes, for example, bilinear and trilinear filtering; pixel filtering includes, for example, filtering, downsampling, and / or upsampling for implementing MSAA or other anti-aliasing techniques. Specifically, the input / output filter unit described herein includes a filter group comprising one or more filter blocks. Each filter block can be dynamically configured to perform one of several types of filtering on the input data. The various types of filtering include one or more types of texture filtering and one or more types of pixel filtering. Input data can be received from an ALU cluster, and the filtering result can be output to memory; or input data can be received (or read) from memory, and the filtering result can be output to an ALU cluster as input for a task performed by the ALU cluster (e.g., a fragment / pixel shader task). In some cases, at least two filter blocks exist, thus allowing multiple filtering tasks to be executed in parallel.
[0047] Now for reference Figure 2 It shows an exemplary GPU 200 including an input / output filter unit 202. Figure 2 The GPU 200 is similar to Figure 1 The GPU 100 is characterized by comprising multiple ALU clusters 102, vertex data master units 104, pixel data master units 106, compute data master units 108, a microcontroller 110, and a scheduler 112, as described above regarding... Figure 1 It works as described. However, Figure 2 The GPU 200 includes a single input / output (I / O) filter unit 202 capable of performing texture filtering and pixel filtering, instead of... Figure 1 Like GPU 100, it includes a separate texture unit 114 and a pixel backend 116. Specifically, the input / output filter unit 202 is capable of (i) texture filtering of texels read from memory to generate data (e.g., texture color) that can be used as input to a task performed by the ALU cluster, and (ii) pixel filtering of pixels / samples (e.g., color values) generated by the ALU cluster 102. For example, the input / output filter unit 202 may include a filter group comprising one or more filter blocks, each of which can be dynamically configured to apply one of several types of filtering to the input data. (Refer to...) Figure 3 An exemplary implementation of the input / output filter unit 202 is described.
[0048] Now for reference Figure 3 The figure shows Figure 2An exemplary implementation of the input / output filter unit 202 is provided. In this example, the input / output filter unit 202 includes a first data buffer, which may be referred to as an ALU-side buffer 302; a second data buffer, which may be referred to as a memory-side buffer 304; a filter group 306; and a weight buffer and a control logic unit 308. In some cases, the input / output filter unit 202 may also include a texture address generator 310 and / or a weight generator 312.
[0049] ALU-side buffer 302 is configured to temporarily store data received from and sent to the ALU cluster 102. The data stored in ALU-side buffer 302 can be used as input to filtering tasks performed by filter group 306, or it can be the output of filtering tasks performed by filter group 306. For example, ALU-side buffer 302 can be configured to store: (i) the results of pixel shader tasks received from the ALU cluster 102, which are used as input to pixel filtering tasks performed by filter group 306, and (ii) the results of texture filtering tasks performed by filter group 306.
[0050] The memory-side buffer 304 is configured to temporarily store data received from and sent to memory. The data stored in the memory-side buffer 304 can be used as input to a filtering task performed by the filter group 306, or it can be the result of a filtering task performed by the filter group 306, which can be written to memory. For example, the memory-side buffer 304 can be configured to store: (i) the result of a pixel filtering task performed by the filter group 306, which is sent to memory (not shown), and (ii) texels read from memory, which are used as input to a texture filtering task performed by the filter group 306. Therefore, the ALU-side buffer 302 and the memory-side buffer 304 can be collectively referred to as data buffers, which store the input and result of the filtering tasks performed by the filter group 306.
[0051] In addition to storing the filtering tasks performed by filter group 306 and the inputs and results of those tasks, the data buffer (ALU-side buffer 302 or memory-side buffer 304) can also be used to store intermediate data generated during the filtering tasks performed by filter group 306. For example, as described in more detail below, some filtering tasks may require multiple rounds of filter group 306 to complete. Specifically, filter group 306 may only be able to perform a certain number of operations at a time, so complex filtering may be performed by filter group 306 for more than a few rounds. In these cases, intermediate data can be generated through one or more rounds of the filter group, which is used as input for subsequent rounds. This intermediate data can be stored in ALU-side buffer 302 or memory-side buffer 304, depending on, for example, which buffer provides the input for that round. For example, if the input to a round of filter group 306 is provided by memory-side buffer 304, the intermediate data generated by that round can be stored in ALU-side buffer 302; if the input to a round of filter group 306 is provided by ALU-side buffer 302, the intermediate data generated by that round can be stored in memory-side buffer 304.
[0052] Filter set 306 is logic that can be dynamically configured to perform any of a variety of filtering types on the input data set. Performing one type of filtering on the input data set may be referred to herein as performing a filtering task. The various types of filtering include at least one type of texture filtering and at least one type of pixel filtering. Texture filtering types include, but are not limited to, bilinear filtering, trilinear filtering, anisotropic filtering, and percentage asymptotic filtering (PCF). Filter set 306 can support any combination of these types of texture filtering. Pixel filtering types include, but are not limited to, downsampling, upsampling, and box filtering to achieve anti-aliasing, such as MSAA. Filter set 306 can support any combination of these types of pixel filtering.
[0053] Filter group 306 may include one or more filter blocks 314, each of which can be configured to perform any of a variety of types of filtering. Each filter block 314 may include a plurality of fixed arithmetic components, which can be individually enabled or disabled to enable the filter block 314 to perform one of the supported types of filtering. For example, each filter block 314 may include a basic filter (e.g., a 2x2 filter) that can generate a weighted sum of the input set; and one or more other arithmetic components that can be selectively enabled to perform more complex filtering. As described in more detail below, the basic filter may include One multiplication unit (of which) (The input value is an integer greater than one). Each of the multiplication units is configured to multiply the input value by a filter weight, followed by multiple adder units that form an adder tree, which generates the sum of the outputs of the multiplication units. Examples of other arithmetic units include, but are not limited to: comparison units that compare two values, minimum value units that calculate the minimum of a set of values, maximum value units that calculate the maximum of a set of values, scaling / offset units that scale values or apply offsets to values, addition units that produce the sum of two values, subtraction units that produce the difference of two values, and shift units that shift the input value by a specific value. The following is in conjunction with... Figure 5 An exemplary implementation of filter block 314 is described.
[0054] As described above, in bilinear filtering, the four texels closest to the relevant location in the texture (e.g., mapped texture coordinates) are read and combined by a weighted average based on distance to produce the texture color of the fragment. Therefore, bilinear filtering can be performed by a basic 2×2 filter by providing the desired texels as input data and using filter weights representing the distance between the texels and the relevant location in the texture. Similarly, as described above, when pixels are oversampled (e.g., more than one sample per pixel—e.g., multiple subsamples exist), the color of a pixel can be determined by combining subsamples (color values) generated by the fragment / pixel shader that are associated with that particular pixel (using a reconstruction filter). A common reconstruction filter is a single-pixel-wide box filter, which essentially generates the average of all subsamples corresponding to (or within) the pixel. With four subsamples per pixel, box filtering can be performed by a combination of a basic 2×2 filter and a shifting component as follows: provide the subsamples as input data to the basic 2×2 filter, then use one of the filter weights, and then divide the output by 4 (this can be done via a shifting operation).
[0055] As described above, each filter block 314 may only be able to perform a certain number of arithmetic operations and / or combinations of arithmetic operations at a time. These limitations can be imposed by the hardware used to implement the filter block 314. However, some filtering tasks may require more than this number of arithmetic operations and / or combinations of arithmetic operations. For example, a filter block may include hardware capable of calculating a weighted sum of four inputs, but the filtering task may require calculating a weighted sum of a first set of inputs and a weighted sum of a second set of inputs. Therefore, the same filter block can be used multiple times to implement or perform more complex filtering tasks. For example, a filter block may first be used to calculate a weighted sum of a first set of inputs and then to calculate a weighted sum of a second set of inputs. Each time the filter block is combined with the same task, a round, or hardware round, is used, referred to herein as a filter block round. Thus, for each round of filter block 314, filter block 314 receives input data from one of the data buffers (ALU-side buffer 302 or memory-side buffer 304) and performs one or more arithmetic operations on the received data. In some cases, each round may take one cycle (e.g., a clock cycle) to complete. However, in other cases, a round may take more than one cycle (e.g., a clock cycle).
[0056] For example, trilinear filtering interpolates between the results of two different bilinear filtering operations, i.e., combining the results of bilinear filtering performed on two mipmaps closest to the location of interest (e.g., the location of a relevant pixel or sample). When filter block 314 can perform one bilinear filtering operation at a time, then during the first round of filter block 314, filter block 314 can be configured to perform bilinear filtering on a first mipmap, and during the second round of filter block 314, filter block 314 can be configured to perform bilinear filtering on a second mipmap, interpolating between the outputs of the two bilinear filtering operations. It will be apparent to those skilled in the art that these are examples of how different filtering techniques or methods can be implemented in multiple rounds, and the number of rounds used to implement a filtering method or technique depends on the components (e.g., basic filters and arithmetic components) and capabilities of each filter block.
[0057] In some cases, the arithmetic component of each filter block 314 can be configured such that each filter group can perform at least bilinear filtering, trilinear filtering, anisotropic filtering, PCF filtering, and box filtering to implement MSAA, wherein: Bilinear filtering can be executed at full speed (e.g., one bilinear filter output can be generated per clock cycle). Trilinear filtering can be performed at half speed (e.g., a trilinear filter output can be generated every two clock cycles). Anisotropic filtering can be performed at a speed of 1 / x, where x is the number of samples (e.g., an anisotropic filter with sixteen samples runs 16 times slower than bilinear filtering); and Box filtering with 4 samples per pixel implementing MSAA can be executed at full speed (e.g., one MSAA box filter output can be generated per clock cycle).
[0058] Generally, the more filter blocks 314 there are, the more filtering tasks the filter group can perform in parallel. The number of filter blocks 314 can be selected to achieve the desired performance level. In some cases, the number of filter blocks 314 can be selected to provide a similar performance level (e.g., the same peak filter rate) to the texture unit 114 and pixel backend 116 that the input / output filter unit 202 is replacing. For example, if the texture unit 114 has a peak rate of 4 outputs per clock cycle, and the pixel backend 116 has a peak rate of 4 pixels (color values) per clock cycle, and each filter block 314 has a peak rate of 1 texture or 1 pixel filter output per cycle, then the filter group 306 can include eight filter blocks 314 such that in any given cycle, four of the filter blocks 314 can be used to perform texture filtering tasks, and four of the filter blocks 314 can be used to perform pixel filtering tasks. However, in other cases, the number of filter blocks 314 can be selected to provide a peak filter rate that is less than (e.g., half the rate) of the peak filter rate provided by the texture unit 114 and pixel backend 116. For example, if texture unit 114 has a peak rate of four outputs per clock cycle and pixel back-end 116 has a peak rate of four output pixels (color values) per clock cycle, then filter group 206 may consist of only four filter blocks 314. This may reduce performance in a few cases, but can have a small impact on overall performance, while resulting in area and / or power savings.
[0059] The weight buffer and control logic unit 308 includes a weight buffer for storing filter weights for the filtering task and control logic for controlling the filter block 314 to perform the filtering task. As described in more detail below, the filter weights stored in the weight buffer can be generated by, for example... Figure 3The weights are generated by the weight generator 312, or they may be loaded from memory. In some cases, after weights have been used in a filtering task, they may not be immediately removed from the weight buffer to allow the filter weights to be reused in subsequent filtering tasks. In other words, in some cases, filter weights may be cached. In some cases, the filtering task may require one or more additional parameters. For example, if a shift is to be performed as part of a filtering task, the shift amount may be a parameter provided to filter block 314. In these cases, additional parameters may also be stored in the weight buffer.
[0060] The control logic is configured to cause filter block 314 to perform filtering tasks. Each filtering task is defined or includes input data (stored in one of the data buffers (ALU-side buffer 302 or memory-side buffer 304)), filter weights (stored in a weight buffer and control logic unit 308), and a filtering type (which is one of a variety of supported filtering types). As mentioned above, in some cases, the filtering task may also include additional parameters (which may also be stored in the weight buffer). The control logic is configured to provide appropriate input data from the appropriate data buffer (ALU-side buffer 302 or memory-side buffer 304) and appropriate filter weights (and optionally other parameters) from the weight buffer to filter block 314, and cause filter block 314 to perform a specific type of filtering. The control logic can be configured to cause filter block 314 to perform a specific type of filtering, for example, by enabling and disabling specific combinations of arithmetic components therein. The control logic can also be configured to cause filter block 314 to perform a specific type of filtering by sending one or more control signals to the filter block.
[0061] In some cases, the input / output filter unit 202 may also include a texture address generator 310. As described above, texture filtering generally involves obtaining or reading one or more texels of the texture near a location of interest in the texture, and performing filtering on the obtained texels. The texture address generator 310 may be configured to generate addresses of relevant texels for the location of interest in the texture. In some cases, the texture address generator 310 may be configured to receive information identifying the location (e.g., x, y coordinates) of a relevant pixel or fragment in rendering space, and to map the received location (e.g., x, y coordinates) to a set of u, v coordinates, which may be referred to as mapped texture coordinates. The mapped texture coordinates identify a specific location in the texture, which may be referred to as a relevant location or location of interest in the texture. In other cases, the texture address generator 310 may simply receive a set of u, v coordinates defining the location of interest as input. For example, each vertex can be associated with a set of u, v coordinates, and when a primitive is rasterized (e.g., converted into one or more fragments), the u, v coordinates of the primitive vertices can be interpolated to generate a set of u, v coordinates for that fragment.
[0062] In either case, the u, v coordinates of the location of interest in the texture are defined to identify the relevant texels and their addresses (e.g., their u, v coordinates). The number of relevant texels at the location of interest can be based on the specific type of texture filtering to be performed. Therefore, in addition to receiving information identifying the location of interest (or receiving information from which the location of interest can be generated), the texture address generator 310 can also be configured to receive information identifying the type of texture filtering to be performed. For example, for bilinear filtering, only the four texels closest to the location of interest in the texture are obtained. However, for trilinear filtering, the texels that form the two mipmaps closest to the point of interest are obtained.
[0063] The texture addresses (e.g., u, v coordinates) generated by texture address generator 310 can then be used to obtain or read the relevant texels from memory. The generated texture addresses (e.g., u, v coordinates) can also be provided to a weight generator (e.g., weight generator 312) to generate appropriate filter weights for those texels. In other cases, input / output filter unit 202 may not include a texture address generator, and the texture addresses may be generated by another component or unit, such as, but not limited to, ALU cluster 102.
[0064] In some cases, the input / output filter unit 202 may also include a weight generator 312. The weight generator 312 is configured to generate filter weights for the filtering task. The number and / or calculation of the filter weights may be based on the type of filtering to be performed. Therefore, the weight generator may be configured to receive information identifying the filtering method or the type of filtering to be performed. For texture filtering, the weight generator 312 may be configured to also receive information identifying the positions of relevant texels in the texture (e.g., texture addresses generated by the texture address generator 310), and calculate the weights of the identified texture filtering method based on said information. For example, for bilinear or trilinear filtering, the weight generator 312 may be configured to generate filter weights based on the distance between relevant texels and positions of interest in the texture. For example, as... Figure 5 As shown, if the texels closest to the point of interest x (texels in the minimum mipmap) are c0, c1, c2, and c3, then the result of bilinear filtering applied to these texels can be expressed as c = (1-t). (1-s) c0 + (1-t) s c1 + t (1-s) c2 + t s c3. Therefore, the filter weights for texels c0, c1, c2, and c3 are w0, w1, w2, and w3, respectively, where w0 = (1-t). (1-s), w1=(1-t) s、w2=t (1-s) and w3 = t It is obvious that this is merely an example, and those skilled in the art will understand how filter weights can be generated for different types of filtering.
[0065] In some cases, weight generator 312 may only be able to generate filter weights for texture filtering. In other cases, weight generator 312 may be able to generate filter weights for one or more other types of filtering (e.g., fixed-weight filtering). A fixed-weight filter type is a filter that always uses the same weights. Examples of fixed-weight filters include, but are not limited to, box filtering, Gaussian filtering, and tent filtering (which may also be called triangle filtering). In contrast, bilinear filtering uses different filter weights depending on the data being filtered, therefore bilinear filtering is not a type of fixed-weight filtering. The filter types for which weight generator 312 can generate filter weights may be only a subset of the supported filter types (i.e., fewer than all supported filter types). In some cases, the filter types for which weight generator 312 can generate filter weights may be hard-coded or dynamically configurable.
[0066] The filter weights generated by weight generator 312 can be output and stored in weight buffer and control logic unit 308. In other cases, input / output filter unit 202 may not include a weight generator, and the filter weights may be generated by another component or unit, such as, but not limited to, ALU cluster 102, or they may be retrieved from memory.
[0067] exist Figure 3 In this configuration, a data path 316 exists between filter group 306 and ALU-side buffer 302, and a data path 318 exists between filter group 306 and memory-side buffer 304, allowing filter group 306 to write data to and read data from the data buffer (ALU-side buffer 302 or memory-side buffer 304). In some cases, the data paths 316 and 318 between filter group 306 and the data buffer (ALU-side buffer 302 or memory-side buffer 304) can be wide enough to allow all filter blocks 314 to simultaneously read and / or write data to the same data buffer (ALU-side buffer 302 or memory-side buffer 304), thus allowing all filter blocks 314 to operate in parallel without stalling. The minimum size of data paths 316 and 318 that allow all filter blocks 314 to simultaneously read and / or write data to the same data buffer (ALU-side buffer 302 or memory-side buffer 304) can be based on the format of the input and output data and the number of filter blocks 314. In some cases, each texel can be in RGBA format, which includes the values of each of the red, green, blue, and opacity channels. When each channel value is a 32-bit floating-point value, each texel will be 128 bits. In a texture filtering task where four texels can be read and processed, and eight texture filtering tasks can be executed in parallel, the data path can be at least 128 x 4 x 8 = 4096 bits wide.
[0068] exist Figure 3 In this context, data paths 320 and 322 also exist between the data buffer (ALU-side buffer 302 or memory-side buffer 304) and the ALU cluster 102 and memory. In some cases, these data paths 320 and 322 may be narrower than the data paths 316 and 318 between the data buffer (ALU-side buffer 302 or memory-side buffer 304) and the filter group 306. This is because, due to the reuse of data between filtering tasks—for example, reusing neighboring values when running a sliding window filter—less data may be transferred between the ALU cluster 102 and the ALU-side buffer 302, and between memory and the memory-side buffer 304, compared to between the data buffer (ALU-side buffer 302 or memory-side buffer 304) and the filter group 306. For example, each bilinear texture filtering task of a segment may read four texels from the memory-side buffer 304; however, bilinear texture filtering tasks of adjacent segments may use some of the same texels, so it may not be necessary to read four texels from memory for each bilinear texture filtering task. In other words, although eight texels can be read from the memory-side buffer 304 to perform two bilinear texture filtering tasks, fewer than eight texels can be read from memory for both tasks because the two tasks can use some of the same texels. Therefore, to perform two bilinear texture filtering tasks, less data needs to be read from memory than from the memory-side buffer.
[0069] Now for reference Figure 5 , showed Figure 3 An exemplary implementation of filter block 314 is provided. The exemplary filter block 314 is implemented as a pipeline of arithmetic components. The pipeline includes five stages numbered 0 to 4. The first pipeline stage (stage 0), which may be referred to as the comparison stage, includes four comparison components 5020, 5021, 5022, and 5023. Comparison units 5020, 5021, 5022, and 5023 are configured to receive input data value D from one of the data buffers (ALU-side buffer 302 or memory-side buffer 304). i and receives the reference value REF from the weight buffer and control logic unit 308. i and input data value D i It compares the value to a reference value and outputs either "0" or "1" based on that comparison. For example, in some cases, if the data value D... i Greater than the reference value REF iIf the comparison components 5020, 5021, 5022, and 5023 can output "1", otherwise they will output "0". However, it will be apparent to those skilled in the art that this is merely an example, and in other examples, if the data value D... i Greater than the reference value REF i If the condition is met, comparison units 5020, 5021, 5022, and 5023 can output "0"; otherwise, they output "1". Comparison levels can be used to implement PCF filtering. Specifically, in PCF filtering, the input data is first compared to a reference value before filtering. In PCF filtering, each input data value is compared to the same reference value (e.g., REF0 = REF1 = REF2 = REF3), but in other types of filtering, different input data values can be compared to different reference values.
[0070] The second pipeline stage (stage 1) can be referred to as the multiplication stage or multiplication stage, and includes four multiplication units: 5040, 5041, 5042, and 5043. Multiplication units 5040, 5041, 5042, and 5043 are configured to receive input data value D. (If the comparison stage or corresponding comparison unit is disabled) or the output of the corresponding comparison units 5020, 5021, 5022 and 5023, and receive the filter weight W from the weight buffer. And generate and output input DW The product. For example, in multiplication unit 504 Receive raw input data value D and weight W In the case of multiplication unit 504 Calculate and output D W Enter DW The product of these can be called weighted data points.
[0071] The third pipeline stage (stage 2), which can be referred to as the first adder stage, includes two adder units 5060 and 5061. Each adder unit 5060 and 5061 receives weighted data points DW. The two adder components in the third pipeline stage receive weighted data points DW0 and DW1 generated by the first multiplier component 5040 and the second multiplier component 5041 of the second pipeline stage, and calculate and output DW0 + DW1; and the second adder component 5061 of the third pipeline stage receives weighted data points DW2 and DW3 generated by the third multiplier component 5042 and the fourth multiplier component 5043, and calculates and outputs DW2 + DW3.
[0072] The fourth pipeline stage (stage 3), which can be referred to as the second adder stage, includes a single adder unit 508 that receives the outputs of adder units 5060 and 5061 from the third pipeline stage and calculates and outputs their sum. It can be seen that the third and fourth pipeline adder stages together form an adder tree, which produces the sum of the outputs of multiplication units 5040, 5041, 5042, and 5043 (i.e., the sum of weighted data points—DW0 + DW1 + DW2 + DW3). It can also be seen that the second, third, and fourth pipeline stages (stage 1, stage 2, and stage 3) together form a filter unit that calculates the weighted sum of four values. The second, third, and fourth pipeline stages (stage 1, stage 2, and stage 3) can alternatively be described as implementing a convolution engine or convolution operation between the four input data points and the four filter weights.
[0073] The fifth pipeline stage (stage 4), which may be referred to as the scaling / offset stage, includes a scaling / offset component 510 configured to receive the output of the fourth pipeline stage (stage 3) and apply scaling or offset to the received value to generate the filtered output F1. In some cases, the scaling or offset applied to the received value by the scaling / offset component 510 can be configurable. For example, the scaling or offset value can be stored in memory (e.g., in a weight buffer and control logic unit 308) and provided to the filter block 314 as part of the control data. The same offset or scaling can be used for a specific texture-filter type combination. For example, a scaling of 2 can be used for any bilinear filtering task related to a specific texture.
[0074] The filtered output F1 can be stored in one of the data buffers (ALU-side buffer 302 or memory-side buffer 304). In some cases, the filtered output F1 can be provided as an input to the filter block 314 in the next cycle (e.g., the next clock cycle). For example, a feedback path can exist between the pipeline output and the pipeline input. For example, a feedback path (not shown) can exist between the output of the scaling / offset unit 510 and, for example, the input D0 of the first comparator unit 5020. Then, if the filtered output F1 will be used in the next round of the filter block 314, the filtered output F1 is provided to the first comparator unit 5020 via the feedback path. This eliminates the need to write the filtered output F1 to memory and then read F1 from memory for the next round.
[0075] The weight buffer and control logic unit 308 are configured to control the filter block 314 to execute or implement a specific filter type. This may include arithmetic components (e.g., comparison components 5020, 502) that selectively enable and / or disable the filter block 314. 1、 5022、 5023, multiplication components 5040, 504 1、 504 2、 5043, adder components 5060, 5061, 508, scaling / offset component 510). The weight buffer and control logic unit 308 can enable or disable an entire level (e.g., all comparison components) and / or enable or disable individual arithmetic components (e.g., a single comparison component). For example, to enable filter block 314 to perform a bilinear filtering task or a box filtering task, the weight buffer and control logic unit 308 can be configured to disable comparison levels (e.g., all comparison components 5020, 5021, 5022, 5023) and enable all other levels (e.g., all other components, such as multiplication components 5040, 5041, 5042, 5043, adder components 5060, 5061, 508, scaling / offset component 510). The difference between a bilinear filtering task and a box filtering task is that for bilinear filtering, each weight (W0, W1, W2, W3) can be different, while for box filtering, all weights are the same (e.g., 1). In another example, in order for filter block 314 to implement PCF filtering, weight buffer and control logic unit 308 can be configured to enable all arithmetic components (e.g., comparison components 5020, 5021, 5022, 5023, multiplication components 5040, 5041, 5042, 5043, adder components 5060, 5061, 508, scaling / offset component 510).
[0076] In some cases, the weight buffer and control logic unit 308 can communicate with each arithmetic component of the filter block 314 (e.g., comparison components 5020, 5021, 5022, 5023, multiplication components 5040, 5041, 5042, 5043, adder components 5060, 5061, 508, scaling / offset component 510), and any arithmetic component can be enabled or disabled by sending an enable or disable signal to that arithmetic component respectively. In other cases, the filter block 314 may include an internal control unit (not shown) that communicates with the weight buffer and control logic unit 308 and each arithmetic component, and the weight buffer and control logic unit 308 is configured to send control signals to the internal control unit indicating which arithmetic components to enable and which to disable, and the internal control unit enables and disables the arithmetic components accordingly. As described above, in some cases, the filtering task can be performed on multiple rounds of the filter block 314. In these cases, the weight buffer and control logic unit 308 can be configured to treat each round as a separate control item. Specifically, the weight buffer and control logic unit 308 can be configured to generate a separate set of control signals for each round.
[0077] In some cases, the weight buffer and control logic unit 308 may receive information (e.g., status information) and / or one or more control signals (e.g., instructions), which cause the weight buffer and control logic unit 308 to cause the filter group 306 to perform a specific filtering task. The information or control signals that cause the weight buffer and control logic unit 308 to cause the filter group 306 to perform a specific filtering task may be generated, for example, by the ALU cluster 102. For example, in some cases, the ALU cluster 102 may be configured to, as part of performing a pixel shader task, issue instructions or instruction sets to the input / output filter unit 202, which cause the input / output filter unit to perform a specific texture filtering task on fragments / pixels and return the result of the texture filtering task to the ALU cluster 102; and / or when the ALU cluster 102 completes the pixel shader task, the ALU cluster may be configured to issue instructions or instruction sets, which cause the input / output filter unit 202 to perform a pixel filtering task on the output of the pixel shader task. In other cases, instead of issuing instructions to the input / output filter unit 202 to cause the filtering task to be executed, the ALU cluster 102 can be configured to store state data along with the data to be filtered, which, when read by the weight buffer and control logic unit 308, causes the weight buffer and control logic unit 308 to cause the filter group 306 to perform filtering operations on the stored data.
[0078] Although Figure 5 The exemplary filter block 314 is configured to implement a 4-input data x 4 weighted filter, but those skilled in the art will appreciate that this is merely an example and other exemplary filter blocks can implement filters of other sizes (e.g., a 2-input data x 2 weighted filter or an 8-input data x 8 weighted filter). However, it should be noted that using a 4-input data x 4 weighted filter as the base filter allows for highly efficient bilinear filtering of four texels and processing of RGBA pixels (e.g., pixels composed of four values), which are respectively... Figure 1 The texture unit 114 and the pixel backend 116 perform the same task. It should also be noted that larger filters can be implemented via multiple rounds of filter block 314.
[0079] It is obvious to those skilled in the art that Figure 5The combinations and arrangements of arithmetic components shown are merely examples. In other examples, filter block 314 may include additional, different arithmetic components and / or different arrangements of arithmetic components. For example, in some cases, filter block 314 may also include a mixing / hybridizing component (not shown) configured to mix or hybridize the output of the scaling / offset component with other data, such as data from previous rounds of filter block 314.
[0080] In some cases, to make the input / output filter unit 202 more useful and / or more versatile, the filter block 314 may be able to perform additional operations, which may not typically be performed by... Figure 1 The operations performed by the texture unit 114 or pixel backend 116 are similar to those performed by them. Specifically, in addition to vertex and pixel processing, the filter block 314 can also be capable of performing generalized computational tasks or functions, such as, but not limited to, image processing, implementing camera ISP algorithms, and neural network processing. Specifically, since neural network operations are very similar to filter operations (e.g., they typically involve performing convolution operations, which involve computing a weighted sum of a set of inputs), the filter block can also be configured to perform neural network operations. In addition to increasing the usefulness and / or versatility of the input / output filter unit 202, this also eliminates the need for a separate neural network accelerator in the GPU. Other similar functions and operations that may be suitable for performance via the input / output filter unit 202 could be other filters using convolution or weighted sums. Such operations include, but are not limited to, color space conversion, Gaussian filters, edge-aware filters / scalers (including advanced edge-aware MSAA filters).
[0081] In some cases, filter block 314 can be configured to perform blending. Specifically, trilinear filtering, and by extending anisotropic filtering, is operationally similar to sampling multiple textures and blending these layers together. Specifically, trilinear filtering takes the output of bilinear filtering performed on two different mipmap levels and blends these results together using weighted blending (e.g., (1-a) x Csource + ax Cdest), a common blending pattern implemented by the ALU cluster. Therefore, filter block 314 can be configured to perform simple texture blending and combination, such as that used in a graphical user interface (GUI), where the composition can include blending multiple layers without complex arithmetic. This increases the complexity of filter block 314, but allows blending as a backend operation when data is written from the ALU cluster to memory (e.g., a tiling buffer). This allows the GPU to enter a much lower power mode, where the ALU cluster does not need to be enabled when blending combinations of several blend surfaces.
[0082] In some cases, filter block 314 can also be configured to perform format conversion. For example, filter block 314 may be able to convert RGB colors (which have values for the red channel R, green channel G, and blue channel B) to YUB (which stores luminance (brightness) as Y values and color (chrominance) as U and V values) or vice versa by using a hard-coded set of weights.
[0083] In the past, the data (texels) input to texture unit 114 was typically in a different format than the data (pixels / samples) output by the ALU cluster, making it difficult to create a universal unit capable of handling both data formats. However, both types of data are now typically in 16-bit floating-point format. Therefore, in some cases, filter group 306 and its filter block 314 can be configured to support 16-bit floating-point operations. In other cases, however, filter group 306 and its filter block 314 can be configured to support multiple data or numeric formats. The multiple data formats supported by filter group 306 and its filter block may include 32-bit floating-point formats (e.g., R16G16B16A16_FLOAT), 8-bit fixed-point or integer formats (e.g., RGBA8888), and 10-bit fixed-point or integer formats (e.g., R10G10B10A2) and / or a smaller format, such as, but not limited to, 444, 565, and 5551. These smaller formats can be supported by decompressing them into wider formats to avoid overcomplicating format support within the filter bank 306 itself. In some cases (e.g., if neural network operations are supported by the filter bank 306), it may also be beneficial to support dual-rate 8-bit fixed-point or integer formats that allow the filter bank 306 to perform 16-bit operations or two 8-bit operations.
[0084] Now for reference Figure 6 It shows the control Figure 3 An exemplary method 600 for an input / output filter unit 202, which may be implemented by control logic of a weight buffer and a control logic unit 308. Method 600 begins at block 602, where the control logic receives information identifying a filtering task and / or control signals. The information identifying the filtering task may include a set of data stored in one of the data buffers (ALU-side buffer 302 or memory-side buffer 304), filter weights stored in the weight buffer and the control logic unit 308, and information about one type of filtering among multiple types. The multiple filtering types include one or more types of texture filtering and one or more types of pixel filtering. As mentioned above, in some cases, the information identifying the filtering task may also include additional parameters, such as, but not limited to, shift amounts or comparison values.
[0085] As described above, the information or control signal identifying a specific filtering task can be generated, for example, by the ALU cluster 102. For instance, in some cases, the ALU cluster 102 can be configured to issue instructions or instruction sets to the control logic, as part of executing a pixel shader task, identifying a specific texture filtering task to be performed on a fragment / pixel; and / or when the ALU cluster 102 completes a pixel shader task, the ALU cluster can be configured to issue instructions or instruction sets identifying a pixel filtering task to be performed on the output of the pixel shader task. In other cases, instead of sending information or control signals to the control logic, the ALU cluster 102 can be configured to store state data along with the data to be filtered, which identifies the filtering task to be performed on the stored data when read by the control logic. Once the control logic has received the information identifying the filtering task, method 600 proceeds to block 604.
[0086] At block 604, the control logic causes filter group 306 to perform the identified filtering task. Specifically, the control logic causes filter group 306 to perform filtering of the identified type on the identified data set using the identified set of weights.
[0087] The control logic can be configured to provide the identified data set from the appropriate data buffer (ALU-side buffer 302 or memory-side buffer 304) and the identified filter weights (and optional other parameters) from the weight buffer and control logic unit 308 to the filter group 306, and to cause the filter group 306 to perform the identified type of filtering on the received data set using the received weight set. The control logic can be configured to cause the filter group 306 to perform a specific type of filtering by sending one or more control signals to the filter group 306. When the filter group 306 includes one or more filter blocks 314 and each filter block has multiple arithmetic components, the control logic can be configured to cause the filter blocks 314 to perform a specific type of filtering, for example, by selectively enabling and / or disabling specific combinations of the arithmetic components of the filter block 314. As described above, when the arithmetic components are divided into multiple levels, the control logic can be able to enable or disable an entire level (e.g., all comparison components) and / or enable or disable individual arithmetic components (e.g., a single comparison component).
[0088] As described above, some filtering tasks may require multiple rounds of filter groups 306 to complete. In these cases, the control logic can be configured to treat each round as a separate control item. Specifically, the control logic can be configured to generate a separate set of control signals for each round. For example, the control logic can be configured to cause one of the filter blocks 314 to perform a first part of the filtering of that type in the first round of filter block 314, and a second part of the filtering of that type in the second round of filter block 314. Once the control logic has caused filter group 306 to perform the identified filtering task, method 600 can end at 608 or method 600 can proceed to block 606.
[0089] At box 606, determine if there is another filtering task to perform. If there is another filtering task to perform, method 600 returns to box 602. If there is no other filtering task to perform, method 600 ends at 608.
[0090] although Figure 6 The description outlines controlling input / output filter units to perform a single filtering task. However, when the filter group of the input / output filter unit comprises multiple filter blocks, and each filter block can perform a filtering task, multiple filtering tasks can be executed in parallel by the input / output filter unit. In these cases, each filtering task can be executed... Figure 6 Method 600.
[0091] Figure 7 A computer system in which the input / output filter unit 202 described herein may be implemented is illustrated. The computer system includes a CPU 702, a GPU 704, memory 706, and other devices 714, such as a display 716, speakers 718, and a camera 720. A processing block 710 (which may be the input / output filter unit 202 described herein) is implemented on the GPU 704. In other examples, the processing block 710 may be implemented on the CPU 702. Components of the computer system may communicate with each other via a communication bus 722.
[0092] Figure 1 , Figure 2 and Figure 3 The input / output filter unit and graphics processing unit are shown as comprising multiple functional blocks or units. This is merely illustrative and not intended to define a strict division between different logical elements of such entities. Each functional block or unit can be provided in any suitable manner. It should be understood that the intermediate values described herein as formed by the blocks or units do not need to be physically generated by the input / output filter unit or graphics processing unit at any point, and may simply represent logical values that conveniently describe the processing performed by the input / output filter unit or graphics processing unit between its inputs and outputs.
[0093] The input / output filter units and / or graphics processing units described herein may be contained in hardware on an integrated circuit. The input / output filter units and / or graphics processing units described herein may be configured to perform any of the methods described herein. Generally, any of the functions, methods, techniques, or components described above may be implemented in software, firmware, hardware (e.g., a fixed logic circuit system), or any combination thereof. The terms “module,” “function,” “component,” “element,” “unit,” “block,” and “logic” may be used herein to generally denote software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, function, component, element, unit, block, or logic represents program code that, when executed on a processor, performs a specified task. The algorithms and methods described herein may be executed by one or more processors executing code that causes the processor to execute the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disk, flash memory, hard disk storage, and other memory devices that may use magnetic, optical, and other techniques to store instructions or other data and may be accessible by a machine.
[0094] As used herein, the terms computer program code and computer-readable instructions refer to any kind of executable code for a processor, comprising code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, bytecode, code defining integrated circuits (e.g., hardware description languages or netlists), and code expressed in programming languages such as C, Java, or OpenCL. Executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, compiled, or run in a virtual machine or other software environment, causes the processor of a computer system that supports the executable code to perform tasks specified by said code.
[0095] A processor, computer, or computer system can be any kind of device, machine, or special-purpose circuit, or a collection or part thereof, that has the processing power to execute instructions. A processor can be any kind of general-purpose or special-purpose processor, such as a CPU, GPU, system-on-a-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), etc. A computer or computer system may include one or more processors.
[0096] This invention also intends to cover software defining the configuration of hardware as described herein, such as hardware description language (HDL) software, for designing integrated circuits or for configuring programmable chips to perform desired functions. That is, a computer-readable storage medium on which computer-readable program code in the form of an integrated circuit definition dataset is encoded may be provided, which, when processed (i.e., run) in an integrated circuit manufacturing system, configures the system to manufacture an input / output filter unit or graphics processing unit configured to perform any of the methods described herein, or to manufacture a processor including any of the devices described herein. The integrated circuit definition dataset may, for example, be an integrated circuit description.
[0097] Therefore, a method for manufacturing input / output filter units or graphics processing units as described herein can be provided in an integrated circuit manufacturing system. Furthermore, an integrated circuit definition dataset can be provided, which, when processed in an integrated circuit manufacturing system, enables the method for manufacturing input / output filter units or graphics processing units to be executed.
[0098] Integrated circuit definition datasets can be in the form of computer code, such as netlists, code for configuring programmable chips, or hardware description languages suitable for manufacturing at any level in integrated circuits, including register-transfer level (RTL) code, high-level circuit representations (such as Verilog or VHDL), and low-level circuit representations (such as OASIS(RTM) and GDSII). Higher-level representations (such as RTL) that logically define hardware suitable for manufacturing in integrated circuits can be processed on a computer system configured to generate manufacturing definitions of integrated circuits within the context of a software environment that includes definitions of circuit elements and rules for combining these elements to generate the manufacturing definition of the integrated circuit defined by that representation. As is typically the case where software executes at a computer system to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to configure the computer system to generate the manufacturing definition of the integrated circuit, executing code that defines the integrated circuit to generate the manufacturing definition of the integrated circuit.
[0099] Now refer to Figure 8 This describes an example of processing integrated circuit definition datasets at an integrated circuit manufacturing system in order to configure the system for manufacturing input / output filter units or graphics processing units.
[0100] Figure 8An example of an integrated circuit (IC) manufacturing system 802 is shown, configured to manufacture input / output filter units and / or graphics processing units as described in any of the examples herein. Specifically, the IC manufacturing system 802 includes a layout processing system 804 and an integrated circuit generation system 806. The IC manufacturing system 802 is configured to receive an IC definition dataset (e.g., defining input / output filter units or graphics processing units as described in any of the examples herein), process the IC definition dataset, and generate an IC (e.g., containing input / output filter units or graphics processing units as described in any of the examples herein) based on the IC definition dataset. The processing of the IC definition dataset configures the IC manufacturing system 802 to manufacture integrated circuits containing memory cell allocators or graphics processing units as described in any of the examples herein.
[0101] The layout processing system 804 is configured to receive and process an IC definition dataset to determine a circuit layout. Methods for determining a circuit layout based on an IC definition dataset are known in the art and may involve, for example, synthesizing RTL code to determine the gate-level representation of the circuit to be generated, for example, in relation to logic components (e.g., NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining the location information of the logic components, the circuit layout can be determined based on the gate-level representation of the circuit. This can be done automatically or with user intervention to optimize the circuit layout. When the layout processing system 804 has determined the circuit layout, it can output the circuit layout definition to the IC generation system 806. The circuit layout definition may be, for example, a circuit layout description.
[0102] IC generation system 806 generates ICs according to circuit layout definitions as known in the art. For example, IC generation system 806 can implement a semiconductor device manufacturing process for generating ICs, which may include a multi-step sequence of photolithography and chemical processing steps, during which electronic circuits are progressively formed on a wafer made of semiconductor material. The circuit layout definition may be in the form of a mask, which can be used in the photolithography process to generate ICs according to the circuit definition. Alternatively, the circuit layout definition provided to IC generation system 806 may be in the form of computer-readable code, which IC generation system 806 can use to form an appropriate mask for generating ICs.
[0103] The different processes performed by the IC manufacturing system 802 can all be implemented in one location, for example, by one party. Alternatively, the IC manufacturing system 802 can be a distributed system, such that some processes can be performed at different locations and by different parties. For example, some of the following stages can be performed at different locations and / or by different parties: (i) synthesizing RTL codes representing the IC definition dataset to form a gate-level representation of the circuit to be generated; (ii) generating a circuit layout based on the gate-level representation; (iii) forming a mask based on the circuit layout; and (iv) fabricating the integrated circuit using the mask.
[0104] In other examples, processing of the integrated circuit definition dataset at an integrated circuit manufacturing system can configure the system to manufacture input / output filter units or graphics processing units where the IC definition dataset is not processed to determine circuit layout. For example, the integrated circuit definition dataset can define the configuration of a reconfigurable processor, such as an FPGA, and processing of the dataset can configure the IC manufacturing system (e.g., by loading the configuration data into the FPGA) to generate a reconfigurable processor with the defined configuration.
[0105] In some embodiments, when processed in an integrated circuit manufacturing system, an integrated circuit manufacturing definition dataset can enable the integrated circuit manufacturing system to generate devices as described herein. For example, using an integrated circuit manufacturing definition dataset, with reference to the above... Figure 8 The configuration of the integrated circuit manufacturing system described herein can produce the equipment as described in this article.
[0106] In some examples, an integrated circuit definition dataset may include software running on hardware defined at the dataset, or software running in combination with hardware defined at the dataset. Figure 8 In the example shown, the IC generation system can also be further configured by the integrated circuit definition dataset to load firmware onto the integrated circuit according to the program code defined in the integrated circuit definition dataset during the manufacturing of the integrated circuit, or otherwise provide the integrated circuit with program code to be used with the integrated circuit.
[0107] Compared to known implementations, the implementation of the concepts set forth in this application in devices, apparatuses, modules, and / or systems (and in the methods implemented herein) can lead to performance improvements. Performance improvements may include one or more of increased computational performance, reduced latency, increased throughput, and / or reduced power consumption. During the manufacture of such devices, apparatuses, modules, and systems (e.g., in integrated circuits), trade-offs can be made between performance improvements and physical implementations, thereby improving manufacturing methods. For example, a trade-off can be made between performance improvements and layout area to match the performance of known implementations but using less silicon. This can be accomplished, for example, by reusing functional blocks serially or sharing functional blocks among elements of a device, apparatus, module, and / or system. Conversely, the concepts set forth in this application that lead to improvements in the physical implementation of devices, apparatuses, modules, and systems (such as reduced silicon area) can be traded off for performance improvements. This can be accomplished, for example, by manufacturing multiple instances of a module within a predefined area budget.
[0108] The applicant has independently disclosed each individual feature described herein, as well as any combination of two or more such features, to the extent that such features or combinations can be implemented based on the specification as a whole, in accordance with the common knowledge of those skilled in the art, regardless of whether such features or combinations of features solve any problem disclosed herein. In view of the foregoing description, those skilled in the art will understand that various modifications can be made within the scope of this invention.
Claims
1. An input / output filter unit for use in a graphics processing unit, the graphics processing unit including a memory and one or more clusters of arithmetic logic units configured to perform shading tasks, the input / output filter unit comprising: A first buffer is configured to store data received from and output to the one or more arithmetic logic unit clusters. A second buffer is configured to store data received from the memory and data output to the memory; A weight buffer, configured to store filter weights; A filter group, which can be configured to perform any of a variety of types of filtering on an input data set, including one or more types of texture filtering and one or more types of pixel filtering; as well as Control logic, configured to cause the filter group to execute a filtering task among a plurality of filtering tasks, the plurality of filtering tasks including a texturing filtering task and a pixel filtering task; The process of enabling the filter group to perform pixel filtering tasks includes: enabling the filter group to use a set of weights stored in the weight buffer through the one or more arithmetic logic unit clusters to perform one of the one or more types of pixel filtering on the data set stored in the first buffer, and storing the result of the pixel filtering in the second buffer for output to the memory; The process of having the filter group perform a texture filtering task includes: having the filter group use a set of weights stored in the weight buffer to perform one of the one or more types of texture filtering on a data set from the memory stored in the second buffer, and storing the result of the texture filtering in the first buffer for output to the one or more arithmetic logic unit clusters.
2. The input / output filter unit according to claim 1, wherein, The filter group includes one or more filter blocks, each filter block including multiple arithmetic components, which can be selectively enabled to enable the filter group to perform one of the multiple types of filtering.
3. The input / output filter unit according to claim 2, wherein, The plurality of arithmetic components are configured to form a pipeline.
4. The input / output filter unit according to claim 2, wherein, The plurality of arithmetic components include forming Enter x A set of arithmetic components of a weighted filter, wherein It is an integer.
5. The input / output filter unit according to claim 4, wherein, The set of arithmetic components includes A multiplier unit, each configured to multiply an input value and a weight, and a plurality of adder units, the plurality of adder units forming an adder tree, the adder tree being configured to generate the input value. The sum of the outputs of each multiplier.
6. The input / output filter unit according to claim 4, wherein, The plurality of arithmetic components also include A comparator, each of which is configured to compare input values and provide the result of the comparison as the input value. Enter x Input to the weighted filter.
7. The input / output filter unit according to claim 4, wherein, The plurality of arithmetic components further include a scaling component, the scaling component being configured to receive the... Enter x The output of the weighted filter is used to generate a scaled version of the output.
8. The input / output filter unit according to any one of claims 2 to 7, wherein, The filter group includes multiple filter blocks.
9. The input / output filter unit according to any one of claims 2 to 7, wherein, The control logic is configured to cause the filter group to perform one of the multiple types of filtering on a data set stored in one of the first and second buffers by having one of the filter blocks perform a first part of the filtering of the type in a first round of the filter block and a second part of the filtering of the type in a second round of the filter block.
10. The input / output filter unit according to claim 9, wherein, Temporary data is generated during at least one of the first and second rounds, and the temporary data is stored in one of the first and second buffers.
11. The input / output filter unit according to any one of claims 1 to 7, wherein, The one or more types of texture filtering include one or more of bilinear filtering, trilinear filtering, anisotropic filtering, and percentage asymptotic filtering.
12. The input / output filter unit according to any one of claims 1 to 7, wherein, The one or more types of pixel filtering include one or more of downsampling, upsampling, and multisampling anti-aliasing box filtering.
13. The input / output filter unit according to any one of claims 1 to 7, wherein, The filter group can be further configured to perform texture blending and / or a set of convolutional operations as part of the convolutional layers of a neural network.
14. The input / output filter unit according to any one of claims 1 to 7, further comprising a texture address generator configured to generate addresses of one or more associated texels for performing a type of texture filtering on a fragment or pixel.
15. The input / output filter unit according to any one of claims 1 to 7, further comprising a weight generator configured to generate the set of weights for performing one or more types of filtering and to store the generated weights in the weight buffer.
16. A method for controlling an input / output filter unit, the input / output filter unit comprising a first buffer, a second buffer, a weight buffer, and a configurable filter group, the first buffer being configured to store data received from and output to one or more arithmetic logic unit clusters, the one or more arithmetic logic unit clusters being configured to perform a coloring task, the second buffer being configured to store data received from and output to a memory, the method comprising: Receive information identifying multiple filtering tasks, including texture filtering tasks and pixel filtering tasks. The information identifying the filtering tasks includes: a data set stored in one of the first and second buffers, a weight set stored in the weight buffer, and information about one type of filtering among multiple types, wherein the multiple types of filtering include one or more types of texture filtering and one or more types of pixel filtering; and The configurable filter group performs the identified filtering task by causing it to perform the following operations: Perform filtering of the identified type on the identified data set using the identified weight set, and The result of the filtering is stored in one of the first buffer and the second buffer; The pixel filtering task includes: using a set of weights stored in the weight buffer, the cluster of one or more arithmetic logic units performs one type of pixel filtering on a data set stored in the first buffer, and stores the result of the pixel filtering in the second buffer for output to the memory; and The texture filtering task includes: using a set of weights stored in the weight buffer, performing one of the one or more types of texture filtering on a data set from the memory stored in the second buffer, and storing the result of the texture filtering in the first buffer for output to the one or more arithmetic logic unit clusters.
17. A graphics processing unit comprising an input / output filter unit according to any one of claims 1 to 7.
18. A computer-readable storage medium storing a computer-readable description of an input / output filter unit according to any one of claims 1 to 7, wherein when processed in an integrated circuit manufacturing system, the computer-readable description causes the integrated circuit manufacturing system to manufacture an integrated circuit including the input / output filter unit.