Weighted parameters for applying upsampling
By applying edge and line filters to determine weighted parameters for upsampling in image processing, the problem of inefficiency of existing super-resolution technologies in resource-limited devices is solved, and an efficient and low-power image upsampling effect is achieved.
Patent Information
- Application Number
- CN202411878741.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
While improving image resolution, existing super-resolution technologies face efficiency problems such as high-performance computing needs, delay, power consumption and silicon area, especially in small battery-powered devices with limited resources.
A method is provided to determine weighted parameters for upsampling by applying horizontal and vertical edge filters, horizontal and vertical line filters to upsample within an image area, reducing dependence on computing resources.
This method can reduce the need for latency, power consumption and silicon area while reducing image blur and artifacts, and is suitable for use in resource-limited devices.
Smart Images

Figure CN120182091A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to UK Patent Application No. 2319649.6 filed on December 20, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates to upsampling. Specifically, upsampling can be applied to input pixel values representing an image region to determine a block of upsampled pixel values, for example, for super-resolution techniques. Background Art
[0004] The term "super-resolution" refers to a technique for upsampling an image to enhance its apparent visual quality, for example, by estimating the appearance of a higher-resolution version of the image. When implementing super-resolution, the system will attempt to find a higher-resolution version of the lower-resolution input image that is, to the greatest extent possible, reasonable and consistent with the lower-resolution input image. Super-resolution is a challenging problem because for each patch in the lower-resolution input image, there are a large number of potential higher-resolution patches that could correspond to it. In other words, super-resolution techniques are attempting to solve an ill-posed problem because although solutions exist, they are not unique.
[0005] Super-resolution has important applications. It can be used to increase the resolution of an image, thereby improving the "quality" of the image perceived by an observer. Additionally, super-resolution can be used as a post-processing step in an image generation process, thereby allowing an image to be generated at a lower resolution (which is typically simpler and faster), while still producing a high-quality, high-resolution image. The image generation process can be, for example, an image capture process using a camera. Alternatively, the image generation process can be an image rendering process in which a computer, such as a graphics processing unit (GPU), renders an image of a virtual scene. Compared to directly rendering a high-resolution image using a GPU, allowing the GPU to render a low-resolution image and then applying super-resolution techniques to upsample the rendered image to produce a high-resolution image has the potential to significantly reduce the GPU's latency, bandwidth, power consumption, silicon area, and / or computational cost. The GPU can implement any suitable rendering technique, such as rasterization or ray tracing. For example, the GPU can render a 960x540 image (i.e., an image having 518,400 pixels arranged in 960 columns and 540 rows), and then the image can be upsampled by a factor of 2 (referred to as "2x upsampling") in both the horizontal and vertical dimensions to produce a 1920x1080 image (i.e., an image having 2,073,600 pixels arranged in 1920 columns and 1080 rows). In this way, to produce a 1920x1080 image, the GPU renders an image having a quarter of the number of pixels. This results in very significant savings during rendering (e.g., in terms of the GPU's latency, power consumption, and / or silicon area), and can, for example, allow a relatively low-performance GPU to render high-quality, high-resolution images within a low power and area budget, provided that an appropriately efficient and high-quality super-resolution implementation is used to perform the upsampling. In other examples, different upsampling factors (other than 2x) can be applied.
[0006] Figure 1 An upsampling process is shown. An input image 102 having a relatively low resolution is processed by a processing module 104 to produce an output image 106 having a relatively high resolution. In some systems, the processing module 104 can be implemented as a neural network to upsample the input image 102 to produce an upsampled output image 106. Implementing the processing module 104 as a neural network can produce a high-quality output image, but typically requires a high-performance computing system (e.g., having large, powerful processing units and memory) to implement the neural network. Therefore, due to processing time, latency, bandwidth, power consumption, memory usage, silicon area, and computational cost, it may be inappropriate to implement the processing module 104 as a neural network for performing upsampling of an image. These efficiency considerations are particularly important in some devices, such as small battery-powered devices with limited computational and bandwidth resources, such as mobile phones and tablets.
[0007] Some systems do not perform super-resolution on images using neural networks, but rather use more conventional processing modules. For example, some systems divide the problem into two stages: (i) upsampling and (ii) adaptive sharpening. In these systems, the upsampling stage can be performed at relatively low cost, for example, using bilinear upsampling, while the adaptive sharpening stage can be used to sharpen the image, i.e., reduce the blur introduced by upsampling. Bilinear upsampling is known in the art and uses linear interpolation of adjacent input pixels in two dimensions to generate output pixels at positions between the input pixels.
[0008] The general objectives of a system implementing super-resolution are: (i) a high-quality output image, i.e., an output image that is as reasonable as possible given the low-resolution input image, (ii) low latency such that the output image is generated quickly, and (iii) a processing module that is low-cost in terms of resources such as power, bandwidth, and silicon area. Summary of the Invention
[0009] This Summary of the Invention is provided to introduce a series of concepts that are further described below in the Detailed Description in a simplified form. This Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0010] A method is provided for determining an indication of one or more weighting parameters for a block that applies upsampling to input pixel values representing an image region to determine one or more upsampled pixel values, the method comprising:
[0011] Applying a horizontal edge filter to two or more of the input pixel values to determine a first filtered value;
[0012] Applying a vertical edge filter to two or more of the input pixel values to determine a second filtered value;
[0013] Applying a horizontal line filter to three or more of the input pixel values to determine a third filtered value;
[0014] Applying a vertical line filter to three or more of the input pixel values to determine a fourth filtered value;
[0015] Using the first filtered value, the second filtered value, the third filtered value, and the fourth filtered value to determine the indication of the one or more weighting parameters, wherein the one or more weighting parameters indicate relative horizontal and vertical variations of the input pixel values within the image region; and
[0016] Outputting the determined indication of the one or more weighting parameters for the block that applies upsampling to the input pixel values representing the image region to determine one or more upsampled pixel values.
[0017] The indication of determining one or more weighted parameters using the first filter value, the second filter value, the third filter value, and the fourth filter value may include combining the first filter value, the second filter value, the third filter value, and the fourth filter value by performing a weighted sum.
[0018] One or more of the weights used in the weighted sum may be trained such that the indication of the one or more weighted parameters indicates one or more weighted parameters indicative of relative horizontal and vertical variations of the input pixels within the image region.
[0019] Before outputting the indication of the one or more weighted parameters, the indication of the one or more weighted parameters may be clamped to be within the range [0, 1].
[0020] The indication of determining one or more weighted parameters using the first filter value, the second filter value, the third filter value, and the fourth filter value may include:
[0021] processing the input pixel values using an implementation of a neural network to determine a residual value; and
[0022] combining the determined residual value with the first filter value, the second filter value, the third filter value, and the fourth filter value to determine the indication of the one or more weighted parameters.
[0023] The neural network may include a first convolutional layer, a second convolutional layer, and a third convolutional layer, and wherein processing the input pixel values using the implementation of the neural network may include:
[0024] using the first convolutional layer to determine a first intermediate tensor based on the input pixel values, wherein the first intermediate tensor extends in only one of the horizontal dimension and the vertical dimension with respect to the dimensions of the image region represented by the input pixel values;
[0025] using the second convolutional layer to determine a second intermediate tensor based on the first intermediate tensor, wherein the second intermediate tensor does not extend in the horizontal dimension or the vertical dimension with respect to the dimensions of the image region represented by the input pixel values; and
[0026] using the third convolutional layer to determine the residual value based on the second intermediate tensor.
[0027] The image region can be a part of the input image, and the upsampling can be iteratively performed for multiple partially overlapping image regions within the input image. The processing of the input pixel values using the neural network embodiment can also include storing the first intermediate tensor in a buffer. The first convolutional layer can operate on the input pixel values within a first portion of the current image region that does not overlap with the previous image region, but may not operate on the input pixel values within a second portion of the current image region that overlaps with the previous image region, to determine the first intermediate tensor.
[0028] The neural network can also include a first activation function implemented between the first convolutional layer and the second convolutional layer, and a second activation function implemented between the second convolutional layer and the third convolutional layer.
[0029] The first activation function can be a first rectified linear unit, and the second activation function can be a second rectified linear unit, and the processing of the input pixel values using the neural network embodiment can also include: (i) setting negative values in the first intermediate tensor to zero using the first rectified linear unit, and (ii) setting negative values in the second intermediate tensor to zero using the second rectified linear unit.
[0030] The first activation function can be an identity function and the second activation function can be an absolute value function, and the processing of the input pixel values using the neural network embodiment can also include setting the values in the second intermediate tensor to their absolute values using the absolute value function.
[0031] The neural network may have been trained using quantization aware training (QAT).
[0032] The horizontal line filter can be configured such that when three or more of the input pixel values to which the horizontal line filter is applied exhibit a purely vertical feature, the third filter value is determined to be zero. The vertical line filter can be configured such that when three or more of the input pixel values to which the vertical line filter is applied exhibit a purely horizontal feature, the fourth filter value is determined to be zero.
[0033] The input pixel values can be the values of an input pixel in a repeating five-point arrangement whose positions correspond to the upsampled pixel positions.
[0034] The input pixel value can be represented by two input blocks. One of the two input blocks can include the input pixel values of the input pixels at positions in the odd rows of the repeated five-point arrangement where the positions correspond to the upsampled pixel positions, and the other input block of the two input blocks can include the input pixel values of the input pixels at positions in the even rows of the repeated five-point arrangement where the positions correspond to the upsampled pixel positions.
[0035] Each filter can have a filter kernel that has weights at the input pixel positions in a 5x5 region corresponding to the upsampled pixel positions, with the upsampled pixel positions centered at the upsampled pixel positions that fall between the positions of adjacent input pixels in the five-point arrangement.
[0036] The horizontal edge filter can have a filter kernel whose weights can be expressed as
[0037] 0.5 0.5 0 1 0 0 0 0 -1 0 -0.5 -0.5
[0038] The vertical edge filter can have a filter kernel whose weights can be expressed as
[0039] 0 0 -0.5 0 0.5 -1 1 -0.5 0 0.5 0 0
[0040] The horizontal line filter can have a filter kernel whose weights can be expressed as
[0041] -0.5 -0.5 0 0 0 1 1 0 0 0 -0.5 -0.5
[0042] The vertical line filter can have a filter kernel whose weights can be expressed as
[0043] 0 0 -0.5 1 -0.5 0 0 -0.5 1 -0.5 0 0
[0044] Each filter can have one or more filter kernels that have weights at the input pixel positions in a 6x6 region corresponding to the upsampled pixel positions centered on a 2x2 block of upsampled pixel positions, where each filter can be configured to determine the filtered values for the top-right upsampled pixel position TR and the bottom-left upsampled pixel position BL of the 2x2 block of upsampled pixel positions. The vertical line filter can have six filter kernels, denoted as kernel 0 to kernel 5, centered on a 2x2 block of upsampled pixel positions, and their weights can be expressed as:
[0045]
[0046] Kernel 0
[0047]
[0048] Kernel 1
[0049]
[0050] Kernel 2
[0051]
[0052] Kernel 3
[0053]
[0054] Kernel 4
[0055]
[0056] Kernel 5
[0057] The filtering value of the vertical line filter for each of the upper right upsampled pixel position and the lower left upsampled pixel position of the 2x2 block of upsampled pixel positions can be determined by finding the weighted sum of the absolute values of the outputs of multiple of six filter kernels (denoted as kernels 0 to 5). The horizontal line filter can have six filter kernels, denoted as kernels 6 to 11, centered on the 2x2 block of upsampled pixel positions, and its weights can be expressed as:
[0058]
[0059] Kernel 6
[0060]
[0061] Kernel 7
[0062]
[0063] Kernel 8
[0064]
[0065] Kernel 9
[0066]
[0067] Kernel 10
[0068]
[0069] Kernel 11
[0070] The filtering value of the horizontal line filter for each of the upper-right upsampled pixel position and the lower-left upsampled pixel position in the 2x2 block of upsampled pixel positions can be determined by finding the weighted sum of the absolute values of the outputs of multiple of the six filter kernels (denoted as kernels 6 to 11). The vertical edge filter can have two filter kernels, denoted as kernels 12 to 13, centered on the 2x2 block of upsampled pixel positions, and its weights can be expressed as:
[0071]
[0072] Kernel 12
[0073]
[0074] Kernel 13
[0075] The filtering value of the vertical edge filter for both the upper-right upsampled pixel position and the lower-left upsampled pixel position in the 2x2 block of upsampled pixel positions can be determined as the sum of the absolute value of the output of kernel 12 and the absolute value of the output of kernel 13. The horizontal edge filter can have two filter kernels, denoted as kernels 14 to 15, centered on the 2x2 block of upsampled pixel positions, and its weights can be expressed as:
[0076]
[0077] Kernel 14
[0078]
[0079] Kernel 15
[0080] The filtering value of the horizontal edge filter for both the upper-right upsampled pixel position and the lower-left upsampled pixel position in the 2x2 block of upsampled pixel positions can be determined as the sum of the absolute value of the output of kernel 14 and the absolute value of the output of kernel 15.
[0081] Input pixels can be represented by values in multiple channels. Upsampled pixels can be represented by values in multiple channels. The input pixel value can be the value of an input pixel in a single channel, and the upsampled pixel value can be the value of an upsampled pixel in the single channel.
[0082] The input pixel value and the upsampled pixel value can be Y-channel values, and the upsampling of Y-channel values can be used in super-resolution techniques.
[0083] The input pixel value and the upsampled pixel value can be green-channel values, and the upsampling of green-channel values can be used in demosaicing techniques.
[0084] The method may further include applying upsampling to input pixel values representing an image region. The upsampling may include determining one or more of the upsampled pixel values of a block of the one or more upsampled pixel values based on relative horizontal and vertical variations of the input pixel values within the image region indicated by the one or more weighting parameters.
[0085] Determining each of the one or more of the upsampled pixel values may include: based on the determined one or more weighting parameters:
[0086] Applying one or more first kernels to at least a first subset of the input pixel values to determine a horizontal component,
[0087] Applying one or more second kernels to at least a second subset of the input pixel values to determine a vertical component, and
[0088] Combining the determined horizontal and vertical components,
[0089] such that each of the one or more of the upsampled pixel values is determined based on relative horizontal and vertical variations of the input pixel values within the image region.
[0090] The combining the determined horizontal and vertical components may include:
[0091] Multiplying the horizontal component by a first weighting parameter of the weighting parameters to determine a weighted horizontal component;
[0092] Multiplying the vertical component by a second weighting parameter of the weighting parameters to determine a weighted vertical component; and
[0093] Summing the weighted horizontal component and the weighted vertical component.
[0094] One or more of the upsampled pixel values of the block of the one or more upsampled pixel values may be unsharp upsampled pixel values.
[0095] One or more of the upsampled pixel values of the block of the one or more upsampled pixel values may be sharpened upsampled pixel values.
[0096] There is provided a processing module configured to determine an indication of one or more weighting parameters for applying upsampling to input pixel values representing an image region to determine a block of one or more upsampled pixel values, the processing module including:
[0097] Horizontal edge filtering logic configured to apply a horizontal edge filter to two or more of the input pixel values to determine a first filtered value;
[0098] Vertical edge filtering logic configured to apply a vertical edge filter to two or more of the input pixel values to determine a second filtered value;
[0099] Horizontal line filtering logic configured to apply a horizontal line filter to three or more of the input pixel values to determine a third filtered value;
[0100] Vertical line filtering logic configured to apply a vertical line filter to three or more of the input pixel values to determine a fourth filtered value; and
[0101] Processing logic configured to:
[0102] Use the first filtered value, the second filtered value, the third filtered value, and the fourth filtered value to determine an indication of one or more weighting parameters, where the one or more weighting parameters indicate relative horizontal and vertical variations of the input pixel values within the image region; and
[0103] Output the determined indication of the one or more weighting parameters for a block that applies upsampling to the input pixel values representing the image region to determine one or more upsampled pixel values.
[0104] The processing logic may be configured to perform a weighted sum of the first filtered value, the second filtered value, the third filtered value, and the fourth filtered value.
[0105] The processing module may further include an implementation of a neural network configured to determine a residual value, where the processing logic may be configured to combine the residual value with the first filtered value, the second filtered value, the third filtered value, and the fourth filtered value to determine an indication of one or more weighting parameters.
[0106] The neural network may include:
[0107] A first convolutional layer configured to determine a first intermediate tensor based on the input pixel values, where the first intermediate tensor extends only in one of the horizontal dimension and the vertical dimension with respect to the dimensions of the image region represented by the input pixel values;
[0108] A second convolutional layer configured to determine a second intermediate tensor based on the first intermediate tensor, where the second intermediate tensor does not extend in the horizontal dimension or the vertical dimension with respect to the dimensions of the image region represented by the input pixel values; and
[0109] A third convolutional layer configured to determine the residual value based on the second intermediate tensor.
[0110] The processing module may further include pixel determination logic configured to determine one or more of the upsampled pixel values of a block of the one or more upsampled pixel values based on relative horizontal and vertical changes of the input pixel values within the image region indicated by the one or more weighting parameters.
[0111] A processing module may be provided that is configured to perform any of the methods described herein.
[0112] Computer-readable code may be provided that is configured to cause any of the methods described herein to be performed when the code runs.
[0113] An integrated circuit definition dataset may be provided that, when processed in an integrated circuit manufacturing system, configures the integrated circuit manufacturing system to manufacture a processing module as described herein.
[0114] The processing module may be embodied in hardware on an integrated circuit. A method of manufacturing a processing module at an integrated circuit manufacturing system may be provided. An integrated circuit definition dataset may be provided that, when processed in an integrated circuit manufacturing system, configures the system to manufacture a processing module. A non-transitory computer-readable storage medium may be provided having stored thereon a computer-readable description of a processing module that, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture an integrated circuit incorporating the processing module.
[0115] An integrated circuit manufacturing system may be provided that includes: a non-transitory computer-readable storage medium having stored thereon a computer-readable description of a processing module; a layout processing system configured to process the computer-readable description to generate a circuit layout description of an integrated circuit incorporating the processing module; and an integrated circuit generation system configured to manufacture the processing module based on the circuit layout description.
[0116] Computer program code for performing any of the methods described herein may be provided. A non-transitory computer-readable storage medium may be provided having stored thereon computer-readable instructions that, when executed in a computer system, cause the computer system to perform any of the methods described herein.
[0117] As will be apparent to those skilled in the art, the above features may be combined as appropriate and may be combined with any aspect of the examples described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0118] Examples will now be described in detail with reference to the accompanying drawings, in which:
[0119] Figure 1 An upsampling process is shown;
[0120] Figure 2 Input pixels in a repeated five-point arrangement whose positions correspond to the upsampled pixel positions are shown;
[0121] Figure 3 Representing the input pixels as two input blocks is shown;
[0122] Figure 4 A processing module is shown that is configured to upsample the input pixel values to determine a block of upsampled pixel values;
[0123] Figure 5 Is a flowchart of a method of applying upsampling to input pixel values representing an image region to determine a block of upsampled pixel values;
[0124] Figure 6 Is a flowchart of the method steps for determining upsampled pixel values;
[0125] Figure 7 A portion of the input pixel values is shown, and how the block of upsampled pixel values is related to the input pixel values is shown;
[0126] Figure 8a A first kernel for applying a first weighting parameter is shown;
[0127] Figure 8b A second kernel for applying a second weighting parameter is shown;
[0128] Figure 9 Weighting parameter determination logic configured to determine an indication of one or more weighting parameters is shown;
[0129] Figure 10 Is a flowchart of a method of determining an indication of (multiple) weighting parameters using the weighting parameter determination logic;
[0130] Figure 11 Shows in more detail an example of step S1012 from the Figure 10 method shown in;
[0131] Figure 12 An implementation of a neural network configured to output residual values for determining an indication of one or more weighting parameters is shown;
[0132] Figure 13 Is a flowchart of a method of determining residual values using an implementation of a neural network;
[0133] Figure 14Shows pixel determination logic configured to determine sharpened upsampled pixel values;
[0134] Figure 15 Is a flowchart of a method of a block for determining sharpened upsampled pixel values;
[0135] Figure 16 Shows a first complete sharpened upsampling kernel as a result of convolving a bilinear interpolation kernel with a sharpening kernel;
[0136] Figure 17 Shows how the first complete sharpened upsampling kernel can be deconstructed to determine the sharpened upsampled pixel value in the upper left corner of a block for determining sharpened upsampled pixel values into: (i) a kernel applied to a first input block of input pixel values, and (ii) a kernel applied to a second input block of input pixel values;
[0137] Figure 18 Shows how the first complete sharpened upsampling kernel can be deconstructed to determine the sharpened upsampled pixel value in the lower right corner of a block for determining sharpened upsampled pixel values into: (i) a kernel applied to a first input block of input pixel values, and (ii) a kernel applied to a second input block of input pixel values;
[0138] Figure 19 Shows a complete horizontal sharpened upsampling kernel as a result of convolving a linear horizontal interpolation kernel with a sharpening kernel;
[0139] Figure 20 Shows a complete vertical sharpened upsampling kernel as a result of convolving a linear vertical interpolation kernel with a sharpening kernel;
[0140] Figure 21 Is a flowchart of a method of combining the results of applying kernels to input blocks to determine sharpened upsampled pixel values at positions falling between corresponding positions of input pixels;
[0141] Figure 22 Shows how the complete horizontal sharpened upsampling kernel can be deconstructed to determine the horizontal component of the sharpened upsampled pixel value in the upper right corner of a block for determining sharpened upsampled pixel values into: (i) a kernel applied to a first input block of input pixel values, and (ii) a kernel applied to a second input block of input pixel values;
[0142] Figure 23 Shows how the complete vertical sharpened upsampling kernel can be deconstructed to determine the vertical component of the sharpened upsampled pixel value in the upper right corner of a block for determining sharpened upsampled pixel values into: (i) a kernel applied to a first input block of input pixel values, and (ii) a kernel applied to a second input block of input pixel values;
[0143] Figure 24Illustrated is how a complete horizontal sharpened upsampling kernel can be decomposed into: (i) a kernel applied to a first input block of input pixel values, and (ii) a kernel applied to a second input block of input pixel values, in order to determine a horizontal component of a sharpened upsampling pixel value in the lower left corner of a block of sharpened upsampling pixel values;
[0144] Figure 25 Illustrated is how a complete vertical sharpened upsampling kernel can be decomposed into: (i) a kernel applied to a first input block of input pixel values, and (ii) a kernel applied to a second input block of input pixel values, in order to determine a vertical component of a sharpened upsampling pixel value in the lower left corner of a block of sharpened upsampling pixel values;
[0145] Figure 26 Illustrated is a computer system in which a processing module is implemented; and
[0146] Figure 27 Illustrated is an integrated circuit manufacturing system for generating an integrated circuit incorporating a processing module.
[0147] The accompanying drawings illustrate various examples. Those skilled in the art will appreciate that the element boundaries shown in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one example of a boundary. In some examples, it may be the case that one element can be designed as multiple elements, or multiple elements can be designed as one element. Where appropriate, common reference numerals are used throughout the drawings to indicate like features. Detailed Description
[0148] The following description is presented by way of example to enable those skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art.
[0149] Embodiments will now be described solely by way of example. Figure 2 Illustrated are input pixel values 202 of input pixels in a repeating five-point arrangement where the positions correspond to upsampled pixel positions. In this arrangement, half of the upsampled (output) pixel positions (e.g., position 204) are Figure 2 shown shaded and correspond to the positions of the input pixels; while the other half of the upsampled pixel positions (e.g., position 206) are Figure 2 not shown shaded and do not correspond to the positions of the input pixels. The repeating five-point arrangement can be thought of as a "checkerboard" pattern. As described herein, the input pixel values 202 are processed to determine a block of upsampled pixel values, the block of upsampled pixel values including those at Figure 2The upsampled pixel values at some of the upsampled pixel positions shown in that do not correspond to the positions of the input pixels. In this way, the density of the upsampled pixels is twice the density of the input pixels. Additionally, using the repeating pentagon arrangement means that the output pixel positions that do not correspond to the positions of the input pixels (not shown shaded in ) are horizontally adjacent to two output pixel positions that do correspond to the positions of the input pixels (shown shaded in ), and vertically adjacent to two output pixel positions that do correspond to the positions of the input pixels (shown shaded in ). Thus, a sparser pattern relative to the input pixels can be achieved, such as the case where only the top left corner of each 2x2 block is available, and the error in the output pixel values at those output pixel positions that do not correspond to the positions of the input pixels (not shown shaded in ) can be kept small. Additionally, since all the missing output pixel positions have the same adjacent pixel arrangement, the upsampling algorithm for the pentagon pattern input can be more regular than the algorithms for sparser input patterns. Figure 2 not shown shaded in Figure 2 shown shaded in Figure 2 shown shaded in Figure 2 not shown shaded in
[0150] Furthermore, the upsampling process described in the examples herein can take into account the horizontal and vertical variations of the input pixel values. The horizontal and vertical variations of the input pixel values can indicate, for example, edges, textures, and / or high-frequency content in the image region represented by the input pixel values. For example, the relative horizontal and vertical variations of the input pixel values can represent the image gradient within the image region. For example, an edge in an image causes a high image gradient in the direction perpendicular to the edge, but a low image gradient in the direction parallel to the edge. When determining the upsampled pixel values by performing a weighted sum of the input pixel values, the upsampling process can weight the input pixel values differently according to the image gradient. For example, this can reduce blurring and staircase artifacts or other anisotropic image features near the edges compared to when using bilinear upsampling techniques. Additionally, the examples described below are, for example, highly efficient to implement in hardware, such that they (when implemented in fixed-function hardware) have low latency, low power consumption, and small silicon area.
[0151] Determine at least one weighting parameter, the at least one weighting parameter indicating weights used in a weighted sum for determining an upsampled pixel value. In the examples described herein, the (multiple) weighting parameters are determined using a manually designed algorithm. According to the algorithm, four filters are applied to input pixel values representing an image region to determine four filtered values. Specifically, a horizontal edge filter, a vertical edge filter, a horizontal line filter, and a vertical line filter are applied to the input pixel values to determine four filtered values. Then, the four filtered values can be used (e.g., combined in a weighted sum) to determine an indication of the (multiple) weighting parameters, which can then be used to apply upsampling to the input pixel values representing the image region. The indication of the (multiple) weighting parameters can be the (multiple) weighting parameters. The indication of the (multiple) weighting parameters can be something other than the (multiple) weighting parameters from which the (multiple) weighting parameters can be determined. In some examples, there are multiple weighting parameters, and the indication of the weighting parameters can be some of the weighting parameters but not all of the weighting parameters, where the (multiple) other weighting parameters can be derived from the (multiple) weighting parameters explicitly included in the indication. For example, there can be two weighting parameters (denoted as a and b in the examples given below), and the indication of the weighting parameters can be the value of the first weighting parameter among the (only) weighting parameters (e.g., the value of a), where the value of the other weighting parameter (e.g., the value of b) can be derived from the value of the first weighting parameter using, for example, a predetermined relationship between the values of the two weighting parameters (e.g., it may be known that a + b = 1).
[0152] (Multiple) weighting parameters indicate relative horizontal and vertical variations of input pixel values within an image region. One or more weighting parameters indicate the directionality of filtering that will be applied when upsampling the input pixel values.
[0153] Using horizontal and vertical filters to determine the (multiple) weighting parameters allows adjustment of upsampling to accommodate the directionality of features within the image region being processed. Additionally, using edge filters and line filters to determine the (multiple) weighting parameters allows adjustment of upsampling to accommodate different types of directional features (e.g., edges and lines) within the image region being processed. Thus, compared to when using other upsampling techniques (e.g., bilinear upsampling), using all four filters can reduce blurring and stair-step artifacts. Additionally, the logic for applying the four filters and for combining the four filtered values (e.g., by performing a weighted sum) can be implemented very efficiently in hardware, such as in fixed-function circuitry, using simple operations such as shifts, additions, and multiplications. The methods described herein (e.g., when implemented in fixed-function hardware) can be implemented with low latency, low power consumption, and small silicon area.
[0154] In the examples described herein, the filters have zero response to features in orthogonal directions. Specifically, the horizontal edge filter and the horizontal line filter have zero response to vertical features, and the vertical edge filter and the vertical line filter have zero response to horizontal features. "Vertical features", such as vertical edges and vertical lines, are constant for different vertical positions but vary for different horizontal positions; while "horizontal features", such as horizontal edges and horizontal lines, are constant for different horizontal positions but vary for different vertical positions.
[0155] In addition, in the examples described herein, an implementation of a neural network can be used to process input pixel values to determine a residual value, which can be combined with a filtering value to determine an indication of the (multiple) weighting parameters. The neural network has been trained to determine the residual value such that the (multiple) weighting parameter indicates the directionality of the filtering that will be applied when upsampling the input pixel values. Using the neural network to determine the residual value allows the determination of the (multiple) weighting parameters to be adapted to give perceptually good results without the need to classify any directional variations in the image region as edges, lines, or textures or any other specific type of high-frequency content. This is advantageous because it is difficult and time-consuming to manually design a system to utilize some or all of these directional cues (e.g., edges, lines, textures, etc.), and furthermore, it is easy to miss subtle useful attributes of the input region that a suitable neural network can better identify and utilize. By using the neural network to only determine the residual value rather than directly determining the weighting parameters or even the upsampled output pixels themselves, the neural network can remain small and simple. This means that the implementation of the neural network can also be small and simple, such that it can have low silicon area, low latency, and low power consumption. In this way, the neural network is used for the part of the process (e.g., determining the residual value for determining the (multiple) weighting parameters) that benefits the most from being implemented in the neural network, while the other parts of the process (e.g., the rest of the upsampling and sharpening processes) are not implemented using the neural network, so that the neural network does not become too large and complex to be implemented, for example, on a low-cost device with strict limitations on power consumption and silicon area.
[0156] In some examples, the input pixel values are received in two input blocks. Figure 3 The input pixel values 302 are shown represented by two input blocks 304 and 306. The first input block 304 includes the input pixel values of the input pixels at positions corresponding to the positions within the odd rows of the repeated five-point arrangement that correspond to the upsampled pixel positions (shown with line shading in Figure 3 ), and the second input block 306 includes the input pixel values of the input pixels at positions corresponding to the positions within the even rows of the repeated five-point arrangement that correspond to the upsampled pixel positions (in Figure 3The input pixel value of the input pixel (shown in cross - hatching). Using two input blocks 304 and 306 means that the input is denser than when the input pixel values are received as blocks having a repeating pentagon arrangement (represented as 302 in Figure 3 ).
[0157] In different examples, the format of the pixel values may vary. For example, pixels can be represented with pixel values in the YUV format (where each pixel has a value in each of the Y, U, and V channels), and upsampling can be applied to each of the Y, U, and V channels separately. The upsampling described herein can be applied to only the Y channel, while upsampling of the U and V channels is performed in a simpler manner, for example, using bilinear interpolation. In other words, in some examples, the input pixel values and the upsampled pixel values are Y - channel values, and the upsampling of the Y - channel values can be used in super - resolution techniques. In other examples, the upsampling described herein can be applied to each of the Y, U, and V channels. In other examples, the results of some or all of the calculations (such as contrast or weighting parameters) from processing the Y pixel values can be applied to the U and V channels to save computation and promote consistency in reconstructed edges. If the input pixel data is in the RGB format, it can be converted to the YUV format (for example, using known color - space conversion techniques) and then processed as data in the Y, U, and V channels. Alternatively, if the input pixel data is in the RGB format (where each pixel has a value in each of the R, G, and B channels), the techniques described herein can be implemented on the R, G, and B channels as described herein, where the G channel can be considered a proxy for the Y channel. If the input data includes an alpha channel, upsampling (for example, using bilinear interpolation) can be applied to the alpha channel independently.
[0158] In addition, the upsampling described herein can be used as part of a demosaicking technique. Demosaicking is a known technique and can be implemented, for example, in a camera pipeline or on a graphics processing unit (GPU). Pixel values of a pixel can have multiple channels, e.g., a pixel can have pixel values in red, green, and blue channels. In some cases, each pixel can have a pixel value in at least one channel but not all channels, and the purpose of the demosaicking process is generally to determine the pixel values in all channels for all pixels. For example, an image sensor can detect pixel values using a color filter array. An example of a typical color filter array is a Bayer filter, where green pixel values are detected for pixels in a five-point arrangement (or checkerboard pattern), and red pixel values are detected for half of the remaining pixels, and blue pixel values are detected for the other half of the remaining pixels. The upsampling described herein can be applied to the pixel values within a channel (e.g., the green channel) to determine the pixel values (e.g., green pixel values) in that channel for all pixels. Some other methods can be used to determine the remaining pixel values in other channels (e.g., red and blue pixel values) for all pixels, and these some other methods can (or may not) be based on the upsampled (e.g., green) pixel values of the pixels. In other words, in some examples, the input pixel values and the upsampled pixel values are green channel values, and the upsampling of the green channel values can be used in the demosaicking technique.
[0159] In the examples described herein, the input pixels can be represented by values in multiple channels, and the upsampled pixels can be represented by values in multiple channels. The upsampling in the examples described herein is applied to the input pixel values, which are the values of the input pixels in a single channel, and the upsampled pixel values are the values of the upsampled pixels in the single channel. For example, upsampling can be applied to the Y pixel values in a super-resolution technique, or to the green pixel values in a demosaicking technique. In an example where upsampling is applied to the pixel values of a specific channel as part of a demosaicking technique, the upsampling increases the number of pixel positions where the pixel values of the specific channel exist. It should be noted that the entire demosaicking process does not change the total number of sampling positions: rather, it changes what is represented at each sampling position, e.g., it can change each monochromatic sampling point to a complete RGB sampling point. Upsampling can be regarded as interpolation.
[0160] Figure 4 Shown is a block 404 configured to apply upsampling to input pixel values 302 representing an image region to determine upsampled pixel values, e.g., a processing module 402 for implementing a super-resolution technique. The block 404 for upsampled pixel values represents at Figure 4The position within the image region indicated by the square 410 in the figure. In this example, the input pixel value 302 is the value of the input pixel at the position corresponding to the 10x10 block of upsampled pixel positions, and the block 404 of upsampled pixel values is a 2x2 block of upsampled pixel values representing the central 2x2 portion of the 10x10 block of upsampled pixel positions. However, it should be noted that in other examples, the shape and / or size of the image region represented by the blocks of input and upsampled pixel values may be different. In some examples, the block of upsampled pixel values may have a single upsampled pixel value, so although we refer to a block of upsampled pixel values in the main examples described herein, more generally, we may refer to one or more blocks of upsampled pixel values. The processing module 402 includes pixel determination logic 406 and weighted parameter determination logic 408. The logic of the processing module 402 can be implemented in hardware, software, or a combination thereof. Compared with a software implementation, a hardware implementation generally provides reduced latency and power consumption as well as higher throughput, but at the cost of less flexibility in operation. The processing module 402 may be used many times in the same way on each upsampled image, and since latency is very important in, for example, real-time super-resolution applications, implementing the logic of the processing module 402 in hardware (e.g., in fixed-function circuitry) may be more preferable than implementing the logic in software. However, a software implementation is still possible and may be preferable in certain cases (e.g., when the microchip area is not sufficient to include additional fixed-function hardware).
[0161] The processing module 402 performs upsampling depending on the relative horizontal and vertical variations of the input pixel value 302 within the image region, rather than using bilinear upsampling (bilinear upsampling is the most common conventional upsampling method). In this way, compared with the case of using bilinear upsampling, the upsampling takes into account the anisotropic features (e.g., edges) in the image to reduce the "staircase" and "fine frill" artifacts and blurring that may occur near the anisotropic features (e.g., the edges of a computer-generated image, especially diagonal edges).
[0162] Reference Figure 5 The flowchart of FIG. describes a method of using the processing module 402 to apply upsampling to the input pixel value 302 to determine the block 404 of one or more upsampled pixel values, for example, for implementing a super-resolution technique.
[0163] In step S502, the input pixel value 302 is received at the processing module 402. In the main example described herein, the input pixel value 302 is represented by two input blocks 304 and 306, as Figure 3As shown. However, in other examples, the input pixel values are not necessarily represented by two input blocks. For example, they can be represented by a single block. As described above, the input pixel value 302 is the value of the input pixels in a repeating five-point arrangement (or "checkerboard pattern") whose positions correspond to the upsampled pixel positions, where the first input block 304 ("input block 1") includes the positions within the odd rows of the repeating five-point arrangement whose positions correspond to the upsampled pixel positions (shown with line shading in Figure 3 ), and the second input block 306 ("input block 2") includes the positions within the even rows of the repeating five-point arrangement whose positions correspond to the upsampled pixel positions (shown with cross shading in Figure 3 ). The input blocks can be regarded as low-resolution representations of the image regions represented by the input pixel values. In the examples described herein, each input block is a 5x5 block of input pixel values, and the block of upsampled pixel values is a 2x2 block of upsampled pixel values. However, in other examples, the input blocks and the blocks of upsampled pixel values can have different shapes and / or sizes.
[0164] The source of the input pixel values does not affect the upsampling process described herein. For example, the input pixel values can be captured (e.g., by a camera), or can be computer-generated (e.g., rendered by a graphics processing unit (GPU) using rendering techniques such as rasterization or ray tracing to represent an image of a scene). For example, the GPU can render all the input pixel values of an image region, i.e., Figure 4 all the input pixel values in the repeating five-point arrangement shown in. In this example, upsampling will double the number of pixel values. Thus, in this example, the GPU can render half the number of pixel values of the pixels in the final (upsampled) image, which allows the GPU to have reduced latency, power consumption, and / or silicon area compared to a GPU that renders the pixel values of all the pixels of the final image.
[0165] As another example, a sequence of frames can be rendered, such as images representing a scene at a sequence of times. For example, this can be useful for rendering images of a computer game when the user is interacting with the scene. In this example, on each frame of the sequence of frames, the input pixel values of only one input block (304 or 306) are rendered, where the input block being rendered alternates across the sequence of frames. Typically, the current frame will be very similar to the previous frame (e.g., the immediately previous frame), and a process of temporal resampling can be performed to estimate the input pixel values of some of the input pixels of the current frame based on the rendered pixel values of the previous frame. As is known to those skilled in the art, the temporal resampling process can use motion vectors to estimate the input pixel values of the current frame based on the rendered pixel values of the previous frame. In this way, when the image region represented by the input pixel values 302 is part of the current frame within the sequence of frames, the input pixel values of the first input block in the input block 304 can be rendered for the current frame, while the input pixel values of the second input block in the input block 306 can be determined by performing temporal resampling on the pixel values of the previous frame using motion vectors. In an example where temporal resampling is used to determine the input pixel values of half of the input pixels of each frame, the number of input pixel values rendered by the GPU for each frame is one quarter of the number of pixel values in the final (upsampled) image, which allows the GPU to have an even further reduced latency, power consumption, and / or silicon area compared to a GPU that renders the pixel values of half of the pixels of the final image for each frame.
[0166] In step S504, the weighted parameter determination logic 408 analyzes the input pixel values 302 to determine one or more weighted parameters. The one or more weighted parameters indicate the relative horizontal and vertical variations of the input pixels 302 within the image region. Thus, the weighted parameters can be referred to as directional weighted parameters. As described above, the one or more weighted parameters can be considered to indicate the directionality of the filtering that will be applied when upsampling is applied to the input pixel values. For example, two weighted parameters (a and b) can be determined in step S504. The weighted parameters can be normalized such that a + b = 1. This means that as one of a or b increases, the other decreases. As will become apparent from the following description, if the parameters are set such that a = b = 0.5, the system will give the same output as a bilinear upsampler. However, in the system described herein, a and b can be different, i.e., a ≠ b can be obtained. Additionally, since b = 1 - a, the weighted parameter determination logic 408 can output an indication of a single weighted parameter (e.g., a), and this can be used to determine the second weighted parameter b as 1 - a. An indication of the one or more weighted parameters is provided to the pixel determination logic 406. More details of how the weighted parameter determination logic 408 determines the one or more weighted parameters are given in the example shown below. Figures 9 to 13 The example shown gives more details of the way in which the weighted parameter determination logic 408 determines the one or more weighted parameters.
[0167] In step S506, pixel determination logic 406 determines one or more of the upsampled pixel values of block 404 of upsampled pixels based on the one or more determined weighting parameters.
[0168] In step S508, block 404 of upsampled pixel values is output from pixel determination logic 406. In some systems, this can be the end of the processing of block 404 of upsampled pixel values, and this can be output from processing module 402, as Figure 4 shown. In other examples, for instance, as described below, sharpening can be applied to the upsampled pixel values (e.g., by blending with sharpened upsampled pixel values) before the upsampled pixel values are output from the processing module.
[0169] Figure 6 is a flowchart that shows how pixel determination logic 406 can determine unsharpened upsampled pixel values in step S506 based on the (multiple) weighting parameters, and thus determine upsampled pixel values, for example, based on the relative horizontal and vertical variations of the input pixel values within an image region. Figure 7 shows a portion of input pixel values 702 and shows how the positions of block 706 of upsampled pixel values are related to the positions of input pixels 704. Each of the upsampled pixel values determined in step S506 based on the relative horizontal and vertical variations of the input pixels within the image region (i.e., based on the one or more determined weighting parameters) is at a corresponding upsampled pixel position that does not correspond to the position of any input pixel. In other words, in step S506, the upper right and lower left upsampled pixel values in the block of upsampled pixel values are determined based on the relative horizontal and vertical variations of the input pixels within the image region (i.e., based on the one or more determined weighting parameters) (represented as “TR” and “BL” in Figure 7 ). If no sharpening is applied, the input pixel values of the input pixels that have a repeated five-point arrangement corresponding to the upsampled pixel positions are used as the upsampled pixel values at those upsampled pixel positions in the block of upsampled pixel values (e.g., the input pixels at the upper left and lower right positions of block 706 of upsampled pixel values are “passed through” directly to the corresponding positions in the output block). If sharpening is applied, some processing is performed on the input pixel values to determine all of the upsampled pixel values in the block of upsampled pixel values, as described in more detail below with reference to Figures 14 to 25 above.
[0170] As described above, input pixel values 704 are represented in two input blocks (e.g., 304 and 306). In Figure 7 the input pixel values 704 shown with line shading 1,1 、704 1,2 、704 1,3 and 704 1,4is part of the first input block 304; and the input pixel values 704 shown with cross - hatching 2,1 、704 2,2 、704 2,3 and 704 2,4 are part of the second input block 306.
[0171] As Figure 6 shown, the step S506 of determining each of the one or more upsampled pixel values among the upsampled pixel values includes applying one or more kernels to at least some of the input pixel values according to the determined one or more weighting parameters. Specifically, in step S602, the pixel determination logic 406 applies one or more first kernels (which may be referred to as horizontal kernels) to at least a first subset of the input pixel values to determine a horizontal component. For example, a kernel of [0.5, 0.5] can be applied to horizontally adjacent input pixel values and on either side of the upsampling pixel position for which the upsampled pixel value is being determined. For example, when performing step S602 on the upsampled pixel value denoted as "TR" in Figure 7 , the horizontal kernel is applied to the pixel values 704 1,1 and 704 1,2 from input block 1. When performing step S602 on the upsampled pixel value denoted as "BL" in Figure 7 , the horizontal kernel is applied to the pixel values 704 2,3 and 704 2,4 .
[0172] In step S604, the pixel determination logic 406 applies one or more second kernels (which may be referred to as vertical kernels) to at least a second subset of the input pixel values to determine a vertical component. For example, a kernel of can be applied to vertically adjacent input pixel values and on either side of the upsampling pixel position for which the upsampled pixel value is being determined. For example, when performing step S604 on the upsampled pixel value denoted as "TR" in Figure 7 , the vertical kernel is applied to the pixel values 704 2,2 and 704 2,4 from input block 2. When performing step S604 on the upsampled pixel value denoted as "BL" in Figure 7 , the vertical kernel is applied to the pixel values 704 1,1 and 704 1,3 from input block 1.
[0173] It should be noted that each of one or more upsampled pixel values to which step S506 is performed has two horizontally adjacent pixel values from the same input block (from input block 1 or input block 2), and has two vertically adjacent pixel values from the other input block (from input block 2 or input block 1). It should also be noted that steps S602 and S604 can be performed in any order, such as sequentially, or can be performed in parallel.
[0174] In steps S606 to S610, the pixel determination logic 406 combines the determined horizontal and vertical components such that each of the one or more upsampled pixel values among the upsampled pixel values is determined according to the relative horizontal and vertical variations of the input pixel values within the image region indicated by one or more weighting parameters determined in step S504. Specifically, in step S606, the pixel determination logic 406 multiplies the horizontal component (determined in step S602) by a first weighting parameter a among the weighting parameters to determine a weighted horizontal component.
[0175] In step S608, the pixel determination logic 406 multiplies the vertical component (determined in step S604) by a second weighting parameter b among the weighting parameters to determine a weighted vertical component. Steps S606 and S608 can be performed in any order, such as sequentially, or can be performed in parallel.
[0176] In step S610, the pixel determination logic 406 sums the weighted horizontal component and the weighted vertical component to determine the upsampled pixel value. In some examples, the kernels applied in S602 and S604 can be multiplied by the first weighting parameter and the second weighting parameter respectively before or during the application to the input pixel values. It is expected that this is less power and area efficient in terms of hardware than the method described in Figure 6 because Figure 8a and Figure 8b the application of the kernels of can be obtained using only inexpensive fixed shift and addition, and the number of variable multipliers required is much lower.
[0177] Figure 8a FIG. shows a 3x3 kernel 802 (which can be implemented as a 1x3 kernel) that can be applied to the input pixel value 702 to represent the effect of performing steps S602 and S606 on the upsampled pixel position being determined. For example, the kernel 802 can be centered on the upsampled pixel value being determined and applied to the input pixel value 702 such that the result will be given by the dot product where h is the vector of horizontally adjacent pixel values. For example, if the upsampled pixel value represented as "TR" is being determined, then h = [p 1,1 , p 1,2 , where p 1,1 is the input pixel value 7041,1 , and p 1,2 is the input pixel value 704 1,2 . Similarly, if the upsampled pixel value represented as "BL" is being determined, then h = [p 2,3 , p 2,4 , where p 2,3 is the input pixel value 704 2,3 , and p 2,4 is the input pixel value 704 2,4 .
[0178] Similarly, Figure 8b shows a 3x3 kernel 804 (which can be implemented as a 3x1 kernel) that can be applied to the input pixel value 702 to represent the effect of performing steps S604 and S608 on the upsampled pixel position being determined. For example, the kernel 804 can be centered on the upsampled pixel value being determined and applied to the input pixel value 702 such that the result will be given by the dot product , where v is a vector of vertically adjacent pixel values. For example, if the upsampled pixel value represented as "TR" is being determined, then v = [p 2,2 , p 2,4 , where p 2,2 is the input pixel value 704 2,2 , and p 2,4 is the input pixel value 704 2,4 . Similarly, if the upsampled pixel value represented as "BL" is being determined, then v = [p 1,1 , p 1,3 , where p 1,1 is the input pixel value 704 1,1 , and p 1,3 is the input pixel value 704 1,3 .
[0179] Therefore, steps S602 to S610 can be summarized as determining the upsampled pixel values represented as TR and BL as the sum of dot products: where h and v are vectors of adjacent pixel values for the TR pixel value and the BL pixel value, respectively, as described above.
[0180] The values of the weighting parameters a and b are set according to the local context such that, by determining the upsampled pixel values according to the determined weighting parameters in step S506, the upsampled pixel values are determined based on the relative horizontal and vertical variations of the input pixel values within the image region. In this way, the weighting parameters can be used to reduce artifacts and blurring that might otherwise be introduced by the upsampling process on anisotropic features (e.g., edges) in the image. As described above, in step S504, the weighting parameter determination logic 408 analyzes the input pixel values to determine one or more weighting parameters. In this way, one or more weighting parameters are determined for the particular image region being processed such that different weighting parameters can be used for the upsampling process of different image regions. In other words, the weighting parameters can be adjusted to adapt to the specific local context of the image region for which the upsampled pixel values are being determined. This allows appropriate weighting parameters to be used for different image regions based on the different anisotropic image features present in the different image regions. The following refers to Figures 9 to 13 the example shown gives more details of how the weighting parameter determination logic 408 determines one or more weighting parameters.
[0181] Figure 9 An example of the weighting parameter determination logic 408 configured to determine an indication of one or more weighting parameters is shown. As described above, the weighting parameter determination logic 408 is implemented on the processing module 402. The weighting parameter determination logic 408 includes a horizontal edge filtering logic 902, a vertical edge filtering logic 904, a horizontal line filtering logic 906, and a vertical line filtering logic 908, all of which are configured to receive input pixel values. The weighting parameter determination logic 408 further includes an implementation of a processing logic 910 and a neural network 912. The implementation of the neural network 912 is configured to receive input pixel values. The processing logic 910 is configured to receive filtering values from the four elements (902, 904, 906, and 908) of the filtering logic and receive an output (a residual value, as described below) from the implementation of the neural network 912. The processing logic 910 is configured to determine an indication of one or more weighting parameters using the filtering values (and in some examples, also using the residual value). Specifically, the processing logic 910 includes four blocks of absolute value determination logic 914, 916, 918, and 920, five blocks of summing logic 922, 926, 930, 932, and 936, three blocks of multiplication logic 924, 928, and 934, and a clamping logic 938.
[0182] Figure 10 is a flowchart of a method for determining an indication of the (multiple) weighting parameters using the weighting parameter determination logic 408. In step S1002, input pixel values are received at the weighting parameter determination logic 408. The input pixel values are the values of the input pixels in a repeated five-point arrangement whose positions correspond to the upsampled pixel positions, e.g., as described above with respect to Figure 2 and3 (which shows input pixel values in a repeated five-point pattern whose positions correspond to upsampled pixel positions). The repeated five-point pattern can be considered a "checkerboard" pattern. Further, as described above with reference to Figure 3 it, the input pixel values can be represented by two input blocks, where one of the two input blocks includes the input pixel values of the input pixels at positions within the odd rows of the repeated five-point pattern whose positions correspond to upsampled pixel positions, and the other of the two input blocks includes the input pixel values of the input pixels at positions within the even rows of the repeated five-point pattern whose positions correspond to upsampled pixel positions.
[0183] In steps S1004, S1006, S1008, and S1010, four filters (a horizontal edge filter, a vertical edge filter, a horizontal line filter, and a vertical line filter) are applied to the input pixel values. Steps S1004, S1006, S1008, and S1010 can be performed in any order, and / or two or more (or possibly all) of steps S1004, S1006, S1008, and S1010 can be performed in parallel. In the example described below, each filter has a filter kernel that has weights at the input pixel positions in a 5x5 region corresponding to the upsampled pixel positions, the upsampled pixel positions being centered at upsampled pixel positions that fall between the positions of adjacent input pixels in the five-point pattern, but in other examples, the filter kernel can have a different size and / or shape. Having a larger filter kernel (i.e., applying the filter to a larger set of pixel value positions) to derive orientation information from a wider region can be considered beneficial, but may risk drowning out signals from closer pixels (which may be more relevant). Thus, there is a trade-off in determining the size of the filter kernel. For example, the filter kernel can correspond to a 6x6 or 7x7 region of the upsampled pixel positions.
[0184] In step S1004, the horizontal edge filtering logic 902 applies a horizontal edge filter to two or more of the input pixel values to determine a first filtered value. In the case where there is an edge in the image, the image gradient is generally perpendicular to the orientation of the edge. For example, around a horizontal edge, the image gradient is generally vertical; around a vertical edge, the image gradient is generally horizontal. The edge filter is configured to identify edges within an image region, that is, to identify the image gradient in a specific direction within the image region. The edge filter represents an odd function. Specifically, the horizontal edge filter is configured to identify horizontal edges within the image region to which the horizontal edge filter is applied. Ideally, when the magnitude of the vertical image gradient within the image region is large, the horizontal edge filter will provide a response with a larger magnitude; and ideally, the horizontal edge filter will not depend on the horizontal image gradient within the image region. For example, the horizontal edge filter can have a 5x5 filter kernel, and its weights can be represented as:
[0185] 0.5 0.5 0 1 0 0 0 0 -1 0 -0.5 -0.5
[0186] As another example, the horizontal edge filter can have a filter kernel, and its weights can be represented as:
[0187] 0 0 0 1 0 0 0 0 -1 0 0 0
[0188] In step S1006, the vertical edge filtering logic 904 applies a vertical edge filter to two or more of the input pixel values to determine a second filtered value. The two or more input pixel values to which the vertical edge filter is applied in step S1006 may or may not be the same as the two or more input pixel values to which the horizontal edge filter is applied in step S1004. The vertical edge filter is configured to identify vertical edges within the image region to which the vertical edge filter is applied. Ideally, when the magnitude of the horizontal image gradient within the image region is large, the vertical edge filter will provide a response with a larger magnitude; and ideally, the vertical edge filter will not depend on the vertical image gradient within the image region. For example, the vertical edge filter can have a 5x5 filter kernel, and its weights can be represented as:
[0189] 0 0 -0.5 0 0.5 -1 1 -0.5 0 0.5 0 0
[0190] As another example, the vertical edge filter can have a filter kernel, and its weights can be represented as:
[0191] 0 0 0 0 0 -1 1 0 0 0 0 0
[0192] In step S1008, the horizontal line filtering logic 906 applies a horizontal line filter to three or more of the input pixel values to determine a third filtered value. The three or more input pixel values among the input pixel values to which the horizontal line filter is applied in step S1008 may or may not include the two or more input pixel values to which the horizontal edge filter is applied in step S1004 and / or the two or more input pixel values to which the vertical edge filter is applied in step S1006. The line filter is configured to identify lines within the image region, i.e., to identify the presence of an image gradient in a specific direction within the image region, where the image gradient varies significantly (e.g., changes sign) as the position within the image region varies in the specific direction. The line filter represents an even function. Specifically, the horizontal line filter is configured to identify horizontal lines within the image region to which the horizontal line filter is applied. Ideally, the horizontal line filter will provide a response with a relatively large magnitude when the vertical image gradient varies to a greater extent at different vertical positions within the image region. Specifically, ideally, the horizontal line filter will provide a response with a relatively large magnitude when there is a significant difference between the vertical image gradient towards the top of the image region and the vertical image gradient towards the bottom of the image region. For example, when the sign of the vertical image gradient towards the top of the image region is different from the sign of the vertical image gradient towards the bottom of the image region, the horizontal line filter will typically provide a response with a relatively large magnitude. However, ideally, the output of the horizontal line filter will not depend on the variation of the image gradient at different horizontal positions within the image region. Additionally, ideally, the output of the horizontal line filter will not depend on a constant image gradient in the horizontal direction. For example, the horizontal line filter may have a 5x5 filter kernel, and its weights may be represented as:
[0193] -0.5 -0.5 0 0 0 1 1 0 0 0 -0.5 -0.5
[0194] In step S1010, the vertical line filtering logic 908 applies a vertical line filter to three or more of the input pixel values to determine a fourth filtered value. The three or more input pixel values among the input pixel values to which the vertical line filter is applied in step S1010 may or may not be the same input pixel values as the three or more input pixel values to which the horizontal line filter is applied in step S1008, and may or may not include the two or more input pixel values to which the horizontal edge filter is applied in step S1004 and / or the two or more input pixel values to which the vertical edge filter is applied in step S1006. The vertical line filter is configured to identify vertical lines within the image region to which the vertical line filter is applied. Ideally, when the horizontal image gradient varies to a large extent at different horizontal positions within the image region, the vertical line filter will provide a response with a relatively large magnitude. Specifically, ideally, when there is a large difference between the horizontal image gradient towards the left side of the image region and the horizontal image gradient towards the right side of the image region, the vertical line filter will provide a response with a relatively large magnitude. For example, when the sign of the horizontal image gradient towards the left side of the image region is different from the sign of the horizontal image gradient towards the right side of the image region, the vertical line filter will generally provide a response with a relatively large magnitude. However, ideally, the output of the vertical line filter will not depend on the variation of the image gradient at different vertical positions within the image region. Additionally, ideally, the output of the vertical line filter will not depend on a constant image gradient in the vertical direction. For example, the vertical line filter may have a 5x5 filter kernel, and its weights can be represented as:
[0195] 0 0 -0.5 1 -0.5 0 0 -0.5 1 -0.5 0 0
[0196] The kernel has weights corresponding to the input pixel positions in a five-point arrangement. Applying the kernel to the input pixel values involves performing multiplications and additions (of the input pixel values with the corresponding kernel weights). It should be noted that in the example given above, many of the weights are zero. Specifically, applying each of the exemplary kernels shown above to the input pixel values involves performing six multiplications and performing additions. For example, implementing the multiplication operations and addition operations using multiply-accumulate (MAC) operations is very efficient. Specifically, the multiplication operations and addition operations can be implemented in hardware with extremely low silicon area, power consumption, and latency, such as in fixed-function circuits (e.g., in fused multiply-add (FMA) logic).
[0197] Horizontal edge filters, vertical edge filters, horizontal line filters, and vertical line filters have zero response to features in orthogonal directions. Configuring the edge filters to have zero response to features in orthogonal directions is relatively straightforward (i.e., it is relatively straightforward for a horizontal edge filter to have zero response to vertical features, and for a vertical edge filter to have zero response to horizontal features) because edge filters represent odd functions. In contrast, line filters represent even functions, and thus it is not as straightforward to configure a line filter such that it has zero response to features in orthogonal directions when applied to a five-point pattern of input pixel values. However, the line filters shown above do have zero response to features that are purely in orthogonal directions. For example, the horizontal line filter shown above is configured such that the sum of the weights for all columns is zero, and the vertical line filter shown above is such that the sum of the weights for all rows is zero. In other words, the horizontal line filter is configured such that when three or more of the input pixel values to which the horizontal line filter is applied exhibit a pure vertical feature, the third filtered value is determined to be zero, and wherein the vertical line filter is configured such that when three or more of the input pixel values to which the vertical line filter is applied exhibit a pure horizontal feature, the fourth filtered value is determined to be zero. For different vertical positions within an image region, pure vertical features such as vertical edges and vertical lines are constant, but vary for different horizontal positions within the image region. For different horizontal positions within an image region, pure horizontal features such as horizontal edges and horizontal lines are constant, but vary for different vertical positions within the image region.
[0198] In step S1012, processing logic 910 uses the first, second, third, and fourth filtered values (determined in steps S1004, S1006, S1008, and S1010) to determine an indication of one or more weighting parameters. As described above, the one or more weighting parameters indicate the relative horizontal and vertical variation of the input pixel values within the image region. It is apparent from the description herein that the one or more weighting parameters indicate the directionality of the filtering that will be applied when upsampling is applied to the input pixel values.
[0199] In step S1014, processing logic 910 outputs the determined indication of the one or more weighting parameters (e.g., the value of a). The indication of the one or more weighting parameters is output from weighting parameter determination logic 408 and provided to pixel determination logic 406 for applying upsampling to the input pixel values representing the image region to determine a block of one or more upsampled pixel values, as described herein.
[0200] In some examples, one or more filters in a filter, such as a line filter, can be implemented using multiple filter kernels. Different filter kernels can be applied to respective image regions, and the outputs of the filter kernels can be combined (e.g., by finding a weighted sum of the absolute values of the outputs of the filter kernels) to determine the output of the corresponding filter. In one example, a 6x6 filter kernel is applied to an image region centered on a 2x2 block of upsampled pixel positions, where the filter kernel is configured to determine the filtered values of the upper-right upsampled pixel position and the lower-left upsampled pixel position of the 2x2 block of upsampled pixel positions. For example, the following six filter kernels centered on a 2x2 block of upsampled pixel positions (shown with a bold border in the kernel below) can be applied to determine the filtered values of the vertical line filter for the upper-right upsampled pixel position (denoted as "TR" in the kernel below) and the lower-left upsampled pixel position (denoted as "BL" in the kernel below) of the 2x2 block of upsampled pixel positions:
[0201]
[0202] Kernel 0
[0203]
[0204] Kernel 1
[0205]
[0206] Kernel 2
[0207]
[0208] Kernel 3
[0209]
[0210] Kernel 4
[0211]
[0212] Kernel 5
[0213] The outputs from the six filter kernels shown above can be used, for example, to determine the filtered values of the vertical line filters for the upper right upsampled pixel position (denoted as "TR") and the lower left upsampled pixel position (denoted as "BL") of the 2x2 block of upsampled pixel positions by finding the weighted sum of the absolute values of the outputs of the filter kernels. For example, the filtered value of the vertical line filter for the upper right upsampled pixel position (denoted as "TR") of the 2x2 block of upsampled pixel positions can be determined by combining the outputs of kernels 0, 1, 3, and 5 (e.g., by setting the weights of the outputs from kernels 2 and 4 to zero in the weighted sum and setting the weights of the outputs from kernels 0, 1, 3, and 5 to non-zero in the weighted sum). The weights of the outputs from kernels 0 and 5 can be equal. The weights of the outputs from kernels 1 and 3 can be equal. The weights of the outputs from kernels 1 and 3 can be greater than the weights of the outputs from kernels 0 and 5. The sum of the weights of the outputs from all six kernels can be 1 such that the weighted sum is normalized.
[0214] As another example, the filtered value of the vertical line filter for the lower left sampled pixel position (denoted as "BL") of the 2x2 block of upsampled pixel positions can be determined by combining the outputs of kernels 0, 2, 4, and 5 (e.g., by setting the weights of the outputs from kernels 1 and 3 to zero in the weighted sum and setting the weights of the outputs from kernels 0, 2, 4, and 5 to non-zero in the weighted sum). The weights of the outputs from kernels 0 and 5 can be equal. The weights of the outputs from kernels 2 and 4 can be equal. The weights of the outputs from kernels 2 and 4 can be greater than the weights of the outputs from kernels 0 and 5. The sum of the weights of the outputs from all six kernels can be 1 such that the weighted sum is normalized.
[0215] In some examples, for both the upper right upsampled pixel position ("TR") and the lower left upsampled pixel position ("BL"), the same filtered value of the vertical line filter is determined. In these examples, due to the symmetry among kernels 0 to 5, the weights of the outputs from kernels 0 and 5 can be equal, the weights of the outputs from kernels 1 and 4 can be equal, and the weights of the outputs from kernels 2 and 3 can be equal.
[0216] The following six filter kernels (kernels 6 to 11), centered on the 2x2 block of upsampled pixel positions (shown with a bold border in the kernel below), can be applied to determine the filtered values of the horizontal line filters for the upper right upsampled pixel position (denoted as "TR") and the lower left upsampled pixel position (denoted as "BL") of the 2x2 block of upsampled pixel positions:
[0217]
[0218] Kernel 6
[0219]
[0220] Kernel 7
[0221]
[0222] Kernel 8
[0223]
[0224] Kernel 9
[0225]
[0226] Kernel 10
[0227]
[0228] Kernel 11
[0229] The outputs from the six filter kernels (kernels 6 to 11) shown above can be used to determine, for example, the filtering values of the horizontal line filters for the upper right upsampled pixel position (denoted as "TR") and the lower left upsampled pixel position (denoted as "BL") of the 2x2 block of upsampled pixel positions by finding the weighted sum of the absolute values of the outputs of the kernels. For example, the filtering value of the horizontal line filter for the upper right upsampled pixel position (denoted as "TR") of the 2x2 block of upsampled pixel positions can be determined by combining the outputs of kernels 6, 8, 10, and 11 (e.g., by setting the weights of the outputs from kernels 7 and 9 to zero in the weighted sum and setting the weights of the outputs from kernels 6, 8, 10, and 11 to non-zero in the weighted sum). The weights of the outputs from kernels 6 and 11 can be equal. The weights of the outputs from kernels 8 and 10 can be equal. The weights of the outputs from kernels 8 and 10 can be greater than the weights of the outputs from kernels 6 and 11. The sum of the weights of the outputs from all six kernels (kernels 6 to 11) can be 1 so that the weighted sum is normalized.
[0230] As another example, the filtering value of the horizontal line filter for the lower left upsampled pixel position (denoted as "BL") of the 2x2 block of upsampled pixel positions can be determined by combining the outputs of kernels 6, 7, 9, and 11 (e.g., by setting the weights of the outputs from kernels 8 and 10 to zero in the weighted sum and setting the weights of the outputs from kernels 6, 7, 9, and 11 to non-zero in the weighted sum). The weights of the outputs from kernels 6 and 11 can be equal. The weights of the outputs from kernels 7 and 9 can be equal. The weights of the outputs from kernels 7 and 9 can be greater than the weights of the outputs from kernels 6 and 11. The sum of the weights of the outputs from all six kernels (kernels 6 to 11) can be 1 so that the weighted sum is normalized.
[0231] In some examples, for both the top-right upsampled pixel position ("TR") and the bottom-left upsampled pixel position ("BL"), the same filtering value of the horizontal line filter is determined. In these examples, due to the symmetry between kernels 6 and 11, the weights of the outputs from kernels 6 and 11 can be equal, the weights of the outputs from kernels 7 and 10 can be equal, and the weights of the outputs from kernels 8 and 9 can be equal.
[0232] In this example, the edge filter can be simplified. The following filter kernels ("kernel 12" and "kernel 13") centered on the 2x2 block of upsampled pixel positions (shown with bold borders) can be applied to determine the filtering values of the vertical edge filter for the top-right upsampled pixel position (denoted as "TR") and the bottom-left upsampled pixel position (denoted as "BL") of the 2x2 block of upsampled pixel positions:
[0233]
[0234] Kernel 12
[0235]
[0236] Kernel 13
[0237] Specifically, the filtering values of the vertical edge filter for both the top-right upsampled pixel position and the bottom-left upsampled pixel position of the 2x2 block of upsampled pixel positions can be determined as the sum of the absolute values of the outputs of kernel 12 and the absolute values of the outputs of kernel 13.
[0238] The following filter kernels ("kernel 14" and "kernel 15") centered on the 2x2 block of upsampled pixel positions (shown with bold borders) can be applied to determine the filtering values of the horizontal edge filter for the top-right upsampled pixel position (denoted as "TR") and the bottom-left upsampled pixel position (denoted as "BL") of the 2x2 block of upsampled pixel positions:
[0239]
[0240] Kernel 14
[0241]
[0242] Kernel 15
[0243] Specifically, the filtering values of the horizontal edge filter for both the top-right upsampled pixel position and the bottom-left upsampled pixel position of the 2x2 block of upsampled pixel positions can be determined as the sum of the absolute values of the outputs of kernel 14 and the absolute values of the outputs of kernel 15.
[0244] Figure 11An example of how step S1012 can be performed is shown. Specifically, step S1012 may include steps S1102, S1104, and S1106, where steps S1102 and S1104 may be performed in any order or in parallel.
[0245] In step S1102, processing logic 910 performs a weighted sum of a first filter value, a second filter value, a third filter value, and a fourth filter value. In other words, processing logic 910 combines the first filter value, the second filter value, the third filter value, and the fourth filter value by performing a weighted sum. The processing logic determines the absolute values of the first filter value, the second filter value, the third filter value, and the fourth filter value because the sign of the filter values is irrelevant to the indication of determining the (multiple) weighted parameters. Specifically, absolute value determination logic 914 determines the absolute value of the first filter value determined by horizontal edge filtering logic 902; absolute value determination logic 916 determines the absolute value of the second filter value determined by vertical edge filtering logic 904; absolute value determination logic 918 determines the absolute value of the third filter value determined by horizontal line filtering logic 906; and absolute value determination logic 920 determines the absolute value of the fourth filter value determined by vertical line filtering logic 908. Summing logic 922 subtracts the absolute value of the second filter value from the absolute value of the first filter value, and multiplication logic 924 multiplies the output from summing logic 922 by parameter β. Summing logic 926 subtracts the absolute value of the fourth filter value from the absolute value of the third filter value, and multiplication logic 928 multiplies the output from summing logic 926 by parameter δ. Summing logic 930 sums the output from multiplication logic 924 and the output from multiplication logic 928. The output from summing logic 930 represents the weighted sum (S) of the absolute values of the four filter values. Specifically, S = β|h e |- β|v e | + δ|h l |- δ|v l |, where h e is the first filter value (i.e., the result of applying the horizontal edge filter to two or more of the input pixel values in step S1004), v e is the second filter value (i.e., the result of applying the vertical edge filter to two or more of the input pixel values in step S1006), h l is the third filter value (i.e., the result of applying the horizontal line filter to three or more of the input pixel values in step S1008), and v l is the fourth filter value (i.e., the result of applying the vertical line filter to three or more of the input pixel values in step S1010).
[0246] In the above example, one or more weighted parameter indications output by the weighted parameter determination logic 408 will be applied to determine the upsampled pixel value at the upsampled pixel position at the center of the filter kernel. In some other examples, the indication of one or more weighted parameters output by the weighted parameter determination logic 408 can be applied to determine the upsampled pixel values at multiple output positions (e.g., both the upper-right output position and the lower-left output position). In these examples, the filtering logic (902, 904, 906, and 908) can be applied multiple times centered on each of the output positions. In these examples, each block of the filtering logic (902, 904, 906, and 908) provides multiple (e.g., two) filtered results to its corresponding absolute value determination logic multiple times, which finds the absolute values of the multiple (e.g., two) filtered values, sums the absolute values together, and then outputs the sum result. For example, referring to Figure 9 , in these examples:
[0247] - The horizontal edge filtering logic 902 applies a horizontal edge filter to the input pixel values centered on multiple (e.g., two) output pixel positions to determine multiple (e.g., two) filtered values; and
[0248] The absolute value determination logic 914 finds the absolute values of these multiple (e.g., two) filtered values, sums the absolute values, and provides the sum result to the summation logic 922;
[0249] - The vertical edge filtering logic 904 applies a vertical edge filter to the input pixel values centered on multiple (e.g., two) output pixel positions to determine multiple (e.g., two) filtered values; and
[0250] The absolute value determination logic 916 finds the absolute values of these multiple (e.g., two) filtered values, sums the absolute values, and provides the sum result to the summation logic 922;
[0251] - The horizontal line filtering logic 906 applies a horizontal line filter to the input pixel values centered on multiple (e.g., two) output pixel positions to determine multiple (e.g., two) filtered values; and the absolute value determination logic 918 finds the absolute values of these multiple (e.g., two) filtered values, sums the absolute values, and provides the sum result to the summation logic 926; and
[0252] - The vertical line filtering logic 908 applies a vertical line filter to the input pixel values centered on multiple (e.g., two) output pixel positions to determine multiple (e.g., two) filtered values; and the absolute value determination logic 920 finds the absolute values of these multiple (e.g., two) filtered values, sums the absolute values, and provides the sum result to the summation logic 926.
[0253] Then, the method continues as described above in these examples.
[0254] In step S1104, an implementation of neural network 912 processes the input pixel values to determine a residual value. In step S1106, processing logic 910 combines the determined residual value with a first filter value, a second filter value, a third filter value, and a fourth filter value to determine an indication of one or more weighting parameters. Specifically, summing logic 932 sums the result of the weighted sum determined by summing logic 930 and the residual value determined by the implementation of neural network 912. Then, in Figure 9 the example shown, multiplication logic 934 multiplies the output of summing logic 932 by parameter α, and summing logic 936 adds 0.5 to the output of multiplication logic 934.
[0255] Before output, clamping logic 938 clamps the indication of one or more weighting parameters to be within the range [0,1]. Then, in step S1014, an indication of one or more weighting parameters (a) is output.
[0256] The values of parameters α, β, and δ can be fixed and can be determined in advance before runtime. For example, they can be hardcoded into the hardware of processing logic 910. Alternatively, one or more of the values of parameters α, β, and δ can be variable and can be set in software. Additionally, one or more of the values of parameters α, β, and δ can be trained (e.g., using backpropagation techniques) such that the indication of one or more weighting parameters indicates one or more weighting parameters that indicate relative horizontal and vertical variations of the input pixel values within the indicated image region.
[0257] For example, the entire system can be trained end-to-end. As part of the training process, a "loss" can be specified based on the difference between the upsampled output pixel values and the ground truth pixel values. The difference can be measured using (e.g.) the L1 or L2 norm. The image can be used as input to the upsampling algorithm by extracting a checkerboard pattern of pixels for training. For the purpose of comparison with the upsampled output, the full image can be used as the ground truth. Typically, the images will be from the target application (e.g., a rendered frame), but synthetic images can also be used to provide examples of specific features, such as edges at a large number of orientations. The value of α can also be configurable such that it can be changed from its trained value. This allows the user to adjust the width of the transition region between full interpolation in one direction and full interpolation in the other direction. This can be useful if intermediate interpolation directions result in staircase artifacts.
[0258] The neural network can be a convolutional neural network. Figure 12An embodiment of a neural network 912 configured to output a residual value is shown. The embodiment of the neural network 912 can be implemented on any suitable hardware, such as a GPU or a neural network accelerator (NNA), or implemented as fixed-function hardware with predetermined fixed weights. The neural network has been trained (e.g., using quantization-aware training (QAT)) to determine the residual value such that one or more weighted parameters indicate the directionality of the filtering to be applied when upsampling is applied to the input pixel values. Specifically, the embodiment of the neural network 912 is configured to receive the input pixel values 1202, process the input pixel values to determine the residual value, and output the determined residual value for the processing logic 910 to use to determine an indication of one or more weighted parameters of the image region represented by the input pixel values. As described above, the input pixel values can be represented by two input blocks, where one of the two input blocks includes the input pixel values at positions within the odd rows of a repeating five-point arrangement whose positions correspond to the upsampled pixel positions, and the other of the two input blocks includes the input pixel values at positions within the even rows of a repeating five-point arrangement whose positions correspond to the upsampled pixel positions. In Figure 12 In the example shown, each input block is a 3x3 block of input pixel values such that the input pixel values 1202 are shown as a 3x3x2 array to indicate that the input pixel values are organized as two 3x3 input blocks of input pixel values.
[0259] As described above, the neural network is trained to output the residual value instead of trying to train the neural network to output an indication of the (multiple) weighted parameters or even the upsampled pixels themselves. This is because the neural network is well-suited to determining the appropriate residual value to make a minor adjustment to the weighted sum of four filter values, and this task can be achieved with a small neural network. Therefore, the embodiment of the neural network 912 has a lower complexity, and thus it can have a small silicon area as well as low latency and power consumption.
[0260] Specifically, an implementation of the neural network 912 includes a first convolutional layer 1206, a first activation function 1208, a second convolutional layer 1214, a second activation function 1216, and a third convolutional layer 1220. It should be noted that convolutional layers and activation functions are simple to implement. The first activation function 1208 is implemented between the first convolutional layer 1206 and the second convolutional layer 1214, and the second activation function 1216 is implemented between the second convolutional layer 1214 and the third convolutional layer 1220. The activation functions 1208 and 1216 can be, for example, rectified linear units (ReLU), absolute value functions, or identity functions. It should be noted that implementing an identity function is equivalent to not implementing a function. The "absolute value function" outputs the absolute value of its input. In a first example, both the first activation function 1208 and the second activation function 1216 are rectified linear units (ReLU). In a second example, the first activation function 1208 is an identity function, and the second activation function 1216 is an absolute value function. It should be noted that ReLU introduces non-linearity. Thus, it is easier to analyze the combined effect of the first two convolutional layers in the second example than in the first example, because no non-linearity is implemented between the first convolutional layer and the second convolutional layer in the second example. The neural network can be implemented in hardware (e.g., fixed function circuitry), software, or a combination thereof.
[0261] Figure 13 FIG. 4 is a flow chart of a method for determining a residual value using the implementation of the neural network 912 in the first example mentioned above, in which the first activation function and the second activation function are ReLU units. In step S1302, an input pixel value 1202 is received at the implementation of the neural network 912.
[0262] In step S1304, the first convolutional layer 1206 is used to determine a first intermediate tensor based on the input pixel value 1202. In step S1306, the first ReLU 1208 is used to set negative values in the first intermediate tensor to zero. In fixed function hardware, the first ReLU 1208 can be easily implemented by combining with the previous convolution, for example, by performing an "OR" operation on the result of the previous convolution and the "NOT" of the sign bit. Figure 12 FIG. 9 shows the first intermediate tensor 1212 output from the first ReLU 1208. The first intermediate tensor 1212 extends in only one of the horizontal and vertical dimensions with respect to the dimensions of the image region represented by the input pixel 1202. In Figure 12 the example shown, the first intermediate tensor 1212 extends only in the horizontal dimension (rather than the vertical dimension), for example, when processing is performed along the rows of the image. However, in other examples, the first intermediate tensor can extend only in the vertical dimension (rather than the horizontal dimension), for example, if processing is to be performed downward along the columns of the image. In Figure 12In the example shown, the first intermediate tensor 1212 is a 3x1x8 array.
[0263] In step S1308, the second convolutional layer 1214 is used to determine a second intermediate tensor based on the first intermediate tensor 1212. To save computation, this layer can be a depthwise convolution, i.e., processing each channel of the input separately, with no cross-channel component in the operations. In step S1310, the second ReLU 1216 is used to set negative values in the second intermediate tensor to zero. Figure 12 The second intermediate tensor 1218 output from the second ReLU 1216 is shown in. The dimensions of the second intermediate tensor with respect to the image region represented by the input pixel values 1202 do not extend in the horizontal or vertical dimensions. In Figure 12 the example shown, the second intermediate tensor 1218 is a 1x1x8 array.
[0264] In step S1312, the third convolutional layer 1220 is used to determine a residual value 1222 based on the second intermediate tensor 1218. The residual value 1222 is shown in Figure 12 as a 1x1x1 array (i.e., a scalar value for each input region).
[0265] In an implementation of the neural network 912, splitting the processing into separate vertical, horizontal, and channel passes (in the first convolutional layer, second convolutional layer, and third convolutional layer respectively) keeps the number of multiply-accumulate (MAC) units low. In this example, the first convolution 1206 performs 1x3x2x8 = 48 MACs (multiply-accumulates) per output block; the second convolution 1214 performs 3x1x8x1 = 24 MACs per output block; and the final convolution 1220 performs 8 MACs per output block. So overall, this network performs 80 MACs per output block, which is significantly lower than a typical upsampling neural network, which typically requires many more orders of magnitude of operations to achieve the same output resolution. By removing unused channels from the trained network and using low-precision integer arithmetic with QAT, significant cost savings can be further achieved. Many different techniques (two examples are, for instance, post-training quantization (PTQ) and manual format selection) can be used to apply quantization to the network. When implemented in fixed-function hardware, such savings directly translate into lower area, power, and latency requirements.
[0266] In step S1314, the residual value is output from the implementation of the neural network 912 and provided to the processing logic 910 for use in step S1106 as described above.
[0267] As described above, Figure 13Shows a first example where the first activation function and the second activation function are ReLU units. The method will proceed in a similar manner in a second example where the first activation function is the identity function and the second activation function is the absolute value function, except that in step S1306, the identity function will simply pass the first intermediate tensor unchanged, and in step S1310, the absolute value function will set the values in the second intermediate tensor to their absolute values (i.e., the sign of any negative value in the second intermediate tensor will be switched to positive while the magnitude of the values in the second intermediate tensor remains unchanged).
[0268] In some examples, the image region is a part of the input image, where upsampling will be performed iteratively for multiple partially overlapping image regions within the input image. In this sense, each of the image regions can be regarded as a window sliding across the input image. In these cases, some of the input pixel values 1202 of the current image region are the same as some of the input pixel values of the previous image region. The "previous image region" can be the image region processed in the previous iteration. In these examples, processing the input pixel values with an implementation of the neural network 912 may also include storing the first intermediate tensor 1212 in a buffer acting as a queue / shift register. This storage step can be performed between steps S1306 and S1308 in the flowchart of Figure 13 . The first convolutional layer 1206 can operate on the input pixel values (shown shaded and denoted as 1204 in Figure 12 ) within the first part of the current image region that does not overlap with the previous image region, but does not need to operate on the input pixel values (i.e., the remaining part of the input pixels 1202 not shown shaded in Figure 12 ) within the second part of the current image region that overlaps with the previous image region. Storing the first intermediate tensor 1212 using the buffer means that the first convolutional layer 1206 can avoid performing a full convolution on all the input pixel values and can instead reuse the values determined in the previous iteration. At each iteration, the window can slide one pixel position in both input blocks, such that the second part of the current image region 904 is a line (e.g., a column or a row) of pixel values in both input blocks. This means that the overlapping region (i.e., the first part of the current image region 1204) is a 3x2 block of input pixel values in each input block. It can be understood that in this case, most of the input pixel values of the current image region overlap with the input pixel values of the previous image region. Additionally, since in Figure 12In the example shown, the second part of the image region 1204 is a column of input pixel values from each input block, and since the first convolutional layer 1206 removes a vertical extent from the first intermediate tensor 1212 but does not remove a horizontal extent from the first intermediate tensor 1212, the first convolutional layer does not need to operate on the input pixel values in the first part of the image region. It should be noted that in other examples, the window may slide vertically instead of horizontally, such that the non-overlapping portions of the input pixel values are a row of input pixel values from each input block, and in such cases, the first convolutional layer may remove a horizontal extent from the first intermediate tensor but not remove a vertical extent from the first intermediate tensor, such that the first convolutional layer will again not need to operate on the input pixel values in the overlapping portions of the input pixel values.
[0269] The implementation of the neural network does not include any line banks. A line bank stores lines (e.g., rows) of pixel values and requires large memory. By avoiding including any line banks in the implementation of the neural network 912, the silicon area and power consumption of the implementation of the neural network can be kept low.
[0270] Furthermore, as described above, the neural network can be trained using quantization-aware training (QAT). In this way, the neural network can be trained for sparsity and low bit-depth. Any MACs for missing weights can be skipped (i.e., not synthesized). The implementation of the neural network 912 does not need to use the same bit-depth for all layers or weights within a layer, and the QAT scheme can be used to appropriately optimize the layers and weights of the neural network.
[0271] It should be noted that Figure 9 An example of an implementation in which the weighted parameter determination logic 408 includes the neural network 912 (which operates as described above with reference to Figures 9 to 13 is described). However, in other examples, the weighted parameter determination logic may not include an implementation of the neural network, and the processing logic may determine an indication of the (multiple) weighted parameters based only on the four filter values it receives from the four blocks of the filter logics 902, 904, 906, and 908. In these other examples, the weighted parameter determination logic may be the same as that shown in Figure 9 except that it does not include an implementation of the neural network 912 or the summation logic 932.
[0272] In the above examples, one or more of the upsampled pixel values in the block 404 of upsampled pixel values are unsharpened upsampled pixel values, i.e., no sharpening is applied when upsampling is performed. In these examples where no sharpening is applied, as described above, step S602 involves applying a single first kernel to the input pixel values horizontally adjacent to the upsampled pixel value being determined to determine the horizontal component, and step S604 involves applying a single second kernel to the input pixel values vertically adjacent to the upsampled pixel value being determined to determine the vertical component. However, in other examples described in detail below with reference to Figures 14 to 25 the upsampled pixel values in the block of upsampled pixel values are sharpened upsampled pixel values, i.e., sharpening is applied when upsampling is performed.
[0273] Figure 14 FIG. shows pixel determination logic 1400 configured to determine sharpened upsampled pixel values. The pixel determination logic 1400 may be implemented on a processing module (e.g., Figure 4 the processing module 402 shown in ). The pixel determination logic 1400 is configured to receive input pixel values 1402, process the input pixel values, and output a block 1410 of sharpened upsampled pixel values. As described above, the input pixel values 1402 are represented by two input blocks (represented as 14041 and 14042 in Figure 14 ) corresponding to the above input blocks 304 and 306. The pixel determination logic 1400 includes six pairs of kernels 1406, where each pair of kernels 1406 includes one kernel to be applied to the first input block 14041 and another kernel to be applied to the second input block 14042. The pixel determination logic 1400 also includes two blocks of weighted sum logic 1408. The logic of the processing module including the pixel determination logic 1400 may be implemented in hardware, software, or a combination thereof. Compared with the software implementation, the hardware implementation generally provides reduced latency and power consumption. The pixel determination logic 1400 may be used many times on each image being upsampled in the same manner, and since latency is very important in, for example, real-time super-resolution applications, implementing the pixel determination logic 1400 in hardware (e.g., in fixed-function circuitry) may be more preferable than implementing the logic in software. However, a software implementation is still possible and may be preferable in some cases.
[0274] Figure 15 is a flowchart of a method for determining a block of sharpened upsampled pixel values using the pixel determination logic 1400. In step S1502, two input blocks 14041 and 14042 of input pixel values are received at the pixel determination logic 1400. Then, the pixel determination logic 1400 determines each sharpened upsampled pixel value in the block 1410 of sharpened upsampled pixel values by performing steps S1504, S1506, and S1508.
[0275] For each sharpened upsampled pixel value in block 1410 that sharpens upsampled pixel values, pixel determination logic 1400: (i) in step S1504, applies a first set of one or more kernels to the input pixel values of first input block 14041; (ii) in step S1506, applies a second set of one or more kernels to the input pixel values of second input block 14042; and (iii) in step S1508, combines the results of applying the first set of one or more kernels and applying the second set of one or more kernels to determine the sharpened upsampled pixel value. Each of kernels 1406 is configured to perform upsampling and sharpening. In other words, sharpening and upsampling (e.g., interpolation) are performed in a single linear kernel. This is possible because these two operations (interpolation and sharpening) are linear operations and thus can be composed (or "folded") into a single linear operation, which is advantageous because it avoids the need for an expensive data intermediate buffer prior to a separate sharpening pass. One of the kernels in the first set of kernels (applied to the input pixel values of first input block 14041) has a first subset of the values of the full sharpened upsampling kernel, and one of the kernels in the second set of kernels (applied to the input pixel values of second input block 14042) has a second subset of the values of the full sharpened upsampling kernel. As described below, the values of the full sharpened upsampling kernel correspond to upsampled pixel positions and represent the result of convolving an interpolation kernel with a sharpening kernel. The first subset of the values of the full sharpened upsampling kernel (in the one of the kernels in the first set of kernels applied to the input pixel values of first input block 14041) includes the values of the full sharpened upsampling kernel at positions corresponding to the input pixel values in first input block 14041; and the second subset of the values of the full sharpened upsampling kernel (in the one of the kernels in the second set of kernels applied to the input pixel values of second input block 14042) includes the values of the full sharpened upsampling kernel at positions corresponding to the input pixel values in second input block 14042.
[0276] In step S1510, block 1410 outputs the sharpened upsampled pixel values from pixel determination logic 1400.
[0277] To determine the values of the full sharpened upsampling kernel at positions corresponding to the input pixel values of the first input block or the second input block (e.g., the top left upsampled pixel value (represented as "TL" in Figure 14 and the bottom right upsampled pixel value (represented as Figure 14Determine the sharpened upsampled pixel value at the position (denoted as "BR" in the text): (i) The first set of one or more kernels (applied to the input pixel values of the first input block 14041 in step S1504) includes a first single kernel applied to the input pixel values of the first input block 14041, (ii) The second set of one or more kernels (applied to the input pixel values of the second input block 14042 in step S1506) includes a second single kernel applied to the input pixel values of the second input block 14042; and (iii) The combination of the results in step S1508 includes summing the result of applying the first single kernel to the input pixel values of the first input block 14041 and the result of applying the second single kernel to the input pixel values of the second input block 14042. For these cases, that is, to determine the sharpened upsampled pixel value at the position corresponding to the position of the input pixel values of the first input block or the second input block, the interpolation kernel convolved with the sharpening kernel is a bilinear interpolation kernel, which is configured to be used with the input blocks described herein, and the input block represents the input pixel values in a repeating pentagon arrangement.
[0278] An example of a bilinear interpolation kernel used with input pixel values whose positions correspond to a repeating pentagon arrangement can be represented as If this bilinear interpolation kernel 1602 is applied and centered at the position of one of the input pixel values in the repeating pentagon arrangement, it is equivalent to applying an identity kernel at that position (since the values horizontally and vertically adjacent to that position in the repeating pentagon arrangement are zero), and the effect is to "pass through" that value to the output. If the bilinear interpolation kernel 1602 is applied and centered at a position between the positions of the input pixel values in the repeating pentagon arrangement, it is equivalent to applying the average of four adjacent input pixel values (since the kernel is centered at zero values in the repeating pentagon arrangement).
[0279] As described above, the value of the complete sharpened upsampled kernel represents the result of convolving an interpolation kernel (e.g., the bilinear interpolation kernel 1602) with a sharpening kernel. In different examples, the sharpening kernel can be different. For example, the sharpening kernel can be an unsharp mask kernel or a Lanczos kernel, both of which are known sharpening kernels.
[0280] Figure 16 The first complete sharpened upsampled kernel 1606, which is the result of convolving the bilinear interpolation kernel 1602 with the sharpening kernel 1604, is shown. Convolving this complete sharpened upsampled kernel with the input pentagon image and replacing the missing input pixel values with zeros will produce a sharpened bilinear interpolation image. In this example, the sharpening kernel 1604 is an unsharp mask kernel. In Figure 16In this case, the bilinear interpolation kernel 1602 and the sharpening kernel 1604 are 5x5 kernels. The sharpening kernel 1604 has a large positive value at the center position (e.g., a value greater than 1, represented as "++++" in Figure 16 ), and has small negative values near the center position, and the magnitude of the small negative values becomes smaller as the distance from the center position increases (represented as "--" and "-" in Figure 16 ). Figure 16 The exact values are not shown in the sharpening kernel 1604 or the complete sharpened upsampling kernel 1606, but it should be understood that positive values are shown with a "+" symbol and negative values are shown with a "-" symbol, where the value represented as "++++" has a greater magnitude than the value represented as "++", and where the value represented as "--" has a greater magnitude than the value represented as "-". It should be understood that this notation indicates the general shape of a linear space-invariant sharpening kernel, and variations in this regard are possible, such as a more "spread out" central maximum.
[0281] Figure 17 Illustrated is how the complete sharpened upsampling kernel 1606 can be decomposed into: (i) a kernel 1704 that is applied to a first input block of input pixel values 304, and (ii) a kernel 1706 that is applied to a second input block of input pixel values 306 to determine the sharpened upsampling pixel value in the upper left corner of the block 404 of sharpened upsampling pixel values. The position of the sharpened upsampling pixel value in the upper left corner of the block 404 of sharpened upsampling pixel values corresponds to one of the input pixel values 1702, specifically one of the input pixel values in the first input block 304. Conceptually, the complete sharpened upsampling kernel 1706 would be applied to the input pixel values. As described above, the positions of the input pixel values correspond to a repeating pentagon arrangement (i.e., a checkerboard pattern) of upsampled pixel positions, e.g., as shown by 302 in Figure 3 and Figure 17 . If the complete sharpened upsampling kernel 1606 is applied to the input pixel values 302 with a repeating pentagon arrangement, the values of the complete sharpened upsampling kernel 1606 corresponding to the zero values in the repeating pentagon arrangement of the input pixel values will not contribute to the result and can therefore be ignored. As described above with reference to Figure 3 , the input pixel values can be received in two input blocks 304 and 306, where the first input block 304 includes the input pixel values of the pixels at positions within the odd rows corresponding to the repeating pentagon arrangement of the upsampled pixel positions (shown with line shading), and the second input block 306 includes the input pixel values of the pixels at positions within the even rows corresponding to the repeating pentagon arrangement of the upsampled pixel positions (shown with cross shading).
[0282] If the kernel 1606 is applied to the input pixel 302 to determine the sharpened upsampled pixel value in the upper left corner of the block 404 for the sharpened upsampled pixel values, the kernel 1606 will be centered on the input pixel value 1702, the values of the kernel 1606 will be multiplied by the input pixel values at the corresponding surrounding positions, and the results will be summed. It can be seen that doing so will mean that the values of the kernel 1606 will be multiplied by Figure 17 some but not all of the input pixel values of the first input block 304 shown with line shading. The kernel 1704 includes those values of the kernel 1606 that will be multiplied by the input pixel values of the first input block 304. In addition, the values of the kernel 1606 will be multiplied by Figure 17 some but not all of the input pixel values of the second input block 306 shown with cross - shading, and the kernel 1706 includes those values of the kernel 1606 that will be multiplied by the input pixel values of the second input block 306. Thus, applying the kernel 1606 to the input pixel value 302 to determine the sharpened upsampled pixel value in the upper left corner of the block 404 for the sharpened upsampled pixel values will give the same result as summing the results of applying the kernel 1704 to the input block 304 and applying the kernel 1706 to the input block 306.
[0283] Therefore, to determine the sharpened upsampled pixel value in the upper left corner of the block 404 for the sharpened upsampled pixel values, step S1504 includes applying the kernel 1704 to the first input block 304, step S1506 includes applying the kernel 1706 to the second input block 306, and step S1508 includes summing the results of steps S1504 and S1506. The kernels 1704 and 1706 correspond to Figure 14 a pair of kernels 1406 in TL .
[0284] Figure 17 shows zero - padding the kernels 1704 and 1706 up to a 5x5 kernel. In practical applications, the zero - padding may not be implemented in hardware (since multiplying by a fixed zero is always zero). Figure 17 The zero - padding shown in Figure 17 is for the purpose of this explanation to be clear, as a way to indicate the correct alignment of the kernels relative to the 5x5 input blocks to which the kernels will be applied. In other words, the kernel 1704 can be implemented as a 3x3 kernel (where the zero - padding around the top edge, bottom edge, left edge, and right edge of the 5x5 kernel 1704 shown in Figure 17 is removed), and then it can be applied to the central 3x3 region of the 5x5 input block 304. Similarly, the kernel 1706 can be implemented as a 4x4 kernel (where the zero - padding around the bottom edge and right edge of the 5x5 kernel 1706 shown in
[0285] Figure 18 Shows how the complete sharpened upsampling kernel 1606 can be decomposed into: (i) a kernel 1804 applied to a first input block of the input pixel values 304, and (ii) a kernel 1806 applied to a second input block of the input pixel values 306 to determine the sharpened upsampling pixel value in the lower right corner of the block 404 of sharpened upsampling pixel values. The position of the sharpened upsampling pixel value in the lower right corner of the block 404 of sharpened upsampling pixel values corresponds to one of the input pixel values 1802, specifically one of the input pixel values of the second input block 306.
[0286] If the kernel 1606 is applied to the input pixel 302 to determine the sharpened upsampling pixel value in the lower right corner of the block 404 of sharpened upsampling pixel values, the kernel 1606 will be centered on the input pixel value 1802, the values of the kernel 1606 will be multiplied by the input pixel values at the corresponding surrounding positions, and the results will be summed. It can be seen that doing so will mean that the values of the kernel 1606 will be multiplied by Figure 18 some but not all of the input pixel values of the first input block 304 shown with line shading. The kernel 1804 includes those values of the kernel 1606 that will be multiplied by the input pixel values of the first input block 304. In addition, the values of the kernel 1606 will be multiplied by Figure 18 some but not all of the input pixel values of the second input block 306 shown with cross shading, and the kernel 1806 includes those values of the kernel 1606 that will be multiplied by the input pixel values of the second input block 306. Therefore, applying the kernel 1606 to the input pixel value 302 to determine the sharpened upsampling pixel value in the lower right corner of the block 404 of sharpened upsampling pixel values will give the same result as summing the results of applying the kernel 1804 to the input block 304 and applying the kernel 1806 to the input block 306.
[0287] Therefore, to determine the sharpened upsampling pixel value in the lower right corner of the block 404 of sharpened upsampling pixel values, step S1504 includes applying the kernel 1804 to the first input block 304, step S1506 includes applying the kernel 1806 to the second input block 306, and step S1508 includes summing the results of steps S1504 and S1506. The kernels 1804 and 1806 correspond to Figure 14 a pair of kernels 1406 BR .
[0288] Figure 18 Shows zero-padding the kernels 1804 and 1806 up to a 5x5 kernel. As described above, in practical applications, zero-padding may not be implemented in hardware (since multiplying by a fixed zero is always zero). Figure 18The zero-padding shown is for clarity of this explanation and serves as a way to indicate the correct alignment of the kernel relative to the 5x5 input block to which the kernel will be applied. In other words, kernel 1804 can be implemented as a 4x4 kernel (where Figure 18 the zero-padding around the top and left edges of the 5x5 kernel 1904 shown in Figure 18 is removed), and then it can be applied to the lower-right 4x4 region of the 5x5 input block 304. Similarly, kernel 1806 can be implemented as a 3x3 kernel (where
[0289] Figure 17 and Figure 18 the zero-padding around the top, bottom, left, and right edges of the 5x5 kernel 1806 shown in
[0290] is removed), and then it can be applied to the central 3x3 region of the 5x5 input block 306. Figure 19 shows the complete horizontal sharpened upsampling kernel 1904 as a result of convolving the linear horizontal interpolation kernel with the sharpening kernel. In this example, the sharpening kernel 1604 is an unsharp mask kernel. In Figure 19 , the linear horizontal interpolation kernel 1902 and the sharpening kernel 1604 are 5x5 kernels. Specifically, the linear horizontal interpolation kernel, described above as a 1x3 kernel, is padded up to the Figure 19 5x5 kernel 1902 shown in Figure 16 by positioning the values of the 1x3 linear horizontal interpolation kernel at the center of the 5x5 kernel and adding zero values to pad it up to the 5x5 kernel 1902. As described above, the sharpening kernel 1604 has a large positive value at the center position (e.g., a value greater than 1, represented as "++++" inFigure 16 represented as "--" and "-" in
[0291] Similarly, a vertical sharpened upsampling kernel that is the result of convolving a linear vertical interpolation kernel with a sharpening kernel can be determined. For example, the linear vertical interpolation kernel can be represented as Figure 20 shows the complete vertical sharpened upsampling kernel 2004 that is the result of convolving the linear vertical interpolation kernel 2002 with the sharpening kernel 1604. In this example, the sharpening kernel 1604 is an unsharp masking kernel. In Figure 20 the linear vertical interpolation kernel 2002 and the sharpening kernel 1604 are 5x5 kernels. Specifically, the linear vertical interpolation kernel described above as a 3x1 kernel is padded up to the 5x5 kernel 2002 shown in Figure 20 by positioning the values of the 3x1 linear vertical interpolation kernel at the center of the 5x5 kernel and adding zero values to pad it up to the 5x5 kernel 2002. It should be noted that the complete vertical sharpened upsampling kernel 2004 is the same as the complete horizontal sharpened upsampling kernel 1904 but rotated by 90°.
[0292] To determine a sharpened upsampling pixel value at a position that falls between the corresponding positions of the input pixel values of the first input block and the second input block (e.g., at the upper right corner or the lower left corner of the block 404 of the sharpened upsampling pixel values), step S1504 includes applying a first horizontal kernel and a first vertical kernel to the input pixel values of the first input block 304 (i.e., the first set of one or more kernels applied in step S1504 includes the first horizontal kernel and the first vertical kernel), and step S1506 includes applying a second horizontal kernel and a second vertical kernel to the input pixel values of the second input block 306 (i.e., the second set of one or more kernels applied in step S1506 includes the second horizontal kernel and the second vertical kernel). In addition, to determine a sharpened upsampling pixel value at a position that falls between the corresponding positions of the input pixel values of the first input block and the second input block (e.g., at the upper right corner or the lower left corner of the block 404 of the sharpened upsampling pixel values), step S1508 (combining the results of steps S1504 and S1506) includes Figure 21 the steps shown in
[0293] Figure 21It is a flowchart of a method for combining the results of applying kernels to input blocks to determine sharpened upsampled pixel values at positions falling between corresponding positions of input pixel values. Specifically, in step S2102, the pixel determination logic 1400 determines the horizontal component by summing the result of applying the first horizontal kernel to the input pixel values of the first input block 304 (in step S1504) and the result of applying the second horizontal kernel to the input pixel values of the second input block 306 (in step S1506). In step S2104, the pixel determination logic 1400 determines the vertical component by summing the result of applying the first vertical kernel to the input pixel values of the first input block 304 (in step S1504) and the result of applying the second vertical kernel to the input pixel values of the second input block 306 (in step S1506). Steps S2102 and S2104 can be executed serially or in parallel. Then, the pixel determination logic 1400 can combine the determined horizontal and vertical components (e.g., by performing steps S2106, S2108, and S2110) to determine the sharpened upsampled pixel values.
[0294] For example, in step S2106, the pixel determination logic 1400 multiplies the horizontal component (determined in step S2102) by a first weighting parameter (e.g., the weighting parameter a described above) to determine a weighted horizontal component. In step S2108, the pixel determination logic 1400 multiplies the vertical component (determined in step S2104) by a second weighting parameter (e.g., the weighting parameter b described above) to determine a weighted vertical component. Steps S2106 and S2108 can be executed serially or in parallel. In step S2110, the pixel determination logic sums the weighted horizontal component and the weighted vertical component to determine the sharpened upsampled pixel values.
[0295] Figure 22 Shows how the complete horizontal sharpened upsampled kernel 1904 can be decomposed into: (i) a kernel 2204 applied to the first input block 304 of input pixel values, and (ii) a kernel 2206 applied to the second input block 306 of input pixel values, for determining the horizontal component of the sharpened upsampled pixel in the upper right corner of block 404 for the sharpened upsampled pixel values. The position of the sharpened upsampled pixel value in the upper right corner of block 404 for the sharpened upsampled pixel values corresponds to position 2202, which is horizontally adjacent to two input pixel values of the first input block 304 and vertically adjacent to two input pixel values of the second input block 306.
[0296] If the kernel 1904 is applied to the input pixel value 302 to determine the horizontal component of the sharpened upsampled pixel value in the upper right corner of the block 404 of sharpened upsampled pixel values (in step S2102), the kernel 1904 will be centered at position 2202, and the values of the kernel 1904 will be multiplied by the input pixel values at the corresponding positions, and the results will be summed. The kernel 2204 includes those values of the kernel 1904 that will be multiplied by the input pixel values of the first input block 304. Additionally, the kernel 2206 includes those values of the kernel 1904 that will be multiplied by the input pixel values of the second input block 306. Applying the kernel 1904 to the input pixel value 302 to determine the horizontal component of the sharpened upsampled pixel value in the upper right corner of the block 404 of sharpened upsampled pixel values will give the same result as summing the results of applying the kernel 2204 to the input block 304 and applying the kernel 2206 to the input block 306.
[0297] Figure 22 It shows zero-padding the kernels 2204 and 2206 up to a 5x5 kernel. As described above, in actual applications, zero-padding may not be implemented in hardware (since multiplying by a fixed zero is always zero). In other words, the kernel 2204 can be implemented as a 3x4 kernel (where Figure 22 the zero-padding around the top edge, bottom edge, and left edge of the 5x5 kernel 2204 shown in is removed), and then it can be applied to the corresponding 3x4 region of the 5x5 input block 304, which is to the right of the second row pixel values, third row pixel values, and fourth row input pixel values within the input block 304. Similarly, the kernel 2206 can be implemented as a 2x3 kernel (where Figure 22 the zero-padding around the top edge, left edge, and right edge of the 5x5 kernel 2206 shown in and along the bottom two rows is removed), and then it can be applied to the corresponding 2x3 region of the 5x5 input block 306, which is in the second, third, and fourth columns of the second row input pixel values and third row input pixel values within the input block 306. The kernels 2204 and 2206 correspond to Figure 14 a pair of kernels 1406 in HTR .
[0298] Figure 23 It shows how the complete vertical sharpened upsampling kernel 2004 can be decomposed into: (i) the kernel 2304 applied to the first input block 304 of input pixel values, and (ii) the kernel 2306 applied to the second input block 306 of input pixel values, for determining the vertical component of the sharpened upsampled pixel in the upper right corner of the block 404 of sharpened upsampled pixel values. The position of the sharpened upsampled pixel value in the upper right corner of the block 404 of sharpened upsampled pixel values corresponds to position 2202.
[0299] If the kernel 2004 is applied to the input pixel value 302 to determine the vertical component of the sharpened upsampled pixel value in the upper right corner of the block 404 of sharpened upsampled pixel values, the kernel 2004 will be centered at position 2202, and the values of the kernel 2004 will be multiplied by the input pixel values at the corresponding surrounding positions, and the results will be summed. The kernel 2304 includes those values of the kernel 2004 that will be multiplied by the input pixel values of the first input block 304. Additionally, the kernel 2306 includes those values of the kernel 2004 that will be multiplied by the input pixel values of the second input block 306. In other words, the first vertical kernel 2304 has a first subset of the values of the complete vertical sharpened upsampled kernel 2004 (i.e., the values at the positions corresponding to the input pixel values in the first input block 304), and the second vertical kernel 2306 has a second subset of the values of the complete vertical sharpened upsampled kernel 2004 (i.e., the values at the positions corresponding to the input pixel values in the second input block 306). Applying the kernel 2004 to the input pixel 302 to determine the vertical component of the sharpened upsampled pixel value in the upper right corner of the block 404 will give the same result as summing the results of applying the kernel 2304 to the input block 304 and applying the kernel 2306 to the input block 306.
[0300] Figure 23 Illustrated is zero-padding the kernels 2304 and 2306 up to a 5x5 kernel. As described above, in a practical application, zero-padding may not be implemented in hardware. The kernel 2304 can be implemented as a 3x2 kernel (where Figure 23 the zero-padding around the top edge, bottom edge, and right edge of the 5x5 kernel 2306 shown in and along the leftmost two columns is removed), and then it can be applied to the corresponding 3x2 region of the 5x5 input block 304, which is in the third and fourth columns of the second row input pixel values, third row input pixel values, and fourth row input pixel values within the input block 304. Similarly, the kernel 2306 can be implemented as a 4x3 kernel (where Figure 23 the zero-padding around the bottom edge, left edge, and right edge of the 5x5 kernel 2304 shown in is removed), and then it can be applied to the corresponding 4x3 region of the 5x5 input block 306, which is at the top of the second column input pixel values, third column input pixel values, and fourth column input pixel values within the input block 306. The kernels 2304 and 2306 correspond to Figure 14 a pair of kernels 1406 in VTR .
[0301] As described above, steps S2106, S2108, and S2110 are performed to determine the sharpened upsampled pixel value in the upper right corner of block 404 of sharpened upsampled pixel values, where: (i) step S2106 includes multiplying the horizontal component (determined as described above with reference to Figure 22 ), by a first weighting parameter a to determine a weighted horizontal component; (ii) step S2108 includes multiplying the vertical component (determined as described above with reference to Figure 23 ), by a second weighting parameter b to determine a weighted vertical component; and (iii) step S2110 includes summing the weighted horizontal and weighted vertical components determined in steps S2106 and S2108 to determine the sharpened upsampled pixel value in the upper right corner of block 404 of sharpened upsampled pixel values.
[0302] Figure 24 Illustrates how the complete horizontal sharpened upsampling kernel 1904 can be decomposed into: (i) a kernel 2404 applied to a first input block 304 of input pixel values, and (ii) a kernel 2406 applied to a second input block 306 of input pixel values, for determining the horizontal component of the sharpened upsampled pixel value in the lower left corner of block 404 of sharpened upsampled pixel values. The position of the sharpened upsampled pixel value in the lower left corner of block 404 of sharpened upsampled pixels corresponds to position 2402, which is vertically adjacent to two input pixel values of the first input block 304 and horizontally adjacent to two input pixel values of the second input block 306.
[0303] If the kernel 1904 is applied to the input pixel value 302 to determine the horizontal component of the sharpened upsampled pixel value in the lower left corner of the block 404 of sharpened upsampled pixel values, the kernel 1904 will be centered at position 2402, and the values of the kernel 1904 will be multiplied by the input pixel values at the corresponding surrounding positions, and the results will be summed. The kernel 2404 includes those values of the kernel 1904 that will be multiplied by the input pixel values of the first input block 304. In addition, the kernel 2406 includes those values of the kernel 1904 that will be multiplied by the input pixel values of the second input block 306. In other words, the first horizontal kernel 2404 has a first subset of the values of the full horizontal sharpened upsampled kernel 1904 (i.e., the values of the full horizontal sharpened upsampled kernel 1904 at the positions corresponding to the input pixel values in the first input block 304), and the second horizontal kernel 2406 has a second subset of the values of the full horizontal sharpened upsampled kernel 1904 (i.e., the values of the full horizontal sharpened upsampled kernel 1904 at the positions corresponding to the input pixel values in the second input block 306). Applying the kernel 1904 to the input pixel value 302 to determine the horizontal component of the sharpened upsampled pixel value in the lower left corner of the block 404 of sharpened upsampled pixel values will give the same result as summing the results of applying the kernel 2404 to the input block 304 and applying the kernel 2406 to the input block 306.
[0304] Figure 24 The zero-padding of the kernels 2404 and 2406 up to a 5x5 kernel is shown. As described above, in practical applications, the zero-padding may not be implemented in hardware. The kernel 2404 can be implemented as a 2x3 kernel (where Figure 24 the zero-padding around the bottom edge, left edge, and right edge of the 5x5 kernel 2404 shown in Figure 24 and along the top two rows is removed), and then it can be applied to the corresponding 2x3 region of the 5x5 input block 304, which is in the second, third, and fourth columns of the third and fourth row input pixel values within the input block 304. Similarly, the kernel 2406 can be implemented as a 3x4 kernel (where Figure 14 the zero-padding around the top edge, bottom edge, and right edge of the 5x5 kernel 2406 shown in HBL is removed), and then it can be applied to the corresponding 3x4 region of the 5x5 input block 306, which is to the left of the second, third, and fourth row input pixel values within the input block 306. The kernels 2404 and 2406 correspond to
[0305] Figure 25Shows how the complete vertical sharpened upsampling kernel 2004 can be decomposed into: (i) a kernel 2504 applied to a first input block 304 of input pixel values, and (ii) a kernel 2506 applied to a second input block 306 of input pixel values, for use in determining the vertical component of the sharpened upsampled pixel in the lower left corner of block 404 of the sharpened upsampled pixel values. The position of the sharpened upsampled pixel value in the lower left corner of block 404 of the sharpened upsampled pixel values corresponds to position 2402.
[0306] If kernel 2004 is applied to input pixel 302 to determine the vertical component of the sharpened upsampled pixel value in the lower left corner of block 404 of the sharpened upsampled pixel values, then kernel 2004 will be centered at position 2402, and the values of kernel 2004 will be multiplied by the input pixel values at the corresponding surrounding positions, and the results will be summed. Kernel 2504 includes those values of kernel 2004 that will be multiplied by the input pixel values of the first input block 304. In addition, kernel 2506 includes those values of kernel 2004 that will be multiplied by the input pixel values of the second input block 306. In other words, the first vertical kernel 2504 has a first subset of the values of the complete vertical sharpened upsampling kernel 2004 (i.e., the values at positions corresponding to the input pixel values in the first input block 304), and the second vertical kernel 2506 has a second subset of the values of the complete vertical sharpened upsampling kernel 2004 (i.e., the values at positions corresponding to the input pixel values in the second input block 306). Applying kernel 2004 to input pixel values 302 to determine the vertical component of the sharpened upsampled pixel value in the lower left corner of block 404 of the sharpened upsampled pixel values will give the same result as summing the results of applying kernel 2504 to input block 304 and applying kernel 2506 to input block 306.
[0307] Figure 25 Shows zero-padding kernels 2504 and 2506 up to a 5x5 kernel. As described above, in practical applications, zero-padding may not be implemented in hardware. Kernel 2504 can be implemented as a 4x3 kernel (where Figure 25 the zero-padding around the top edge, left edge, and right edge of the 5x5 kernel 2504 shown in is removed), and then it can be applied to the corresponding 4x3 region of the 5x5 input block 304, which is at the bottom of the second column input pixel values, the third column input pixel values, and the fourth column input pixel values within the input block 304. Similarly, kernel 2506 can be implemented as a 3x2 kernel (where Figure 25The zero-padding around the top edge, bottom edge, and left edge of the 5x5 kernel 2506 shown in [Figure] and along the rightmost two columns is removed), and then it can be applied to the corresponding 3x2 region of the 5x5 input block 306, which is in the second and third columns of the second row input pixel values, the third row input pixel values, and the fourth row input pixel values within the input block 306. Kernels 2504 and 2506 correspond to Figure 14 a pair of kernels 1406 in VBL .
[0308] As described above, steps S2106, S2108, and S2110 are performed to determine the sharpened upsampled pixel value in the lower left corner of the block 404 of sharpened upsampled pixel values, where: (i) step S2106 includes multiplying the horizontal component (determined as described above with reference to Figure 24 ) by a first weighting parameter a to determine a weighted horizontal component; (ii) step S2108 includes multiplying the vertical component (determined as described above with reference to Figure 25 ) by a second weighting parameter b to determine a weighted vertical component; and (iii) step S2110 includes summing the weighted horizontal component and the weighted vertical component determined in steps S2106 and S2108 to determine the sharpened upsampled pixel value in the lower left corner of the block 404 of sharpened upsampled pixel values.
[0309] Figure 26 [Figure] shows a computer system in which the processing module described herein can be implemented. The computer system includes a CPU 2602, a GPU 2604, a memory 2606, a neural network accelerator (NNA) 2608, and other devices 2614, such as a display 2616, a speaker 2618, and a camera 2622. The processing block 2610 (corresponding to the processing module described herein) is implemented on the GPU 2604. In other examples, one or more of the depicted components may be omitted from the system, and / or the processing block 2610 may be implemented on the CPU 2602 or within the NNA 2608 or in a separate block of the computer system. The components of the computer system can communicate with each other via a communication bus 2620.
[0310] The processing module described herein is shown as including a plurality of functional blocks. This is merely illustrative and is not intended to define a strict division between different logical elements of such entities. Each functional block can be provided in any suitable manner. It should be understood that the intermediate values described herein as being formed by the processing module need not be physically generated by the processing module at any point and may only represent logical values that conveniently describe the processing performed by the processing module between its input and output.
[0311] The processing modules described herein can be included in hardware on an integrated circuit. The processing modules described herein can be configured to perform any of the methods described herein. In general, any of the functions, methods, techniques, or components described above can be implemented in software, firmware, hardware (e.g., fixed logic circuitry), or any combination thereof. Terms such as "module", "function", "component", "element", "unit", "block", and "logic" can be used herein to generically represent software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, function, component, element, unit, block, or logic represents program code that, when executed on a processor, performs the specified task. The algorithms and methods described herein can be executed by one or more processors executing code that causes the processors to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical discs, flash memory, hard disk memory, and other memory devices that can use magnetic, optical, and other technologies to store instructions or other data and that are accessible by a machine.
[0312] As used herein, the terms computer program code and computer-readable instructions refer to any kind of executable code for execution by a processor, including code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, byte code, code defining an integrated circuit (such as a hardware description language or netlist), and code expressed in a programming language code such as C, Java(RTM), or OpenCL(RTM). Executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, compiled, or run in a virtual machine or other software environment, causes a processor of a computer system that supports the executable code to perform the task specified by the code.
[0313] A processor, computer, or computer system can be any kind of device, machine, or dedicated circuit, or a collection or part thereof, having processing capabilities such that it can execute instructions. A processor can be or include any kind of general-purpose or special-purpose processor, such as a CPU, GPU, NNA, system-on-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), etc. A computer or computer system can include one or more processors.
[0314] The present invention also aims to cover software that defines the configuration of hardware as described herein, such as HDL (Hardware Description Language) software, for example for designing integrated circuits or for configuring programmable chips to perform desired functions. That is, a computer-readable storage medium encoded with computer-readable program code in the form of an integrated circuit definition data set can be provided, which, when processed (i.e., run) in an integrated circuit manufacturing system, configures the system to manufacture a processing module configured to perform any of the methods described herein, or to manufacture a processing module including any of the devices described herein. The integrated circuit definition data set can be, for example, an integrated circuit description.
[0315] Accordingly, a method of manufacturing a processing module as described herein at an integrated circuit manufacturing system can be provided. In addition, an integrated circuit definition data set can be provided, which, when processed in an integrated circuit manufacturing system, causes the method of manufacturing a processing module to be executed.
[0316] The integrated circuit definition data set can be in the form of computer code, such as a netlist, code for configuring a programmable chip, a hardware description language defining hardware suitable for manufacturing at any level in an integrated circuit, including as register transfer level (RTL) code, as high-level circuit representation (such as Verilog or VHDL), and as low-level circuit representation (such as OASIS(RTM) and GDSII). A higher-level representation (such as RTL) that logically defines hardware suitable for manufacturing in an integrated circuit can be processed at a computer system configured to generate a manufacturing definition of an integrated circuit in the context of a software environment that includes definitions of circuit elements and rules for combining those elements to generate a manufacturing definition of the integrated circuit so defined by the representation. As is typically the case where software is executed at a computer system to define a machine, one or more intermediate user steps (such as providing commands, variables, etc.) may be required to configure the computer system to generate a manufacturing definition of an integrated circuit to execute code that defines the integrated circuit to generate a manufacturing definition of the integrated circuit.
[0317] Reference will now be made to Figure 27 describe an example of processing an integrated circuit definition data set at an integrated circuit manufacturing system to configure the system to manufacture a processing module.
[0318] Figure 27An example of an integrated circuit (IC) manufacturing system 2702 is shown, which is configured to manufacture a processing module as described in any example herein. Specifically, the IC manufacturing system 2702 includes a layout processing system 2704 and an integrated circuit generation system 2706. The IC manufacturing system 2702 is configured to receive an IC definition data set (e.g., defining a processing module as described in any example herein), process the IC definition data set, and generate an IC according to the IC definition data set (e.g., which includes a processing module as described in any example herein). The processing of the IC definition data set configures the IC manufacturing system 2702 to manufacture an integrated circuit that includes a processing module as described in any example herein.
[0319] The layout processing system 2704 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout from an IC definition data set are known in the art and may, for example, involve synthesizing RTL code to determine a gate-level representation of the circuit to be generated, e.g., in terms of logic components (such as NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining the location information of the logic components, the circuit layout can be determined from the gate-level representation of the circuit. This can be done automatically or with user participation to optimize the circuit layout. When the layout processing system 2704 has determined the circuit layout, it can output a circuit layout definition to the IC generation system 2706. The circuit layout definition can be, for example, a circuit layout description.
[0320] As is known in the art, the IC generation system 2706 generates an IC according to the circuit layout definition. For example, the IC generation system 2706 can implement a semiconductor device manufacturing process for generating an IC, which can involve a multi-step sequence of lithography and chemical processing steps, during which an electronic circuit is gradually formed on a wafer made of semiconductor material. The circuit layout definition can be in the form of a mask, which can be used in a lithography process to generate an IC according to the circuit definition. Alternatively, the circuit layout definition provided to the IC generation system 2706 can be in the form of computer-readable code, which the IC generation system 2706 can use to form a suitable mask for generating the IC.
[0321] The different processes performed by the IC fabrication system 2702 can all be implemented in one location, e.g., by one party. Alternatively, the IC fabrication system 2702 can be a distributed system such that some processes can be performed at different locations and can be performed by different parties. For example, some of the following stages can be performed at different locations and / or by different parties: (i) synthesizing RTL code representing an IC definition data set to form a gate-level representation of the circuit to be generated; (ii) generating a circuit layout based on the gate-level representation; (iii) forming a mask according to the circuit layout; and (iv) manufacturing an integrated circuit using the mask.
[0322] In other examples, the processing of an integrated circuit definition data set at an integrated circuit fabrication system can configure the system to fabricate a processing module without processing the IC definition data set to determine a circuit layout. For example, the integrated circuit definition data set can define the configuration of a reconfigurable processor such as an FPGA, and the processing of the data set can configure the IC fabrication system to generate (e.g., by loading configuration data into the FPGA) a reconfigurable processor with the defined configuration.
[0323] In some embodiments, when processed in an integrated circuit fabrication system, the integrated circuit fabrication definition data set can cause the integrated circuit fabrication system to generate the devices described herein. For example, configuring the integrated circuit fabrication system in the manner described above with reference to the integrated circuit fabrication definition data set can fabricate the devices described herein. Figure 27
[0324] In some examples, the integrated circuit definition data set can include software that runs on hardware defined at the data set, or software that runs in combination with the hardware defined at the data set. In the Figure 27 example shown, the IC generation system can be further configured by the integrated circuit definition data set to load firmware onto the integrated circuit according to program code defined at the integrated circuit definition data set during the fabrication of the integrated circuit, or otherwise provide program code for use with the integrated circuit.
[0325] When compared to known embodiments, the implementation of the concepts set forth in this application in apparatuses, devices, modules, and / or systems (and in the methods implemented herein) can improve performance. Performance improvements can include one or more of increased computing performance, reduced latency, increased throughput, and / or reduced power consumption. During the manufacture of such apparatuses, devices, modules, and systems (e.g., in an integrated circuit), a trade-off can be made between performance improvements and physical implementation to improve the manufacturing method. For example, a trade-off can be made between performance improvements and layout area to match the performance of known embodiments but use less silicon. For example, this can be done by reusing functional blocks serially or sharing functional blocks among the elements of an apparatus, device, module, and / or system. Conversely, the concepts set forth in this application that result in improvements in the physical implementation of apparatuses, devices, modules, and systems (e.g., reduced silicon area) can be traded off against performance improvements. This can be done, for example, by manufacturing multiple instances of a module within a predefined area budget.
[0326] The applicant hereby independently discloses each individual feature described herein and any combination of two or more such features to the extent that such features or combinations are capable of being implemented based on the whole of this specification in view of the common general knowledge of a person skilled in the art, regardless of whether such features or combinations of features solve any of the problems disclosed herein. In view of the foregoing description, it will be apparent to a person skilled in the art that various modifications can be made within the scope of the present invention.
Claims
1. A method of determining an indication of one or more weighting parameters for applying upsampling to input pixel values representing an image region to determine a block of one or more upsampled pixel values, the method comprising: applying a horizontal edge filter to two or more of the input pixel values to determine a first filtered value; applying a vertical edge filter to two or more of the input pixel values to determine a second filtered value; applying a horizontal line filter to three or more of the input pixel values to determine a third filtered value; applying a vertical line filter to three or more of the input pixel values to determine a fourth filtered value; determining the indication of one or more weighting parameters using the first filter value, the second filter value, the third filter value, and the fourth filter value, wherein the one or more weighting parameters indicate relative horizontal and vertical variations of the input pixel values within the image region; as well as The determined indications of the one or more weighting parameters for applying upsampling to the input pixel values representing the image region to determine a block of one or more upsampled pixel values are outputted.
2. The method according to claim 1, wherein: The indicating of determining one or more weighting parameters using the first filtered value, the second filtered value, the third filtered value, and the fourth filtered value includes combining the first filtered value, the second filtered value, the third filtered value, and the fourth filtered value by performing a weighted sum.
3. The method according to claim 2, wherein: One or more of the weights used in the weighted sum are trained such that the indication of one or more weighting parameters indicates one or more weighting parameters indicative of relative horizontal and vertical variations of the input pixels within the image region.
4. A method according to any preceding claim, wherein: Prior to outputting the indication of the one or more weighting parameters, the indication of the one or more weighting parameters is clamped to be within the range [0, 1].
5. A method according to any preceding claim, wherein: The indicating of using the first filter value, the second filter value, the third filter value, and the fourth filter value to determine one or more weighting parameters includes: processing the input pixel values using an implementation of a neural network to determine a residual value; and The determined residual value is combined with the first filtered value, the second filtered value, the third filtered value, and the fourth filtered value to determine the indication of one or more weighting parameters.
6. The method according to claim 5, wherein: The neural network includes a first convolutional layer, a second convolutional layer, and a third convolutional layer, and wherein the processing of the input pixel value using the implementation of the neural network includes: determining, using the first convolutional layer, a first intermediate tensor based on the input pixel values, wherein the first intermediate tensor extends in only one of a horizontal dimension and a vertical dimension relative to a dimension of the image region represented by the input pixel values; determining, using the second convolutional layer, a second intermediate tensor based on the first intermediate tensor, wherein the second intermediate tensor does not extend in the horizontal dimension or the vertical dimension relative to a dimension of the image region represented by the input pixel values; and The residual value is determined based on the second intermediate tensor using the third convolutional layer.
7. The method according to claim 6, wherein: The image region is a portion of an input image, and wherein the upsampling is iteratively performed on a plurality of partially overlapping image regions within the input image, Wherein, the processing of the input pixel value using the implementation of the neural network further comprises storing the first intermediate tensor in a buffer, and The first convolutional layer operates on input pixel values within a first part of a current image region that does not overlap with a previous image region, but does not operate on input pixel values within a second part of the current image region that overlaps with the previous image region, so as to determine the first intermediate tensor.
8. The method according to claim 6 or 7, wherein: The neural network further includes a first activation function implemented between the first convolutional layer and the second convolutional layer, and a second activation function implemented between the second convolutional layer and the third convolutional layer, Any of the following is true: The first activation function is a first rectified linear unit and the second activation function is a second rectified linear unit, and the processing of the input pixel value using the embodiment of the neural network further comprises: (i) setting negative values in the first intermediate tensor to zero using the first rectified linear unit, and (ii) setting negative values in the second intermediate tensor to zero using the second rectified linear unit; or The first activation function is an identity function and the second activation function is an absolute value function, and the processing of the input pixel values using the implementation of the neural network further comprises setting values in the second intermediate tensor to absolute values using the absolute value function.
9. The method according to any one of claims 5 to 8, wherein: The neural network has been trained using quantization aware training (QAT).
10. A method according to any preceding claim, wherein: The horizontal line filter is configured to determine the third filter value to be zero when three or more of the input pixel values to which the horizontal line filter is applied exhibit pure vertical features, and wherein the vertical line filter is configured to determine the fourth filter value to be zero when three or more of the input pixel values to which the vertical line filter is applied exhibit pure horizontal features.
11. A method according to any preceding claim, wherein: The input pixel values are values of input pixels of a repeating quincuncial arrangement whose positions correspond to the upsampled pixel positions.
12. The method according to claim 11, wherein: The input pixel values are represented by two input blocks, wherein one of the two input blocks includes input pixel values of input pixels located within odd rows of a repeated pentad arrangement corresponding to upsampled pixel positions, and the other of the two input blocks includes input pixel values of input pixels located within even rows of a repeated pentad arrangement corresponding to upsampled pixel positions.
13. A method according to any preceding claim, wherein: An input pixel is represented by values in multiple channels, wherein an upsampled pixel is represented by values in multiple channels, wherein the input pixel value is the value of the input pixel in a single channel, and wherein the upsampled pixel value is the value of the upsampled pixel in the single channel.
14. A method according to any preceding claim, wherein: The input pixel values and the upsampled pixel values are Y channel values, and wherein the upsampling of the Y channel values is used in a super-resolution technique.
15. The method according to any one of claims 1 to 13, wherein: The input pixel value and the upsampled pixel value are green channel values, and wherein the upsampling of the green channel value is used in a demosaicing technique.
16. A method according to any preceding claim, further comprising applying upsampling to the input pixel values representing the image area, the upsampling comprising determining one or more of the upsampled pixel values of a block of the one or more upsampled pixel values based on relative horizontal and vertical changes in the input pixel values within the image area indicated by the one or more weighting parameters.
17. A method according to any preceding claim, wherein: One or more upsampled pixel values in the block of one or more upsampled pixel values are unsharp upsampled pixel values, or wherein one or more upsampled pixel values in the block of one or more upsampled pixel values are sharp upsampled pixel values.
18. A processing module configured to determine an indication of one or more weighting parameters for applying upsampling to input pixel values representing an image region to determine a block of one or more upsampled pixel values, the processing module comprising: horizontal edge filtering logic configured to apply a horizontal edge filter to two or more of the input pixel values to determine a first filtered value; vertical edge filtering logic configured to apply a vertical edge filter to two or more of the input pixel values to determine a second filtered value; horizontal line filtering logic configured to apply a horizontal line filter to three or more of the input pixel values to determine a third filtered value; vertical line filtering logic configured to apply a vertical line filter to three or more of the input pixel values to determine a fourth filtered value; as well as Processing logic, the processing logic being configured to: determining the indication of one or more weighting parameters using the first filter value, the second filter value, the third filter value, and the fourth filter value, wherein the one or more weighting parameters indicate relative horizontal and vertical variations of the input pixel values within the image region; and The determined indications of the one or more weighting parameters for applying upsampling to the input pixel values representing the image region to determine a block of one or more upsampled pixel values are outputted.
19. A computer-readable storage medium having computer-readable codes stored thereon, the computer-readable codes being configured to enable the method according to any one of claims 1 to 17 to be executed when the codes are executed.
20. A computer readable storage medium having stored thereon an integrated circuit definition data set, which when processed in an integrated circuit manufacturing system configures the integrated circuit manufacturing system to manufacture the process module of claim 18.