Similarity-driven downsampling

By incorporating a similarity function in the downsampling process, the method improves the representation quality of downsampled images, addressing the issue of weak representations in existing downsamplers.

WO2026104836A1PCT designated stage Publication Date: 2026-05-21V NOVA INT LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
V NOVA INT LTD
Filing Date
2025-11-14
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing downsamplers often produce downsampled images that are not strong representations of the original images, lacking sufficient detail and accuracy.

Method used

Implement similarity-driven downsampling by using a similarity function in the convolution process to enhance the downsampled signal, ensuring it retains more information and improves representation quality.

Benefits of technology

The similarity-driven downsampling method produces downsampled signals that are stronger representations of the original signals, enhancing the accuracy and detail of reconstructed images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025052494_21052026_PF_FP_ABST
    Figure GB2025052494_21052026_PF_FP_ABST
Patent Text Reader

Abstract

A downsampling method is provided. An input set of pixels (404) is obtained. A downsampled set of pixels (410) is generated by downsampling the input set of pixels (404). The downsampling uses a similarity function (418). The similarity function (418) is a function of pixel value differences of pixels of the input set of pixels (404).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SIMILARITY-DRIVEN DOWNSAMPLING

[0002] Technical Field

[0003] The present disclosure relates to similarity-driven downsampling.

[0004] Background

[0005] A downsampler may be used to downsample a signal, such as an image. For an image with width ‘IV’ and heightlH’, a downsampled image may have width

[0006]

[0007] and height

[0008]

[0009] and k are downsampling factors. When k = λ = 2, the downsampled image is a 2D-downsampled image. When λ = 2 and k = 1, the downsampled image is a ID-downsampled image. With some downsamplers, the downsampled image is not a particularly strong representation of the original image.

[0010] Summary

[0011] Various aspects of the present disclosure are set out in the appended claims. Further features and advantages will become apparent from the following description of preferred embodiments, given by way of example only, which is made with reference to the accompanying drawings.

[0012] Brief Description of the Drawings

[0013] Figure 1 shows a schematic block diagram of an example of an image processing system;

[0014] Figures 2A and 2B show schematic block diagrams of another example of an image processing system;

[0015] Figure 3 shows a schematic diagram of an example of downsampling an input signal;

[0016] Figure 4 shows a schematic diagram of an example of similarity-driven downsampling of an input signal;

[0017] Figure 5 shows a graph of example similarity functions; and

[0018] Figure 6 shows a schematic block diagram of an example of an apparatus. Detailed Description

[0019] Referring to Figure 1, there is shown an example of a signal processing system 100. The signal processing system 100 is used to process signals. Examples of types of signal include, but are not limited to, video signals, image signals, audio signals, volumetric signals such as those used in medical, scientific or holographic imaging, or other multidimensional signals.

[0020] The signal processing system 100 includes a first apparatus 102 and a second apparatus 104. The first apparatus 102 and second apparatus 104 may have a clientserver relationship, with the first apparatus 102 performing the functions of a server device and the second apparatus 104 performing the functions of a client device. The signal processing system 100 may include at least one additional apparatus (not shown). The first apparatus 102 and / or second apparatus 104 may comprise one or more components. The one or more components may be implemented in hardware and / or software. The one or more components may be co-located or may be located remotely from each other in the signal processing system 100. Examples of types of apparatus include, but are not limited to, computerised devices, handheld or laptop computers, tablets, mobile devices, games consoles, smart televisions, set-top boxes, Extended Reality (XR) headsets (including Augmented Reality (AR) and / or Virtual Reality (VR) headsets) etc.

[0021] The first apparatus 102 is communicatively coupled to the second apparatus 104 via a data communications network 106. Examples of the data communications network 106 include, but are not limited to, the Internet, a Local Area Network (LAN) and a Wide Area Network (WAN). The first and / or second apparatus 102, 104 may have a wired and / or wireless connection to the data communications network 106.

[0022] In this example, the first apparatus 102 comprises an encoder 108. The encoder 108 is configured to encode data comprised in and / or derived based on the signal, which is referred to hereinafter as “signal data”. For example, where the signal is a video signal, the encoder 108 is configured to encode video data. Video data comprises a sequence of multiple images or frames. The encoder 108 may perform one or more further functions in addition to encoding signal data. The encoder 108 may be embodied in various different ways. For example, the encoder 108 may be embodied in hardware and / or software. The encoder 108 may encode metadata associated with the signal. The first apparatus 102 may use one or more than one encoder 108.

[0023] Although in this example the first apparatus 102 comprises the encoder 108, in other examples the first apparatus 102 is separate from the encoder 108. In such examples, the first apparatus 102 is communicatively coupled to the encoder 108. The first apparatus 102 may be embodied as one or more software functions and / or hardware modules.

[0024] In this example, the second apparatus 104 comprises a decoder 110. The decoder 110 is configured to decode signal data. The decoder 110 may perform one or more further functions in addition to decoding signal data. The decoder 110 may be embodied in various different ways. For example, the decoder 110 may be embodied in hardware and / or software. The decoder 110 may decode metadata associated with the signal. The second apparatus 104 may use one or more than one decoder 110.

[0025] Although in this example the second apparatus 104 comprises the decoder 110, in other examples the second apparatus 104 is separate from the decoder 110. In such examples, the second apparatus 104 is communicatively coupled to the decoder 110. The second apparatus 104 may be embodied as one or more software functions and / or hardware modules.

[0026] The encoder 108 encodes signal data and transmits the encoded signal data to the decoder 110 via the data communications network 106. The decoder 110 decodes the received, encoded signal data and generates decoded signal data. The decoder 110 may output the decoded signal data, or data derived using the decoded signal data. For example, the decoder 110 may output such data for display on one or more display devices associated with the second apparatus 104. The one or more display devices may be components of the second apparatus 104 or may otherwise be associated with the second apparatus 104. The one or more display devices may be operable to display XR content and may, therefore, be referred to as XR display devices.

[0027] In some examples described herein, the encoder 108 transmits to the decoder 110 a representation of a signal at a given level of quality and information the decoder 110 can use to reconstruct a representation of some or all of the signal at one or more higher levels of quality. Such information may be referred to as “reconstruction data”. In some examples, “reconstruction” of a representation involves obtaining a representation that is not an exact replica of an original representation. The extent to which the representation is the same as the original representation may depend on various factors including, but not limited to, quantisation levels. A representation of a signal at a given level of quality may be considered to be a rendition, version or depiction of data comprised in the signal at the given level of quality. In some examples, the reconstruction data is included in the signal data that is encoded by the encoder 108 and transmitted to the decoder 110. For example, the reconstruction data may be in the form of metadata. In some examples, the reconstruction data is encoded and transmitted separately from the signal data.

[0028] The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may comprise residual data, as described in more detail below. Residual data is an example of reconstruction data. The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may also comprise configuration data relating to processing of the residual data. The configuration data may indicate how the residual data has been processed by the encoder 108 and / or how the residual data is to be processed by the decoder 110. The configuration data may be signalled to the decoder 110, for example in the form of metadata.

[0029] The first and / or second apparatuses 102, 104 may be configured to perform some or all of the techniques described herein. A computer program may be configured to perform some or all of the techniques described herein.

[0030] Referring to Figures 2A and 2B, there is shown schematically an example of a signal processing system 200. The signal processing system 200 includes a first apparatus 202 and a second apparatus 204. In this example, the first apparatus 202 comprises an encoder and the second apparatus 204 comprises a decoder. However, as explained above, in other examples, the encoder is not comprised in the first apparatus 202 and / or the decoder is not comprised in the second apparatus 204. In each of the first apparatus 202 and the second apparatus 204, items are shown on two logical levels. The two levels are separated by a dashed line. Items on the first, highest level relate to data at a first level of quality. Items on the second, lowest level relate to data at a second level of quality. The first level of quality is higher than the second level of quality. The first and second levels of quality relate to a tiered hierarchy having multiple levels of quality. In some examples, the tiered hierarchy comprises more than two levels of quality. In such examples, the first apparatus 202 and the second apparatus 204 may include more than two different levels. There may be one or more other levels above and / or below those depicted in Figures 2A and 2B. As described herein, in certain cases, the levels of quality may correspond to different spatial resolutions.

[0031] Referring first to Figure 2A, the first apparatus 202 obtains a first representation of an image at the first level of quality 206. A representation of a given image is a representation of data comprised in the image. The image may be a given frame of a video. The first representation of the image at the first level of quality 206 will be referred to as “input data” hereinafter as, in this example, it is data provided as an input to the encoder in the first apparatus 202. The first apparatus 202 may receive the input data 206. For example, the first apparatus 202 may receive the input data 206 from at least one other apparatus. The first apparatus 202 may be configured to receive successive portions of input data 206, e.g. successive frames of a video, and to perform the operations described herein to each successive frame. For example, a video may comprise frames Fi, F2,... FT and the first apparatus 202 may process each of these in turn.

[0032] The first apparatus 202 derives data 212 based on the input data 206. In this example, the data 212 based on the input data 206 is a representation 212 of the image at the second, lower level of quality. In this example, the data 212 is derived by performing a downsampling operation on the input data 206 and will therefore be referred to as “downsampled data” hereinafter. In other examples, the data 212 is derived by performing an operation other than a downsampling operation on the input data 206, or the data 212 is the same as the input data 206 (i.e. the input data 206 is not processed, e.g. downsampled).

[0033] In this example, the downsampled data 212 is processed to generate processed data 213 at the second level of quality. In other examples, the downsampled data 212 is not processed at the second level of quality. As such, the first apparatus 202 may generate data at the second level of quality, where the data at the second level of quality comprises the downsampled data 212 or the processed data 213.

[0034] In some examples, generating the processed data 213 involves the downsampled data 212 being encoded. Such encoding may occur within the first apparatus 202, or the first apparatus 202 may output the processed data 213 to an external encoder. Encoding the downsampled data 212 produces an encoded image at the second level of quality. The first apparatus 202 may output the encoded image, for example for transmission to the second apparatus 204. A series of encoded images, e.g. forming an encoded video, as output for transmission to the second apparatus 204 may be referred to as a “base” stream or “base” layer. As explained above, instead of being produced in the first apparatus 202, the encoded image may be produced by an encoder that is separate from the first apparatus 202. The encoded image may be part of an H.264 or H.265 encoded video, or otherwise. Generating the processed data 213 may, for example, comprise generating successive frames of video as output by a separate encoder such as an H.264 or H.265 video encoder. An intermediate set of data for the generation of the processed data 213 may comprise the output of such an encoder, as opposed to any intermediate data generated by the separate encoder.

[0035] Generating the processed data 213 at the second level of quality may further involve decoding the encoded image at the second level of quality. The decoding operation may be performed to emulate a decoding operation at the second apparatus 204, as will become apparent below. Decoding the encoded image produces a decoded image at the second level of quality. In some examples, the first apparatus 202 decodes the encoded image at the second level of quality to produce the decoded image at the second level of quality. In other examples, the first apparatus 202 receives the decoded image at the second level of quality, for example from an encoder and / or decoder that is separate from the first apparatus 202. The encoded image may be decoded using an H.264 or H.265 decoder. The decoding by a separate decoder may comprise inputting encoded video, such as an encoded data stream configured for transmission to a remote decoder, into a separate black-box decoder implemented together with the first apparatus 202 to generate successive decoded frames of video. Processed data 213 may thus comprise a frame of video data that is generated via a complex non-linear encoding and decoding process, where the encoding and decoding process may involve modelling spatio-temporal correlations as per a particular encoding standard such as H.264 or H.265. However, because the output of any encoder is fed into a corresponding decoder, this complexity is effectively hidden from the first apparatus 202. In an example, generating the processed data 213 at the second level of quality further involves obtaining correction data based on a comparison between the downsampled data 212 and the decoded image obtained by the first apparatus 202, for example based on the difference between the downsampled data 212 and the decoded image. The correction data can be used to correct for encoder-decoder errors (which may also be referred to as “encode-decode errors”), namely errors introduced in encoding and decoding the downsampled data 212. In some examples, the first apparatus 202 outputs the correction data, for example for transmission to the second apparatus 204, as well as the encoded signal. This allows the recipient to correct for the encoder-decoder errors introduced in encoding and decoding the downsampled data 212. This correction data may also be referred to as a “first enhancement” stream. As the correction data may be based on the difference between the downsampled data 212 and the decoded image it may be seen as a form of residual data (e.g. that is different from the other set of residual data described later below). An item of residual data may be referred to as “a residual”.

[0036] In some examples, generating the processed data 213 at the second level of quality further involves correcting the decoded image using the correction data. For example, the correction data as output for transmission may be placed into a form suitable for combination with the decoded image, and then added to the decoded image. This may be performed on a frame-by-frame basis. In other examples, rather than correcting the decoded image using the correction data, the first apparatus 202 uses the downsampled data 212. For example, in certain cases, just the encoded then decoded data may be used and, in other cases, encoding and decoding may be replaced by other processing.

[0037] In some examples, generating the processed data 213 involves performing one or more operations other than the encoding, decoding, obtaining, and correcting acts described above.

[0038] The first apparatus 202 obtains data 214 based on the data at the second level of quality. As indicated above, the data at the second level of quality may comprise the processed data 213, or the downsampled data 212 where the downsampled data 212 is not processed at the lower level. As described above, in certain cases, the processed data 213 may comprise a reconstructed video stream (e.g. from an encoding-decoding operation) that is corrected using correction data. In the example of Figures 2A and 2B, the data 214 is a second representation of the image at the first level of quality, the first representation of the image at the first level of quality being the input data 206. The second representation at the first level of quality may be considered to be a preliminary or predicted representation of the image at the first level of quality. In this example, the first apparatus 202 derives the data 214 by performing an upsampling operation on the data at the second level of quality. The data 214 will be referred to hereinafter as “upsampled data”. However, in other examples one or more other operations could be used to derive the data 214, for example where data 212 is not derived by downsampling the input data 206.

[0039] The input data 206 and the upsampled data 214 are used to obtain residual data 216. The residual data 216 is associated with the image. The residual data 216 may be in the form of a set of residual elements, which may be referred to as a “residual frame” or a “residual image”. A residual element may be referred to as “a residual”. A residual element in the set of residual elements 216 may be associated with a respective image element in the input data 206. An example of an image element is a pixel.

[0040] In this example, a given residual element is obtained by subtracting a value of an image element in the upsampled data 214 from a value of a corresponding image element in the input data 206. As such, the residual data 216 is useable in combination with the upsampled data 214 to reconstruct the input data 206. The residual data 216 may also be referred to as “reconstruction data” or “enhancement data”. In one case, the residual data 216 may form part of a “second enhancement” stream. The residual data 216 may therefore result from upsampler-downsampler asymmetry. Upsampler-downsampler asymmetry may also be referred to as “upsamping-downsamping asymmetry”, “upsample-downsample” asymmetry or the like.

[0041] The first apparatus 202 obtains configuration data relating to processing of the residual data 216. The configuration data indicates how the residual data 216 has been processed and / or generated by the first apparatus 202 and / or how the residual data 216 is to be processed by the second apparatus 204. The configuration data may comprise a set of configuration parameters. The configuration data may be useable to control how the second apparatus 204 processes data and / or reconstructs the input data 206 using the residual data 216. The configuration data may relate to one or more characteristics of the residual data 216. The configuration data may relate to one or more characteristics of the input data 206. Different configuration data may result in different processing being performed on and / or using the residual data 216. The configuration data is therefore useable to reconstruct the input data 206 using the residual data 216. As described below, in certain cases, configuration data may also relate to the correction data described herein.

[0042] In this example, the first apparatus 202 transmits to the second apparatus 204 data based on the downsampled data 212, data based on the residual data 216, and the configuration data (or data based on the configuration data), to enable the second apparatus 204 to reconstruct the input data 206.

[0043] Turning now to Figure 2B, the second apparatus 204 receives data 220 based on (e.g. derived from) the downsampled data 212. The second apparatus 204 also receives data based on the residual data 216. For example, the second apparatus 204 may receive a “base” stream (data 220), a “first enhancement stream” (any correction data) and a “second enhancement stream” (residual data 216). The base stream may be referred to as a “base layer”. The first and / or second enhancement stream may be referred to as, and / or may be comprised in, an “enhancement layer”. The second apparatus 204 also receives the configuration data relating to processing of the residual data 216. The data 220 based on the downsampled data 212 may be the downsampled data 212 itself, the processed data 213, or data derived from the downsampled data 212 or the processed data 213. The data based on the residual data 216 may be the residual data 216 itself, or data derived from the residual data 216.

[0044] In some examples, the received data 220 comprises the processed data 213, which may comprise the encoded image at the second level of quality and / or the correction data. In some examples, for example where the first apparatus 202 has processed the downsampled data 212 to generate the processed data 213, the second apparatus 204 processes the received data 220 to generate processed data 222. Such processing by the second apparatus 204 may comprise decoding an encoded image (e.g. that forms part of a “base” encoded video stream) to produce a decoded image at the second level of quality. In some examples, the processing by the second apparatus 204 comprises correcting the decoded image using obtained correction data. Hence, the processed data 222 may comprise a frame of corrected data at the second level of quality. In some examples, the encoded image at the second level of quality is decoded by a decoder that is separate from the second apparatus 204. The encoded image at the second level of quality may be decoded using an H.264 decoder.

[0045] In other examples, the received data 220 comprises the downsampled data 212 and does not comprise the processed data 213. In some such examples, the second apparatus 204 does not process the received data 220 to generate processed data 222.

[0046] The second apparatus 204 uses data at the second level of quality to derive the upsampled data 214. As indicated above, the data at the second level of quality may comprise the processed data 222, or the received data 220 where the second apparatus 204 does not process the received data 220 at the second level of quality. The upsampled data 214 is a preliminary representation of the image at the first level of quality. The upsampled data 214 may be derived by performing an upsampling operation on the data at the second level of quality.

[0047] The second apparatus 204 obtains the residual data 216. The residual data 216 is useable with the upsampled data 214 to reconstruct the input data 206. The residual data 216 is indicative of a comparison between the input data 206 and the upsampled data 214.

[0048] The second apparatus 204 also obtains the configuration data related to processing of the residual data 216. The configuration data is useable by the second apparatus 204 to reconstruct the input data 206. For example, the configuration data may indicate a characteristic or property relating to the residual data 216 that affects how the residual data 216 is to be used and / or processed, or whether the residual data 216 is to be used at all. In some examples, the configuration data comprises the residual data 216.

[0049] There are several considerations relating to such processing. One such consideration is the amount of information that is generated, stored, transmitted and / or processed. The more information that is used, the greater the amount of resources that may be involved in handling such information. Examples of such resources include transmission resources, storage resources and processing resources. Some signal processing techniques allow a relatively small amount of information to be used. This may reduce the amount of data transmitted via the data communications network 106. The savings may be particularly relevant where the data relates to high quality video data, where the amount of information transmitted can be especially high.

[0050] Another consideration is latency. Complex image processing may introduce latency, which may negatively impact performance.

[0051] Other considerations include the ability of the decoder to perform image reconstruction accurately, reliably, and / or efficiently. Performing image reconstruction accurately and reliably may affect the ultimate visual quality of the displayed image and consequently may affect a viewer’s engagement with the image and / or with a video comprising the image. This can be especially relevant to XR. Efficient reconstruction is especially effective for mobile computing devices, which may readily be used in XR applications.

[0052] Referring to Figure 3, there is shown an example of a signal processing system 300.

[0053] The example signal processing system 300 comprises a downsampler 302. In this example, the downsampler 302 obtains an input signal 304, denoted ‘x’. The downsampler 302 may obtain the input signal 304 by receiving the input signal 304, by generating the input signal 304, by retrieving the input signal 304 (for example, from memory), or otherwise. For example, the downsampler 302 may obtain the input signal 304 from a sensor. Examples of sensors include, but are not limited to, cameras, thermal sensors and microphones.

[0054] In examples described herein, the input signal 304 is generally an input image. However, the input signal 304 may take other forms in other examples. For instance, the input signal 304 may comprise an audio signal, a temperature signal, and so on. Similarly, where examples described herein concern images, such examples may be applicable to signals in general, including other types of signal.

[0055] In examples, signals comprise signal elements. Where the signal is an image, the signal elements may correspond to pixels. The signal elements may have respective signal element values. Where the signal is an image, the signal element values may correspond to pixel values.

[0056] In this example, the downsampler 302 obtains a kernel 306, denoted ‘X’. The downsampler 302 may obtain the kernel 306 by receiving the kernel 306, by generating the kernel 306, by retrieving the kernel 306 (for example, from memory), or otherwise. The kernel 306 may be referred to as a “downsampling kernel”.

[0057] In this example, the downsampler 302 convolves the input signal 304 with the kernel 306, denoted x * K, as shown by convolution 308 in Figure 3.

[0058] In this example, the downsampler 302 generates a downsampled signal 310, denoted ‘d’, by downsampling the input signal 304 using the kernel 306. The downsampled signal 310 may also be referred to as a “filtered signal”, a “base signal”, a “downscaled signal” or the like.

[0059] The downsampled signal 310 may be encoded by an encoder to generate an encoded signal. The encoded signal may be output.

[0060] The encoded signal may be decoded by a decoder to generate a decoded version of the encoded signal. Correction data may be generated based on differences between the downsampled signal 310 and the decoded version of the encoded version of the downsampled signal 310. The correction data may therefore correct for encode-decode errors. The correction data may be output.

[0061] The downsampled signal 310 may be upsampled by an upsampler to generate an upsampled signal 312, denoted ‘u’. The upsampled signal 312 may be considered to be a predicted version of the input signal 304. A predicted version may also be referred to as a “predicted rendition”.

[0062] A residual generator 314 may obtain the input signal 304 and the upsampled signal 312 and may generate a set of residuals 316 based on the input signal 304 and the upsampled signal 312. The set of residuals 316 may represent differences between the input signal 304 and the upsampled signal 312, where the upsampled signal 312 is a predicted version of the input signal 304. Each residual in the set of residuals 316 may represent a difference between a value of a signal element of the input signal 304 and a value of a corresponding signal element of the upsampled signal 312. The set of residuals 316 may therefore account for downsample-upsample asymmetries. The set of residuals 316 may be transformed, quantized and / or encoded. The set of residuals 316 may be output.

[0063] The input signal 304, the kernel 306, the convolution 308, the downsampled signal 310, the upsampled signal 312, the residual generator 314, and the residuals 316 will be described in more detail herein. Reference is made in particular to Figures 2A and 2B and to the description of Figures 2A and 2B above, for example in relation to correction data, upscaling, and residuals.

[0064] A specific example of an input signal 304 and a specific example of a kernel 306 will now be described to facilitate an understanding of signal processing in the example signal processing system 300.

[0065] In this example, the input signal 304 comprises input data, x, in the form of a row signal. The input data, x, may be referred to as an “input signal”, an “input image”, an “input row”, an “input row signal”, or the like.

[0066] In this specific example, the input data, x, is defined as: x = [30,31,32,33,34,35,36,248,249,250,251,252,253,254].

[0067] The value of the input data, x, at position i is denoted x[i]. The position may be referred to as an “index”. In this example, i is defined in the range [0,13]. Thus, in this example, x[0] = 30, x[1] = 31, ..., x

[0012] = 253, and x

[0013] = 254.

[0068] In this example, to handle out-of-range values of x[i] when i < 0, x[i] is defined for i < 0 as: x[i] = x[— i]. Thus, in this example, x[— 1] = x[l] = 31, x[— 2] = x[2] = 32, and so on.

[0069] In this example, to handle out-of-range values of x[i] when i > 13, x[i] is defined for i > 13 as: x[i] = x[14 — (imodl4) — 2], Thus, in this example, x

[0014] = x[14 - (14modl4) - 2] = x[14 - 0 - 2] = x

[0012] = 253, x

[0015] = x[14 - (15modl4) — 2] = x[14 — 1 — 2] = x[14 — 3] = x[ll] = 252, and so on.

[0070] An extended set of signal element values (for example, pixel values) may thereby be defined for when a calculated index is not a defined index within the input data, x, i.e. when the calculated index is not in the range [0,13] in this example. A signal element value from the extended set of signal element values may thereby be used when the calculated index is not a defined index within the input data, x. Examples of extended sets of signal element values have been provided above for when the calculated index is outside the range [0,13].

[0071] In this specific example, the signal element values in the extended set of signal element values are mirrored with respect to the input data, x. Such mirrored values may be referred to as “reflected” values. However, the signal element values in the extended set of signal element values may be defined differently in other examples.

[0072] In a first such example, the signal element values in the extended set of signal element values may be zero-filled. Using the above example input data, x, defined in the range i E [0,13], a zero-filled extended set may be defined such that x[i] is defined for i < 0 and for i > 13 as x[i] = 0.

[0073] In a second such example, the first signal element value in the input data, x, may be used for all calculated indices before the first defined index, i.e. before i = 0. Similarly, the final signal element value in the input data, x, may be used for all calculated indices after the final defined index, i.e. after i = 13. Using the above example input data, x, defined in the range i E [0,13], such an extended set may be defined such that x[i] is defined for i < 0 as x[i] = x[0] and for i > 13 as x[i] = x

[0013] .

[0074] In a third such example, the signal element values in the extended set of signal element values are wrapped with respect to the input data, x. Using the above example input data, x, defined in the range i E [0,13], such an extended set may be defined such that x[i] is defined for i < 0 as x[i] = x[(i + 14)modl4] and for i > 13 as x[i] = x[imodl4]. For example, for i = — 1, x[— 1] = x[((— 1) + 14)modl4] = x[13(modl4)] = x

[0013] . For i = 14, x

[0014] = x[14modl4] = x[0].

[0075] Wrapped signal element values may also be referred to as “circular” signal element values.

[0076] In this specific example, the downsampling kernel, K, is a row signal, defined as: K = [60,247,-557,-1092,2220,7314,7314,2220, -1092, -557,247,60].

[0077] Signal elements of the downsampling kernel, K, may be referred to as “kernel coefficients”.

[0078] The value of the downsampling kernel, K, at position j is denoted K[j], j may be defined in the range [— — 1], where N is the length of the kernel, K. Thus, j

[0079]

[0080] may have N different values.

[0081] [ 12 12 1 — — — 11 = [—6,5]. In

[0082]

[0083] this example, K[ ^6 ] = 60, K[ -5 ] = 247,..., K

[0004] = 247, K

[0005] = 60.

[0084] ;=-6 7=-5 7=4 7=5 Convolution may be computed as a dot product between signal element values of the input signal, x, and kernel coefficients of the downsampling kernel, K. An example of such convolution, which may be performed by the downsampler 302, will now be described. The downsampler 302 may perform the convolution using a convolution operation defined as follows:

[0085] —i

[0086] d[i] =; - S2Nx[2i -j] ■ K[j].

[0087] ~

[0088] In this formula, d [i] is the 1 downsampled signal, d, obtained by the downsampler rj

[0089] 302 at the Ithposition, x[i] represents the input signal, x, at position i, K[j] is the kernel coefficient of the kernel, K, at index j, and s is the kernel sum. This convolution may be written as d = x * K. The kernel sum, s, may also be referred to as a “normalisation factor”.

[0090] In this example, the kernel sum, s, is defined as:

[0091] —i

[0092]

[0093] In this specific example in which N = 12:

[0094] d

[0095]

[0096] [i] = I ■ S>-6x[2i - J] ’ ^[ / l

[0097] In this example, the input signal, x, has fourteen values. In this example, the input signal, x, is being 2D-downsampled, resulting in the downsampled signal, d, having seven values.

[0098] In another example, the downsampler 302 may perform the convolution using an alternative convolution operation defined as follows:

[0099] d[i] = \x[2i -j] ■ K[j]

[0100]

[0101] In such an example, [■] represents a floor function.

[0102] An explanation of how the downsampled signal, d, may be computed will now be provided.

[0103] Firstly, for ease of understanding, a table showing the mapping between the input signal index, i, and the corresponding input signal value, x[i], is provided below:

[0104] i -5 -4 -3 -1 0 1 2 3 4 5 6

[0105]

[0106] x[l] 35 34 33 32 31 30 31 32 33 34 35 36

[0107] i 7 8 9 10 11 12 13 14 15 16 17 18 x[i] 248 249 250 251 252 253 254 253 252 251 250 249

[0108]

[0109] The table above includes an extended set of signal element values, i.e. with values of i outside the range [0,13].

[0110] Secondly, and again for ease of understanding, a table showing the mapping between the kernel index, j, and the corresponding kernel value, K [ / ], is provided below:

[0111] J -6 -5 -4 -3 —2 -1 0 1 2 3 4 5 K[j] 60 247 -557 -1092 2220 7314 7314 2220 -1092 -557 247 60

[0112]

[0113] In this example, the kernel, K, is a symmetric kernel. An example of a symmetric kernel is aLanczos kernel. Using a symmetric kernel can increase efficiency since a special case of convolution may be used in which kernel inversion is not used. Kernel inversion is not used in such examples because inversion has no effect on a symmetric kernel.

[0114] The kernel sum, S, may be computed as:

[0115] s

[0116]

[0117] = £5=-6K[j] = K[-6] + K[-5] + ••• + K[4] + K[5] = 60 + 247 + ••• + 247 + 60 = 16384.

[0118] To compute the first value of the downsampled signal, d, i.e. to compute d[i] where i = 0, the following formula may be used:

[0119] d[0] = | - S75=-6 *[( 2 X 0 ] -j] - K[j].

[0120] \ i=0 / Thus: d[0] = ■ ((x[6] ■ K[-6]) + (x[5J ■ K[-5]) + - + (x[-4] ■ K[4]) + (x[— 5] ■ K[5])) = ■ ((36 x 60) + (35 X 247) + ••• + (34 X 247) + (35 X 60)) = ■ (2160 + 8645 + ••• + 8398 + 2100) = ■ (499018) « 30.

[0121]

[0122] Thus, in this example, the final value of d[0] = 30. Using the above-indicated alternative convolution operation, the first value of the downsampled signal, d, i.e. d[i] where i = 0, may be computed as follows:

[0123] d[L0]J= | 2 x 0 ) -; ■ K[j] ) +- = I ■ 499018) + \ s 6 I / J 2 1V16384 7 L\ L \ i=o / / J |] « [(30.46) + 1] = [30.96] = 30.

[0124]

[0125] Thus, using this alternative convolution operation, again the final value of d[0] = 30.

[0126] To compute the second value of the downsampled signal, d, i.e. to compute d [i] where i = 1, the following formula may be used:

[0127] d[l] = | - S75=-6X[2 x 1 -j] • K[j], i=l Thus: d [1] = — ■ ((x[8] ■ K[-6]) + (x[7] ■ K[-5]) + - + (x[-2] ■ K[4]) + (x[— 3] ■ K[5])) = ■ ((249 x 60) + (248 X 247) + ••• + (32 x 247) + (33 x 60)) = ■ (14940 + 61136 + ••• + 9884 + 1980) = ■ (597491) « 36.

[0128]

[0129] 16384k 716384k 7Thus, in this example, the final value of d[l] = 36.

[0130] Using the above-indicated alternative convolution operation, the first value of the downsampled signal, d, i.e. d[i] where i = 0, may be computed as follows:

[0131] d[l] = f S?._6x[2 x 1 -)] • / <[ / ] +1 ■ 597491) + -I 16384 7 21

[0132]

[0133] i=l [(36.47) + ] = [33.97] = 36.

[0134] Thus, using this alternative convolution operation, again the final value of d[l] = 36.

[0135] The values of d[2], d[3], d[4], d[5], and d[6] may be computed in a similar manner, giving d[2] = 17, d[3] = 142, d[4] = 267, d[5] = 248, d[6] = 254. In this example, the computed value for d[4] of ‘267’ is higher than the maximum permitted value of ‘255’. In this example, this computed value is clipped at ‘255’ to correspond to the maximum permitted value.

[0136] Thus, in this example, the final downsampled row, d, is d = [30,36,17,142,255,248,254], In examples that will now be described, this downsampling process may be improved by considering signal element value differences when filtering, i.e. downsampling, a signal. In particular, a modified convolution operation may be used to provide similarity-driven downsampling. Similarity-driven downsampling may result in downsampled signals that are stronger representations of input signals compared to downsampled signals that are not based on similarity-driven downsampling. Thus, such examples concern improved downsamplers and improved downsampling.

[0137] Referring to Figure 4, there is shown another example of a signal processing system 400.

[0138] The example signal processing system 400 shown in Figure 4 corresponds generally to the example signal processing system 300 shown in Figure 3. Reference signs used in Figure 4 are the same as those used in Figure 3 for the same or similar features, but incremented by 100.

[0139] In this example, in addition to the downsampler 402 obtaining the input signal 404 and the kernel 406, the downsampler 402 obtains a similarity function 418, denoted ‘h’. The downsampler 402 may obtain the similarity function 418 by receiving the similarity function 418, by generating the similarity function 418, by retrieving the similarity function 418 (for example, from memory), or otherwise. The downsampler 402 uses the similarity function 418 to perform a convolution 408 that is different from the convolution 308 described above with reference to Figure 3 and that will be explained in more detail below. The downsampler 402 outputs a downsampled signal 410, which is denoted ‘£>’ to differentiate from the downsampled signal 310, ‘d’, generated by the downsampler 302 described above with reference to Figure 3. The downsampled signal 410 may be encoded and decoded, and correction data may be generated based on differences between the downsampled signal 410 and the decoded version of the encoded version of the downsampled signal 410. The downsampled signal 410 may be upsampled to generate an upsampled signal 412, denoted ‘I / ’. A residual generator 414 may generate residuals 416 based on the input signal 404 and the upsampled signal 412. Reference is made again to Figures 2A and 2B and their associated description above. In an example that will now be described, the example downsampler functionality described above with reference to Figure 3 is modified to have similarity-driven downsampler functionality.

[0140] In this example, an example convolution operation used by the similarity-driven downsampler 402 is defined as:

[0141] —i

[0142] 0[i] = 7777' I2Nx[2i - j] ■ K[j] ■ h(x[2i -j],x[2i]),

[0143]

[0144] where £)[i] is the downsampled signal obtained by the similarity-driven downsampler 402 at the Ithposition, where h(x[2i — j],x[2i]) denotes a similarity function, h, between signal element values x[2i — j] and x[2i], and where S[i] denotes a modified kernel sum at the Ithposition. The similarity function, h, may be referred to as a “weighting function” etc. The modified kernel sum, 5[i], may be referred to as a “normalisation factor”.

[0145] In this example, the modified kernel sum, 5[i], is defined as:

[0146] — -1

[0147] S[i] = S2_NK[j - (x[2i -j],x[2i]).

[0148] The convolution above may be written as D = x * K', where K' represents a modified kernel and where K' = K ■ h.

[0149] This convolution operation differs from the convolution operation performed by the downsampler 302 described above with reference to Figure 3 in that this convolution involves the similarity function 418 and in that a modified kernel sum, 5[i], which is dependent on the similarity function 418, is also used.

[0150] The convolution operation may be defined more generally as:

[0151] --1

[0152] 0[i] = 7777' I2Nx[Ai - j] ■ K[j] ■ h(x[Ai - j],x[Ai]),

[0153] SpJJ =-- — -1

[0154] where S[i] = £2N ^[ / ] ’ h(x[Ai — j],x[Ai]), and where A is a downsampling

[0155]

[0156] factor.

[0157] In another example, the similarity-driven downsampler 402 may perform the convolution using an alternative convolution operation defined as follows:

[0158] £>[i] = ( 77777 ’ S2NX[AI -j] ■ K[j] ■ h(x[Ai -j],x[Ai]) + 7.

[0159]

[0160] \sLlJ ~ In this specific example, = 2. However, the downsampling factor,, may be a different value in other examples.

[0161] The similarity function, h, may take various different forms.

[0162] In this specific example, the similarity function, h, is a linear mapping, defined as:

[0163] h

[0164]

[0165] (x[2i - y], x[2i]) = 1 - The similarity function, h, thus relates to similarity between the signal element values x[2i — j] and x[2i], In this example, the similarity function, h, has a value in the range [0,1], where a value of ‘0’ means no similarity at all and a value of ‘1’ represents total similarity, i.e. the same value. In this example, the similarity function, h, uses a linear formula. Such a similarity function, h, may be particularly effective for a greyscale image having 256 different possible values.

[0166] More generally however, the similarity function, h x[ i — j],x[Ai]) between the input signal, x, at a calculated index i — j, and the input signal, x, at an index i may 1 7 z r-i * * r-i,_i\ -4

[0167]

[0168] — / I—I > • *i 1 be II(X[AI — JJ, X IJ) = 1 - — -, where m represents a maximum possible image element value difference.

[0169] In some examples, m = 2b— 1, where b is a bit-length of the signal elements (e.g. pixels) in the input signal, x. For instance, for a greyscale image having a bitlength of eight bits (corresponding to 256 possible pixel values), m = 28— 1 = 255.

[0170] Although, in this specific example, the similarity function, h, uses an absolute value of the difference between the signal element values x[2i — j] and x[2i], the signed difference may be used in other examples. For instance, the similarity function, h, may be defined as:

[0171] h(x[2i — j],x[2i]) = 1 -x[2i~j^2i\

[0172]

[0173] In this specific example in which N = 12:

[0174] f>[i] = Sj=-6*[2i- J] ‘ ‘ h(x[2i -j],x[2i]).

[0175]

[0176] In this example, the input signal, x, again has fourteen values. Since the input signal, x, is again being 2D-downsampled in this example, the downsampled signal, D, has seven values. To compute the first value of the downsampled signal, D, i.e. to compute £)[i] where i = 0, the following formula may be used:

[0177] D2 x 0 j — j ■ K [j] ■ h I x 2 x 0 x 2 x t°J = S ' 5-6 * t=o / \ 1=0

[0178]

[0179] For ease of understanding, a table showing how h(x[— j],x[0]) may be computed for each value of j E [—5,6] is provided below. The table shows the value of |x[— j] — x[0] |, and then the corresponding value of h(x[— j], x[0]):

[0180] -j -5 -4 -3 —2 -1 0 1 2 3 4 5 6 *[-;] 35 34 33 32 31 30 31 32 33 34 35 36 j 5 4 3 2 1 0 1 2 3 4 5 -6 K[j] 60 247 -557 -1092 2220 7314 7314 2220 -1092 -557 247 60 |x[−j] − x[0]| 5 4 3 2 1 0 1 2 3 5 5 6 x[0]) 0.980 0.984 0.988 0.992 0.996 1.0 0.996 0.992 0.988 0.984 0.98 0.976

[0181]

[0182] To demonstrate how the value of (x[— j], x[0]) may be computed, for j = —6: |x[— j] — x[0] | = |x[— (— 6)] — x[0] | = |x[6] — x[0] | = 136 — 30| = 6. Thus, for j = —6:

[0183] h(x[-j],x[0]) = 1 -|x[~7']~x[0] l= 1 -|x[6]-*[0] l= 1 - — « 1 - 0.024 =

[0184]

[0185] k L L J y225 225 255 0.976.

[0186] To demonstrate how the value of S[i] may be computed for i = 0:

[0187] 5[0] = 'h\X

[0188] 0 - ( ^6 ),x[0]) + K;-5 ■ h ^x,x[0]) + - + K 4 7=-6..7^5. 7=-5..7=4.

[0189] 0 - ( 4 ),x[0]) + K 5 ■ h ^x0- ( 5 ),x[0] = K[-6] ■

[0190]

[0191] 7=4..7=5. 7=5. h(x[6],x[0]) + / <[— 5] ■ h(x[5],x[0]) + — I- 7<[4] ■ (x[— 4],x[0]) + 7<[5] ■ (x[— 5],x[0]) = 60 x 0.976 + 247 X 0.98 + ••• + 247 X 0.984 + 60 x 0.98 « 16355.

[0192] Thus, first value of the downsampled signal, i.e. D[0], may be calculated as:D [0] =nfe’ ((x[6]■ *[_6]■ ^Cx[6],x[0])) + x([5] ■ 7<[-5] ■ h(x[5],x[0])) + - + (x[— 4] ■ K[4] ■ h(x[— 4],x[0])) + (x[— 5] ■ / C[5] ■ h(x[-5],x[0]))} = ■ ((36 x 60 x 0.976) + (35 X 247 X 0.98) + ••• + (34 X 247 X 0.984) + (35 X 60 X 0.98))

[0193] 7« ■ (2108 + 8472 + ••• + 8264 + 16355 2058) = ■ (498136) « 30.

[0194] 16355 Thus, in this example, the final value of D[0] = 30. Using the above-indicated alternative convolution operation, the first value of the downsampled signal, D, i.e. D [i] where i = 0, may be computed as follows:

[0195] £)[0] = ■ (498136) V -I « 1(30.46) + -I = [30.96] = 30.

[0196]

[0197] L\16355 J 2J L 2J Thus, using this alternative convolution operation, again the final value of [0] = 30.

[0198] To compute the second value of the downsampled signal, D, i.e. to compute D [i] where i = 1, the following formula may be used:

[0199] D2 x 1^ j — j ■ K [j] ■ h I x 2 x 1 2 x W = ^7 ' 5-6 *

[0200] 1=1 / \ i=l = S;5=-6*[2 ~j] ■ K[j] ■ h(x[2 -j],x[2]).

[0201]

[0202] For ease of understanding, a table showing how h(x[2 — j],x[2]) may be computed for each value of j E [—5,6] is provided below. The table shows the value of |x[2 — j] — x[2] |, and then the corresponding value of h(x[2 — j],x[2]):

[0203] 2 -7 -3 —2 -1 0 1 2 3 4 5 6 7 8 x[2 -j] 33 32 31 30 31 32 33 34 35 36 248 249 j 5 4 3 2 1 0 -1 —2 -3 -4 -5 -6 K[j] 60 247 -557 -109 2220 7314 7314 2220 -109 -557 247 60 |x[2 -j] — x[2]| 1 0 1 2 1 0 1 2 3 4 216 217 h(x[2−j],x[2]) 0.996 1 0.996 0.992 0.996 1 0.996 0.992 0.988 0.984 0.153 0.149

[0204]

[0205] To demonstrate how the value of h(x[2 — j],x[2]) may be computed, for j = -6:

[0206] |x[2 - j] - x[2]| = |x[2 - (-6)] - x[2]| = |x[8] - x[2]| = |249 - 32| = 217.

[0207] Thus, for j = —6:

[0208]

[0209] 0.851 = 0.149.

[0210] To demonstrate how the value of S[i] may be computed for i = 1:

[0211] S[l] = S75=-6 K[j] ■ h x = 7<[-6] ■

[0212] + ••• + K [4] ■

[0213] x ^2 — 3 = 60 x 0.149 + 247 x 0.153 +

[0214]

[0215] =2 — (4) =2 — (5) ••• + 247 X 1.0 + 60 X 0.996 « 16101.

[0216] Thus, the second value of the downsampled signal, i.e. D[l], may be calculated as:

[0217] [1] = 77777’ ((x[8] ■ 7<[-6] ■ h(x[8],x[2])) + (x[7] ■ 7<[-5] ■

[0218]

[0219] h(x[7],x[2])) + — I- (x[— 2] ■ 7<[4] ■ h(x[— 2], x[2])) + (x[— 3] ■ 7<[5] ■

[0220] h(x[-3],x[2])) = 77777' ((249 X 60 X 0.149) + (248 X 247 x 0.153) + ••• + (32 X 247 X 1.0) + (33 X 60 X 0.996)) « 77777 ■ (2226 + 9372 + ••• + 7904 + 1972) = 77777- (532223) « 33.

[0221] Thus, in this example, the final value of D[l] = 33.

[0222] Using the above-indicated alternative convolution operation, the second value of the downsampled signal, D, i.e. £)[i] where i = 1, may be computed as follows:

[0223] D[l] = [(77777’ (532223)) + 7] « [(32.54) +7] = |33.04] = 33.

[0224]

[0225] Thus, using this alternative convolution operation, again the final value of [l] = 33. The final values of D[2], [3], D[4], D[5], and D[6] may be computed in a similar manner, giving D [2] = 32, D [3] = 67, D [4] = 252, D [5] = 251, £)[6] = 254.

[0226] Thus, in this example, the downsampled row, D, is D = [30,33,32,67,252,251,254]. The example method may be performed in a similar manner, on a row-by-row basis, for each row of a multi-row input image.

[0227] Thus, in this example, the similarity function, h, is a function of signal element value differences of a set of signal elements of an input set of signal elements, x, that are at most a predetermined value of the index, i, away from each other in the input set of signal element, x. In particular, in this example, such signal element are never more than N indices away from each other.

[0228] As explained above, generating the downsampled set of signal element, D, may comprise clipping a calculated signal element value for a signal element in the downsampled set of signal element, D, in response to the calculated signal element value being outside a permitted signal element value range. For instance, in this example with 8-bit image content, the permitted pixel value range is [0,255]. An ‘overflowing’ calculated pixel value above ‘255’, such as the value ‘257’, may be clipped to the maximum pixel value in the permitted pixel value range, namely ‘255’. Similarly, an ‘underflowing’ calculated pixel value below ‘O’, such as the value ‘—4’, may be clipped to the minimum pixel value in the permitted pixel value range, namely ‘O’.

[0229] To demonstrate the differences between the downsampled row, d, resulting from the downsampler 302 described above with reference to Figure 3, and the downsampled row, D, resulting from the similarity-driven downsampler 402, both will be compared to the values of the input signal, x, having even indices in the range i E [0,13], i.e. the values x[2i] where i E [0,6]:

[0230] i 0 1 2 3 4 5 6

[0231] 2i 0 2 4 6 8 10 12

[0232] x[2i] 30 32 34 36 249 251 253

[0233] d[i] 30 36 17 142 255 258 254

[0234]

[0235] |d[i] — x[2i]| 0 4 17 106 6 7 1

[0236] D[i] 30 33 32 67 252 251 254

[0237] |D[i] - x[2i]| 0 1 2 31 3 0 1

[0238] |d[i] — x[2i]| — |£)[i] — x[2i]| 0 3 15 75 3 7 0

[0239]

[0240] Thus, for all values of i E [0,6], £)[i] is closer than d[i] to, or is the same distance as d[i] from, the corresponding value of the input row signal, x, i.e. the value x[2i], For i = 0 and i = 6, £)[i] is the same distance as d[i] from x[0] and x

[0012] respectively. For all other values of i E [0,6], £)[i] is closer than d[i] to the corresponding value of the input row signal, x.

[0241] In principle, to minimise the differences between signal element values of a downsampled signal and corresponding signal element values of an input signal, the values of those corresponding signal elements could be used directly for the signal element values of the downsampled signal. For example, for an input signal in the form [0,255,0,255,0,255,0,255], the signal element values in even positions could be selected and used for the downsampled signal, giving [0,0, 0,0]. Such an input signal may correspond to an image with alternating black and white vertical lines. Such a downsampled image may correspond to an entirely black image. The pixel values of such a downsampled image are the same as the corresponding pixel values of the input image such that the differences between those values is zero. However, such a downsampled signal does not correspond to an effective downsampled representation of the input signal. This is because the content relating to the white lines is lost in the downsampled image. The similarity-driven downsampler provides effective downsampling in this regard.

[0242] With reference again to the above definition of a modified kernel, K', in which K' = K ■ h, up to W x H different modified kernels, K', may be used in respect of input data, x, having a width ‘IV’ and height ‘H’, or more generally having W X H signal elements. In particular, a modified kernel, K', may be determined on a per-pixel basis (or, more generally, a per-signal-element basis).

[0243] This differs from downsampling that downsamples input data multiple times using a small number (for example, at most two) of different kernels and selects the preferred downsampled pixel on a pixel -by pixel basis. In contrast, the similarity-driven downsampler 402 provides content-adaptive modification of a kernel, K, using the similarity function, h. Additionally, significantly more kernels may be used per frame, for fine-tuning.

[0244] This also differs from a downsampler in which coefficients of a kernel, K, are changed on a per-frame basis. In such a downsampler, the kernel, K, does not change across any given frame, for example on a per-row or per-pixel basis. The similarity-driven downsampler 402 optimises downsampling of the input signal, x, based on the content of the input signal, x, for example by modifying a Lanczos kernel, K, per-pixel based on similarity of surrounding pixels to a target pixel in an input signal, x.

[0245] Thus, various example downsampling methods are provided. The example methods may be implemented in the example signal processing system 400 or otherwise.

[0246] In examples, an input set of pixels 404, x is obtained, for example by the downsampler 402. More generally, an input set of signal elements 404 may be obtained. The input set of pixels 404, x, may correspond to the input signal 404. The input set of pixels 404, x, may comprise all pixels of a source image. Alternatively, the input set of pixels 404, x, may include only a subset of pixels of the source image. For example, the subset of pixels may correspond to a row of pixels of the source image, a group of neighboring pixels of the source image, or otherwise.

[0247] In examples, a downsampled set of pixels 410, D, is generated, for example by the downsampler 402, by downsampling the input set of pixels 404, x.

[0248] In examples, the downsampling uses a similarity function 412, h. For example, the downsampler 402 may obtain the similarity function 412, h. The similarity function 412, h, may be a function of pixel value differences of pixels of the input set of pixels 404, x. Such pixels may be nearby pixels (in relation to a target pixel). The number of such pixels may be based on the size, N, of a kernel 406, K.

[0249] The downsampled set of pixels 410, D, may be upsampled to generate a predicted rendition 412, U, of the input set of pixels 404, x. A set of residuals 416 may be generated (for example, by a residual generator 414) based on differences between the input set of pixels 404, x, and the predicted rendition 412, U, of the input set of pixels 404, x.

[0250] The downsampled set of pixels 410, D, may be output to an encoder. The encoder may be configured to generate an encoded set of pixels by encoding the downsampled set of pixels 410, D.

[0251] A decoded version of the encoded set of pixels may be obtained. Correction data may be generated based on differences between the downsampled set of pixels 410, D, and the decoded version of the encoded set of pixels.

[0252] The downsampled set of pixels 410, D, may be generated according to a - N 1

[0253] convolution operation 408: £)[i] = —2Nx[Ai — j] ■ K[j] ■ h x[Ai — j],x[ i]). As S[t] j = --explained above, £)[i] may be the downsampled set of pixels 410, D, at index i. x[Ai — j] may be the input set of pixels 404, x, at calculated index Ai — j. The index Ai — j may be considered to be a “calculated” index in that the value of Ai — j may be outside the range of indices for the input set of pixels, x. K[j may be a kernel coefficient of a kernel 406, K, at index j. h x[Ai — j],x[ i]) may be the similarity function 412, h, of the input set of pixels 404, x, at the calculated index Ai — j, and the input set of pixels 404, x, at index Ai. A may be a downsampling factor. A may be a length of the kernel 406, K. The (modified) kernel sum at index i, 5[i], may be defined as: S

[0254]

[0255] [i] = K[j] ■ h(x[Ai — j],x[Ai]).

[0256] In some examples, A = 2.

[0257] The kernel 406, K, may be a symmetric kernel. The kernel 406, K, may be a Lanczos kernel. A Lanczos kernel is an example of a symmetric kernel.

[0258] An extended set of pixel values for x[Ai — j may be defined for when the calculated index Ai — j is not a defined index within the input set of pixels 404, x. A pixel value from the extended set of pixel values for x[Ai — j may be used when the calculated index Ai — j is not a defined index within the input set of pixels 404, x.

[0259] The pixel values in the extended set of pixel values may be zero-filled.

[0260] The first pixel value in the input set of pixels 404, x, may be used for all calculated indices Ai — j before the first defined index. The final pixel value in the input set of pixels 404, x, may be used for all calculated indices Ai — j after the final defined index.

[0261] The pixel values in the extended set of pixel values may be mirrored with respect to the input set of pixels 404, x.

[0262] The pixel values in the extended set of pixel values may be wrapped with respect to the input set of pixels 404, x.

[0263] The similarity function 418, h, may be a function of pixel value differences of a set of pixels of the input set of pixels 404, x, that are at most a predetermined value of the index, i, away from each other in the input set of pixels 404, x.

[0264] The similarity function 418, h x[Ai — j],x[Ai]) between the input set of pixels 404, x, at calculated index Ai — j, and the input set of pixels 404, x, at index i may be 7 / " r 1 ■ *n r 1 -4— / I I 1 • • 1 1 • 1 II(X[AI —

[0265]

[0266] X[AI]J = 1 - — -, where m represents a maximum possible pixel value difference.

[0267] In some examples, m = 2b— 1. b is a bit-length of the pixels in the input set of pixels 404, x. In some examples, m = 255. However, m can take other values. For example, for 16-bit signal element values, m=216— 1 = 65535.

[0268] In some examples, the similarity function 418, h, is a linear similarity function. In some examples, the similarity function 418, h, is a non-linear similarity function. For example, the non-linear similarity function may be a logarithmic function, an exponential function, a sigmoid function, or a step function.

[0269] Generating the downsampled set of pixels 410, D, may comprise clipping a calculated pixel value for a pixel in the downsampled set of pixels 410, D, in response to the calculated pixel value being outside a permitted pixel value range.

[0270] The input set of pixels 404, x, may correspond to a row of an input image, and the method may be performed on a row-by-row basis for the rows of the input image.

[0271] Another example similarity-driven downsampling method is provided. An input signal 404, x, is obtained. A downsampled signal 410, D, is generated according to a i — ±i

[0272] convolution operation 408 defined as: £)[i] = — ‘ S2wx[Ai -j] ■ K[j] ■ (x[Ai -

[0273]

[0274] SpJ j=—

[0275] 1 1

[0276] j],x[Ai]) 1 + -. £)[i] is the downsampled signal 410, D, at index i. x[Ai — j] is the

[0277]

[0278] input signal 404, x, at calculated index Ai — j. K[j] is a kernel coefficient of a kernel 406, K, at index j. h x[ i — j],x[Ai]) is a similarity function 418 between the input signal 404, x, at calculated index i — j, and the input signal 404, x, at index Ai. is a downscaling factor. N is a length of the kernel 406, K. S[i] = j ^[ / ] ‘ h(x[ i — j],x[Ai]), where S[i] is a normalisation factor at index i.

[0279] Another example similarity-driven downsampling method is provided. An input signal 404, x, is obtained. The input signal 404, x, is downsampled to generate a downsampled signal 410, D. The downsampling is dependent on signal element value differences of signal elements of the input signal 404, x.

[0280] Apparatus configured to perform any such method(s) is provided.

[0281] A computer program configured to perform any such method(s) is provided. Referring to Figure 5, there is shown an example of a graph 500 of example similarity functions, h. An example linear similarity function, h, is shown using a solid line 502. An example non-linear similarity function, h, is shown using a broken line 504. Examples of non-linear similarity functions include, but are not limited to, logarithmic functions, exponential functions, sigmoid functions, and step functions.

[0282] Referring to Figure 6, there is shown a schematic block diagram of an example of an apparatus 600.

[0283] In an example, the apparatus 600 comprises an encoder. In another example, the apparatus 600 comprises a decoder. In other examples, the apparatus 600 comprises neither an encoder nor a decoder but is configured to communicate with an encoder and / or a decoder.

[0284] Examples of apparatus 600 include, but are not limited to, a mobile computer, a personal computer system, a wireless device, base station, phone device, desktop computer, laptop, notebook, netbook computer, mainframe computer system, handheld computer, workstation, network computer, application server, storage device, a consumer electronics device such as a camera, camcorder, mobile device, video game console, handheld video game device, an XR headset, or in general any type of computing or electronic device.

[0285] In this example, the apparatus 600 comprises one or more processors 601 configured to process information and / or instructions. The one or more processors 601 may comprise a CPU. The one or more processors 601 are coupled with a bus 602. Operations performed by the one or more processors 601 may be carried out by hardware and / or software. The one or more processors 601 may comprise multiple colocated processors or multiple disparately located processors.

[0286] In this example, the apparatus 600 comprises computer-useable volatile memory 603 configured to store information and / or instructions for the one or more processors 601. The computer-useable volatile memory 603 is coupled with the bus 602. The computer-useable volatile memory 603 may comprise random access memory (RAM).

[0287] In this example, the apparatus 600 comprises computer-useable non-volatile memory 604 configured to store information and / or instructions for the one or more processors 601. The computer-useable non-volatile memory 604 is coupled with the bus 602. The computer-useable non-volatile memory 604 may comprise read-only memory (ROM).

[0288] In this example, the apparatus 600 comprises one or more data-storage units 605 configured to store information and / or instructions. The one or more data-storage units 605 are coupled with the bus 602. The one or more data-storage units 605 may for example comprise a magnetic or optical disk and disk drive or a solid-state drive (SSD).

[0289] In this example, the apparatus 600 comprises one or more input / output (I / O) devices 606 configured to communicate information to and / or from the one or more processors 601. The one or more I / O devices 606 are coupled with the bus 602. The one or more I / O devices 606 may comprise at least one network interface. The at least one network interface may enable the apparatus 600 to communicate via one or more data communications networks. Examples of data communications networks include, but are not limited to, the Internet and a Local Area Network (LAN). The one or more I / O devices 606 may enable a user to provide input to the apparatus 600 via one or more input devices (not shown). The one or more input devices may include for example a remote control, one or more physical buttons etc. The one or more I / O devices 606 may enable information to be provided to a user via one or more output devices (not shown). The one or more output devices may for example include a display screen.

[0290] Various other entities are depicted for the apparatus 600. For example, when present, an operating system 607, image processing module 608, one or more further modules 609, and data 610 are shown as residing in one, or a combination, of the computer-usable volatile memory 603, computer-usable non-volatile memory 604 and the one or more data-storage units 605. The data signal processing module 608 may be implemented by way of computer program code stored in memory locations within the computer-usable non-volatile memory 604, computer-readable storage media within the one or more data-storage units 605 and / or other tangible computer-readable storage media. Examples of tangible computer-readable storage media include, but are not limited to, an optical medium (e.g., CD-ROM, DVD-ROM or Blu-ray), flash memory card, floppy or hard disk or any other medium capable of storing computer -readable instructions such as firmware or microcode in at least one ROM or RAM or Programmable ROM (PROM) chips or as an Application Specific Integrated Circuit (ASIC).

[0291] The apparatus 600 may therefore comprise a data signal processing module 608 which can be executed by the one or more processors 601. The data signal processing module 608 can be configured to include instructions to implement at least some of the operations described herein. During operation, the one or more processors 601 launch, run, execute, interpret or otherwise perform the instructions in the signal processing module 608.

[0292] Although at least some aspects of the examples described herein with reference to the drawings comprise computer processes performed in processing systems or processors, examples described herein also extend to computer programs, for example computer programs on or in a carrier, adapted for putting the examples into practice. The carrier may be any entity or device capable of carrying the program.

[0293] It will be appreciated that the apparatus 600 may comprise more, fewer and / or different components from those depicted in Figure 6.

[0294] The apparatus 600 may be located in a single location or may be distributed in multiple locations. Such locations may be local or remote.

[0295] The techniques described herein may be implemented in software or hardware, or may be implemented using a combination of software and hardware. They may include configuring an apparatus to carry out and / or support any or all of techniques described herein. It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

[0296] For example, while certain mathematical notation has been used herein, other computations using different mathematical notation may serve the same purpose and may not depart from the scope of the invention.

Claims

CLAIMS1. A downsampling method comprising:obtaining an input set of pixels, x; andgenerating a downsampled set of pixels, D, by downsampling the input set of pixels, x,wherein the downsampling uses a similarity function, h, andwherein the similarity function, h, is a function of pixel value differences of pixels of the input set of pixels, x.

2. A method according to claim 1, comprising:upsampling the downsampled set of pixels, D, to generate a predicted rendition of the input set of pixels, x; andgenerating a set of residuals based on differences between the input set of pixels, x, and the predicted rendition of the input set of pixels, x.

3. A method according to claim 1 or 2, comprising:outputting the downsampled set of pixels, D, to an encoder, wherein the encoder is configured to generate an encoded set of pixels by encoding the downsampled set of pixels, D.

4. A method according to claim 3, comprising:obtaining a decoded version of the encoded set of pixels; andgenerating correction data based on differences between the downsampled set of pixels, D, and the decoded version of the encoded set of pixels.

5. A method according to any of claims 1 to 4, wherein the downsampled set of pixels, D, is generated according to a convolution operation--i0[i] = 77- 12Nx[M - j] ■ K[j] ■ h(x[Ai -j],x[Ai]),SpJJ =--wherein £)[i] is the downsampled set of pixels, D, at index i,wherein x[ i — j] is the input set of pixels, x, at calculated index i — j, wherein K[j] is a kernel coefficient of a kernel, K, at index j,wherein h(x[Ai — j],x[Ai]) is the similarity function, h, of the input set of pixels, x, at the calculated index Ai — j, and the input set of pixels, x, at index Ai, wherein is a downsampling factor,wherein N is a length of the kernel, K, andwherein S[i] = K[j] ■ h(x[Ai — j],x[Ai]).

6. A method according to claim 5, wherein the kernel, K, is a symmetric kernel.

7. A method according to claim 6, wherein the kernel, K, is a Lanczos kernel.

8. A method according to any of claims 5 to 7, comprising:defining an extended set of pixel values for x[ i — j] for when the calculated index Ai — j is not a defined index within the input set of pixels, x; andusing a pixel value from the extended set of pixel values for x[ i — j] when the calculated index Ai — j is not a defined index within the input set of pixels, x.

9. A method according to claim 8, wherein the pixel values in the extended set of pixel values are zero-filled.

10. A method according to claim 8, wherein the first pixel value in the input set of pixels, x, is used for all calculated indices Ai — j before the first defined index, and wherein the final pixel value in the input set of pixels, x, is used for all calculated indices Ai — j after the final defined index.

11. A method according to claim 8, wherein the pixel values in the extended set of pixel values are mirrored with respect to the input set of pixels, x.

12. A method according to claim 8, wherein the pixel values in the extended set of pixel values are wrapped with respect to the input set of pixels, x.

13. A method according to any of claims 5 to 12, wherein the similarity function, h, is a function of pixel value differences of a set of pixels of the input set of pixels, x, that are at most a predetermined value of the index, i, away from each other in the input set of pixels, x.

14. A method according to any of claims 5 to 14, wherein the similarity function, h x[ i — j],x[Ai]) between the input set of pixels, x, at calculated index i — j, andthe input set of pixels, x, at index i is / i(x[ i — j],x[Ai]) = 1 — wherein m represents a maximum possible pixel value difference.

15. A method according to claim 14, wherein m = 2b— 1, wherein b is a bit-length of the pixels in the input set of pixels, x.

16. A method according to any of claims 5 to 15, wherein = 2.

17. A method according to any of claims 1 to 16, wherein the similarity function, h, is a linear similarity function.

18. A method according to any of claims 1 to 13, wherein the similarity function, h, is a non-linear similarity function.

19. A method according to claim 18, wherein the non-linear similarity function is a logarithmic function.

20. A method according to claim 18, wherein the non-linear function is an exponential function.

21. A method according to claim 18, wherein the non-linear function is a sigmoid function.

22. A method according to claim 18, wherein the non-linear function is a step function.

23. A method according to any of claims 1 to 22, wherein generating the downsampled set of pixels, D, comprises clipping a calculated pixel value for a pixel in the downsampled set of pixels, D, in response to the calculated pixel value outside a permitted pixel value range.

24. A method according to any of claims 1 to 23, wherein the input set of pixels, x, corresponds to a row of an input image, and wherein the method is performed on a row-by-row basis for the rows of the input image.

25. A similarity-driven downsampling method comprising:obtaining an input signal, x; andgenerating a downsampled signal, D, according to a convolution operation£>[*] = 77^’ S2Nx[Ai — j] - K[f\ - h(x[Ai — j],x[Ai]') + -|,j=-~ y2wherein £)[i] is the downsampled signal, D, at index i,wherein x[ i — j] is the input signal, x, at calculated index i — j, wherein K[j] is a kernel coefficient of a kernel, K, at index j,wherein / i(x[ i — j],x[Ai]) is a similarity function between the input signal, x, at calculated index i — j, and the input signal, x, at index Ai,wherein is a downsampling factor,wherein A is a length of the kernel, K, andwherein S[i] = K[j] ■ h(x[Ai — j],x[Ai]).

26. A similarity-driven downsampling method comprising:obtaining an input signal, x; anddownsampling the input signal, x, to generate a downsampled signal, D, wherein the downsampling is dependent on signal element value differences of signal elements of the input signal, x.

27. Apparatus configured to perform a method according to any of claims 1 to 26.

28. A computer program configured to perform a method according to any of claims 1 to 26.