Edge-driven image generation

Edge-driven downsampling addresses halos in image processing by adaptively selecting pixel processing based on edge strength, enhancing XR image quality and reducing computational demands.

WO2026104839A1PCT designated stage Publication Date: 2026-05-21V NOVA INT LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
V NOVA INT LTD
Filing Date
2025-11-14
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Visible artefacts, such as halos, arise during image processing in tiered hierarchy signal processing systems, particularly in high-efficiency video coding, which can be exacerbated in XR applications due to magnification and amplification of even small artefacts.

Method used

Implement edge-driven downsampling by extracting an edge map from the input image and processing it twice with different filters and downsamplers to select the appropriate processed version for each pixel based on edge strength, reducing halos while maintaining visual quality.

Benefits of technology

Reduces visible artefacts like halos effectively, especially in XR applications, by optimizing image processing to maintain high visual quality with reduced computational resources and latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025052497_21052026_PF_FP_ABST
    Figure GB2025052497_21052026_PF_FP_ABST
Patent Text Reader

Abstract

An edge map 506 of a first set of image elements 502 is obtained. An edge map element 508 indicates an edge characteristic of one or more corresponding image elements 504 in the first set 502. A second set of image elements 510 is generated. For an image element 512 in the second set 510: (i) one or more corresponding edge map elements 508 are identified; (ii) one or more corresponding image elements 518, 520 of a third and / or fourth set of image elements 514, 516 are selected based on the identified one or more corresponding edge map elements 508; and (iii) the image element 512 in the second set 510 is generated based on the selected one or more corresponding image elements 518, 520. The third and fourth sets 514, 516 have been generated as a result of some or all of the first set 502 having been processed using first and second, different processing respectively.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] EDGE-DRIVEN IMAGE GENERATION

[0002] Technical Field

[0003] The present disclosure relates to edge-driven image generation.

[0004] Background

[0005] Visible artefacts can arise when images are processed. An example of such an artefact is a halo artefact, or simply a “halo”. A halo is a type of ringing. A halo can be in the form of a light or dark ring around an identifiable object. Such artefacts can appear in a base layer in a tiered hierarchy signal processing system that uses hierarchical codecs. An example of such abase layer is a High-Efficiency Video Coding (HEVC) base layer. HEVC is also known as H.265.

[0006] Summary

[0007] Various aspects of the present disclosure are set out in the appended claims. Further features and advantages will become apparent from the following description of preferred embodiments, given by way of example only, which is made with reference to the accompanying drawings.

[0008] Brief Description of the Drawings

[0009] Figure 1 shows a schematic block diagram of an example of an image processing system;

[0010] Figures 2A and 2B show schematic block diagrams of another example of an image processing system;

[0011] Figure 3 shows a flowchart of an example of an image processing method; Figure 4 shows a flowchart of another example of an image processing method; Figure 5 shows a schematic diagram of an example of edge-driven downsampling;

[0012] Figures 6A and 6B show schematic diagrams of another example of edge-driven downsampling;

[0013] Figures 7A and 7B show schematic diagrams of another example of edge-driven downsampling; Figure 8 shows a schematic diagram of another example of edge-driven downsampling; and

[0014] Figure 9 shows a schematic block diagram of an example of an apparatus.

[0015] Detailed Description

[0016] Referring to Figure 1, there is shown an example of a signal processing system 100. The signal processing system 100 is used to process signals. Examples of types of signal include, but are not limited to, video signals, image signals, audio signals, volumetric signals such as those used in medical, scientific or holographic imaging, or other multidimensional signals.

[0017] The signal processing system 100 includes a first apparatus 102 and a second apparatus 104. The first apparatus 102 and second apparatus 104 may have a clientserver relationship, with the first apparatus 102 performing the functions of a server device and the second apparatus 104 performing the functions of a client device. The signal processing system 100 may include at least one additional apparatus (not shown). The first apparatus 102 and / or second apparatus 104 may comprise one or more components. The one or more components may be implemented in hardware and / or software. The one or more components may be co-located or may be located remotely from each other in the signal processing system 100. Examples of types of apparatus include, but are not limited to, computerised devices, handheld or laptop computers, tablets, mobile devices, games consoles, smart televisions, set-top boxes, Extended Reality (XR) headsets (including Augmented Reality (AR) and / or Virtual Reality (VR) headsets) etc.

[0018] The first apparatus 102 is communicatively coupled to the second apparatus 104 via a data communications network 106. Examples of the data communications network 106 include, but are not limited to, the Internet, a Local Area Network (LAN) and a Wide Area Network (WAN). The first and / or second apparatus 102, 104 may have a wired and / or wireless connection to the data communications network 106.

[0019] In this example, the first apparatus 102 comprises an encoder 108. The encoder 108 is configured to encode data comprised in and / or derived based on the signal, which is referred to hereinafter as “signal data”. For example, where the signal is a video signal, the encoder 108 is configured to encode video data. Video data comprises a sequence of multiple images or frames. The encoder 108 may perform one or more further functions in addition to encoding signal data. The encoder 108 may be embodied in various different ways. For example, the encoder 108 may be embodied in hardware and / or software. The encoder 108 may encode metadata associated with the signal. The first apparatus 102 may use one or more than one encoder 108.

[0020] Although in this example the first apparatus 102 comprises the encoder 108, in other examples the first apparatus 102 is separate from the encoder 108. In such examples, the first apparatus 102 is communicatively coupled to the encoder 108. The first apparatus 102 may be embodied as one or more software functions and / or hardware modules.

[0021] In this example, the second apparatus 104 comprises a decoder 110. The decoder 110 is configured to decode signal data. The decoder 110 may perform one or more further functions in addition to decoding signal data. The decoder 110 may be embodied in various different ways. For example, the decoder 110 may be embodied in hardware and / or software. The decoder 110 may decode metadata associated with the signal. The second apparatus 104 may use one or more than one decoder 110.

[0022] Although in this example the second apparatus 104 comprises the decoder 110, in other examples the second apparatus 104 is separate from the decoder 110. In such examples, the second apparatus 104 is communicatively coupled to the decoder 110. The second apparatus 104 may be embodied as one or more software functions and / or hardware modules.

[0023] The encoder 108 encodes signal data and transmits the encoded signal data to the decoder 110 via the data communications network 106. The decoder 110 decodes the received, encoded signal data and generates decoded signal data. The decoder 110 may output the decoded signal data, or data derived using the decoded signal data. For example, the decoder 110 may output such data for display on one or more display devices associated with the second apparatus 104. The one or more display devices may be components of the second apparatus 104 or may otherwise be associated with the second apparatus 104. The one or more display devices may be operable to display XR content and may, therefore, be referred to as XR display devices.

[0024] In some examples described herein, the encoder 108 transmits to the decoder 110 a representation of a signal at a given level of quality and information the decoder 110 can use to reconstruct a representation of some or all of the signal at one or more higher levels of quality. Such information may be referred to as “reconstruction data”. In some examples, “reconstruction” of a representation involves obtaining a representation that is not an exact replica of an original representation. The extent to which the representation is the same as the original representation may depend on various factors including, but not limited to, quantisation levels. A representation of a signal at a given level of quality may be considered to be a rendition, version or depiction of data comprised in the signal at the given level of quality. In some examples, the reconstruction data is included in the signal data that is encoded by the encoder 108 and transmitted to the decoder 110. For example, the reconstruction data may be in the form of metadata. In some examples, the reconstruction data is encoded and transmitted separately from the signal data.

[0025] The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may comprise residual data, as described in more detail below. Residual data is an example of reconstruction data. The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may also comprise configuration data relating to processing of the residual data. The configuration data may indicate how the residual data has been processed by the encoder 108 and / or how the residual data is to be processed by the decoder 110. The configuration data may be signalled to the decoder 110, for example in the form of metadata.

[0026] The first and / or second apparatuses 102, 104 may be configured to perform some or all of the techniques described herein. A computer program may be configured to perform some or all of the techniques described herein.

[0027] Referring to Figures 2A and 2B, there is shown schematically an example of a signal processing system 200. The signal processing system 200 includes a first apparatus 202 and a second apparatus 204. In this example, the first apparatus 202 comprises an encoder and the second apparatus 204 comprises a decoder. However, as explained above, in other examples, the encoder is not comprised in the first apparatus 202 and / or the decoder is not comprised in the second apparatus 204. In each of the first apparatus 202 and the second apparatus 204, items are shown on two logical levels. The two levels are separated by a dashed line. Items on the first, highest level relate to data at a first level of quality. Items on the second, lowest level relate to data at a second level of quality. The first level of quality is higher than the second level of quality. The first and second levels of quality relate to a tiered hierarchy having multiple levels of quality. In some examples, the tiered hierarchy comprises more than two levels of quality. In such examples, the first apparatus 202 and the second apparatus 204 may include more than two different levels. There may be one or more other levels above and / or below those depicted in Figures 2A and 2B. As described herein, in certain cases, the levels of quality may correspond to different spatial resolutions.

[0028] Referring first to Figure 2A, the first apparatus 202 obtains a first representation of an image at the first level of quality 206. A representation of a given image is a representation of data comprised in the image. The image may be a given frame of a video. The first representation of the image at the first level of quality 206 will be referred to as “input data” hereinafter as, in this example, it is data provided as an input to the encoder in the first apparatus 202. The first apparatus 202 may receive the input data 206. For example, the first apparatus 202 may receive the input data 206 from at least one other apparatus. The first apparatus 202 may be configured to receive successive portions of input data 206, e.g. successive frames of a video, and to perform the operations described herein to each successive frame. For example, a video may comprise frames Fi, F2, ... FT and the first apparatus 202 may process each of these in turn.

[0029] The first apparatus 202 derives data 212 based on the input data 206. In this example, the data 212 based on the input data 206 is a representation 212 of the image at the second, lower level of quality. In this example, the data 212 is derived by performing a downsampling operation on the input data 206 and will therefore be referred to as “downsampled data” hereinafter. In other examples, the data 212 is derived by performing an operation other than a downsampling operation on the input data 206, or the data 212 is the same as the input data 206 (i.e. the input data 206 is not processed, e.g. downsampled).

[0030] In this example, the downsampled data 212 is processed to generate processed data 213 at the second level of quality. In other examples, the downsampled data 212 is not processed at the second level of quality. As such, the first apparatus 202 may generate data at the second level of quality, where the data at the second level of quality comprises the downsampled data 212 or the processed data 213.

[0031] In some examples, generating the processed data 213 involves the downsampled data 212 being encoded. Such encoding may occur within the first apparatus 202, or the first apparatus 202 may output the processed data 213 to an external encoder. Encoding the downsampled data 212 produces an encoded image at the second level of quality. The first apparatus 202 may output the encoded image, for example for transmission to the second apparatus 204. A series of encoded images, e.g. forming an encoded video, as output for transmission to the second apparatus 204 may be referred to as a “base” stream or “base” layer. As explained above, instead of being produced in the first apparatus 202, the encoded image may be produced by an encoder that is separate from the first apparatus 202. The encoded image may be part of an H.264 or H.265 encoded video, or otherwise. Generating the processed data 213 may, for example, comprise generating successive frames of video as output by a separate encoder such as an H.264 or H.265 video encoder. An intermediate set of data for the generation of the processed data 213 may comprise the output of such an encoder, as opposed to any intermediate data generated by the separate encoder.

[0032] Generating the processed data 213 at the second level of quality may further involve decoding the encoded image at the second level of quality. The decoding operation may be performed to emulate a decoding operation at the second apparatus 204, as will become apparent below. Decoding the encoded image produces a decoded image at the second level of quality. In some examples, the first apparatus 202 decodes the encoded image at the second level of quality to produce the decoded image at the second level of quality. In other examples, the first apparatus 202 receives the decoded image at the second level of quality, for example from an encoder and / or decoder that is separate from the first apparatus 202. The encoded image may be decoded using an H.264 or H.265 decoder. The decoding by a separate decoder may comprise inputting encoded video, such as an encoded data stream configured for transmission to a remote decoder, into a separate black-box decoder implemented together with the first apparatus 202 to generate successive decoded frames of video. Processed data 213 may thus comprise a frame of video data that is generated via a complex non-linear encoding and decoding process, where the encoding and decoding process may involve modelling spatio-temporal correlations as per a particular encoding standard such as H.264 or H.265. However, because the output of any encoder is fed into a corresponding decoder, this complexity is effectively hidden from the first apparatus 202.

[0033] In an example, generating the processed data 213 at the second level of quality further involves obtaining correction data based on a comparison between the downsampled data 212 and the decoded image obtained by the first apparatus 202, for example based on the difference between the downsampled data 212 and the decoded image. The correction data can be used to correct for encoder-decoder errors (which may also be referred to as “encode-decode errors”), namely errors introduced in encoding and decoding the downsampled data 212. In some examples, the first apparatus 202 outputs the correction data, for example for transmission to the second apparatus 204, as well as the encoded signal. This allows the recipient to correct for the encoder-decoder errors introduced in encoding and decoding the downsampled data 212. This correction data may also be referred to as a “first enhancement” stream. As the correction data may be based on the difference between the downsampled data 212 and the decoded image it may be seen as a form of residual data (e.g. that is different from the other set of residual data described later below). An item of residual data may be referred to as “a residual”.

[0034] In some examples, generating the processed data 213 at the second level of quality further involves correcting the decoded image using the correction data. For example, the correction data as output for transmission may be placed into a form suitable for combination with the decoded image, and then added to the decoded image. This may be performed on a frame-by-frame basis. In other examples, rather than correcting the decoded image using the correction data, the first apparatus 202 uses the downsampled data 212. For example, in certain cases, just the encoded then decoded data may be used and, in other cases, encoding and decoding may be replaced by other processing.

[0035] In some examples, generating the processed data 213 involves performing one or more operations other than the encoding, decoding, obtaining, and correcting acts described above.

[0036] The first apparatus 202 obtains data 214 based on the data at the second level of quality. As indicated above, the data at the second level of quality may comprise the processed data 213, or the downsampled data 212 where the downsampled data 212 is not processed at the lower level. As described above, in certain cases, the processed data 213 may comprise a reconstructed video stream (e.g. from an encoding-decoding operation) that is corrected using correction data. In the example of Figures 2A and 2B, the data 214 is a second representation of the image at the first level of quality, the first representation of the image at the first level of quality being the input data 206. The second representation at the first level of quality may be considered to be a preliminary or predicted representation of the image at the first level of quality. In this example, the first apparatus 202 derives the data 214 by performing an upsampling operation on the data at the second level of quality. The data 214 will be referred to hereinafter as “upsampled data”. However, in other examples one or more other operations could be used to derive the data 214, for example where data 212 is not derived by downsampling the input data 206.

[0037] The input data 206 and the upsampled data 214 are used to obtain residual data 216. The residual data 216 is associated with the image. The residual data 216 may be in the form of a set of residual elements, which may be referred to as a “residual frame” or a “residual image”. A residual element may be referred to as “a residual”. A residual element in the set of residual elements 216 may be associated with a respective image element in the input data 206. An example of an image element is a pixel.

[0038] In this example, a given residual element is obtained by subtracting a value of an image element in the upsampled data 214 from a value of a corresponding image element in the input data 206. As such, the residual data 216 is useable in combination with the upsampled data 214 to reconstruct the input data 206. The residual data 216 may also be referred to as “reconstruction data” or “enhancement data”. In one case, the residual data 216 may form part of a “second enhancement” stream. The residual data 216 may therefore result from upsampler-downsampler asymmetry. Upsampler-downsampler asymmetry may also be referred to as “upsampling-downsampling asymmetry”, “upsample-downsample” asymmetry or the like.

[0039] The first apparatus 202 obtains configuration data relating to processing of the residual data 216. The configuration data indicates how the residual data 216 has been processed and / or generated by the first apparatus 202 and / or how the residual data 216 is to be processed by the second apparatus 204. The configuration data may comprise a set of configuration parameters. The configuration data may be useable to control how the second apparatus 204 processes data and / or reconstructs the input data 206 using the residual data 216. The configuration data may relate to one or more characteristics of the residual data 216. The configuration data may relate to one or more characteristics of the input data 206. Different configuration data may result in different processing being performed on and / or using the residual data 216. The configuration data is therefore useable to reconstruct the input data 206 using the residual data 216. As described below, in certain cases, configuration data may also relate to the correction data described herein.

[0040] In this example, the first apparatus 202 transmits to the second apparatus 204 data based on the downsampled data 212, data based on the residual data 216, and the configuration data (or data based on the configuration data), to enable the second apparatus 204 to reconstruct the input data 206.

[0041] Turning now to Figure 2B, the second apparatus 204 receives data 220 based on (e.g. derived from) the downsampled data 212. The second apparatus 204 also receives data based on the residual data 216. For example, the second apparatus 204 may receive a “base” stream (data 220), a “first enhancement stream” (any correction data) and a “second enhancement stream” (residual data 216). The base stream may be referred to as a “base layer”. The first and / or second enhancement stream may be referred to as, and / or may be comprised in, an “enhancement layer”. The second apparatus 204 also receives the configuration data relating to processing of the residual data 216. The data 220 based on the downsampled data 212 may be the downsampled data 212 itself, the processed data 213, or data derived from the downsampled data 212 or the processed data 213. The data based on the residual data 216 may be the residual data 216 itself, or data derived from the residual data 216.

[0042] In some examples, the received data 220 comprises the processed data 213, which may comprise the encoded image at the second level of quality and / or the correction data. In some examples, for example where the first apparatus 202 has processed the downsampled data 212 to generate the processed data 213, the second apparatus 204 processes the received data 220 to generate processed data 222. Such processing by the second apparatus 204 may comprise decoding an encoded image (e.g. that forms part of a “base” encoded video stream) to produce a decoded image at the second level of quality. In some examples, the processing by the second apparatus 204 comprises correcting the decoded image using obtained correction data. Hence, the processed data 222 may comprise a frame of corrected data at the second level of quality. In some examples, the encoded image at the second level of quality is decoded by a decoder that is separate from the second apparatus 204. The encoded image at the second level of quality may be decoded using an H.264 decoder.

[0043] In other examples, the received data 220 comprises the downsampled data 212 and does not comprise the processed data 213. In some such examples, the second apparatus 204 does not process the received data 220 to generate processed data 222.

[0044] The second apparatus 204 uses data at the second level of quality to derive the upsampled data 214. As indicated above, the data at the second level of quality may comprise the processed data 222, or the received data 220 where the second apparatus 204 does not process the received data 220 at the second level of quality. The upsampled data 214 is a preliminary representation of the image at the first level of quality. The upsampled data 214 may be derived by performing an upsampling operation on the data at the second level of quality.

[0045] The second apparatus 204 obtains the residual data 216. The residual data 216 is useable with the upsampled data 214 to reconstruct the input data 206. The residual data 216 is indicative of a comparison between the input data 206 and the upsampled data 214.

[0046] The second apparatus 204 also obtains the configuration data related to processing of the residual data 216. The configuration data is useable by the second apparatus 204 to reconstruct the input data 206. For example, the configuration data may indicate a characteristic or property relating to the residual data 216 that affects how the residual data 216 is to be used and / or processed, or whether the residual data 216 is to be used at all. In some examples, the configuration data comprises the residual data 216.

[0047] There are several considerations relating to such processing. One such consideration is the amount of information that is generated, stored, transmitted and / or processed. The more information that is used, the greater the amount of resources that may be involved in handling such information. Examples of such resources include transmission resources, storage resources and processing resources. Some signal processing techniques allow a relatively small amount of information to be used. This may reduce the amount of data transmitted via the data communications network 106. The savings may be particularly relevant where the data relates to high quality video data, where the amount of information transmitted can be especially high.

[0048] Another consideration is latency. Complex image processing may introduce latency, which may negatively impact performance.

[0049] Other considerations include the ability of the decoder to perform image reconstruction accurately, reliably, and / or efficiently. Performing image reconstruction accurately and reliably may affect the ultimate visual quality of the displayed image and consequently may affect a viewer’s engagement with the image and / or with a video comprising the image. This can be especially relevant to XR. Efficient reconstruction is especially effective for mobile computing devices, which may readily be used in XR applications.

[0050] Referring to Figure 3, there is shown schematically an example of an image processing method 300.

[0051] An input image 302 is obtained. The input image 302 may be obtained in various different ways. For example, obtaining may comprise receiving, generating, retrieving, or otherwise.

[0052] At item 304, the input image 302 is pre-processed to generate a pre-processed image 306. An example type of pre-processing is filtering. An example type of filtering is m-filtering using an m-filter. M-filtering can perform blurring and / or sharpening.

[0053] At item 308, the pre-processed image 306 is downscaled to generate a downscaled image 310.

[0054] At item 312, the downscaled image 310 is base-encoded to generate an encoded base 314.

[0055] At item 316, the encoded base 314 is decoded 316 to generate a decoded base 318.

[0056] At item 320, the decoded base 318 is upscaled 320 to generate an upscaled image 322.

[0057] At item 324, residual generation is performed to generate a set of residuals 326. In this example, the residual generation 324 uses the input image 302 and the upscaled image 322. At item 328, the set of residuals 326 is encoded to generate an encoded set of residuals 330.

[0058] At item 332, the encoded set of residuals 330 is decoded to generate a decoded set of residuals 334.

[0059] At item 336, residual addition is performed to generate a reconstructed image 338. In this example, the residual addition 336 uses the upscaled image 322 and the decoded set of residuals 334.

[0060] At item 340, the reconstructed image 338 is post-processed to generate a final output image 342. Such post-processing may comprise dithering and / or s-filtering.

[0061] Referring to Figure 4, there is shown schematically another example of an image processing method 400.

[0062] The example image processing method 400 shown in Figure 4 corresponds closely to the example image processing method 300 shown in Figure 3. Reference signs used in Figure 4 are the same as those used in Figure 3 for the same or similar features, but incremented by 100.

[0063] In the example image processing method 300 shown in Figure 3, the residual generation 324 uses the input image 302 and the upscaled image 322.

[0064] In contrast, in the example image processing method 400 shown in Figure 4, the residual generation 424 uses the pre-processed image 406 and the upscaled image 422.

[0065] Using the pre-processed image 406 for the residual generation 424 may be particularly effective in some scenarios. For example, the input image 402 may be noisy and may have visual artefacts (for example, salt-and-pepper noise), whereas the pre-processed image 406 may be filtered and cleaner. In such an example, residuals 426 based on the pre-processed image 406 and the upscaled image 422 do not include any information about the noise. As a result, the residuals 426 are less complicated and use fewer bits to encode them. Additionally, the reconstructed image 438 does not carry any of the noise artefacts that were present in the input image 402.

[0066] Thus, with reference to both Figures 3 and 4, where pre-processing (such as filtering) is performed at item 304 or 404, the pre-processed (for example, filtered) image 306, 406 is downscaled (which may also be referred to as “downsampled”) to create a filtered base image 310, 410. The filtered base image 310, 410 is upscaled (which may also be referred to as “upsampled”) at items 320 and 420 to create an upscaled image 322, 422. The upscaled image 322, 422 is a prediction of the input image 302, 402. With reference to Figure 3, to calculate the residuals 326, 426, the upscaled prediction 320, 420 may be compared to the original input image 302, 402, i.e., the image 302, 402 before pre-processing (for example, filtering). With reference to Figure 4, to calculate the residuals 326, 426, the upscaled prediction 320, 420 may be compared to the pre-processed (for example, filtered) image 306, 406.

[0067] As explained above, visible artefacts such as halos can arise when images are processed in a manner such as that shown in Figures 3 and 4. Such artefacts may be particularly visible in the downscaled image 310, 410, the decoded base 318, 418, the upscaled image 322, 422, the reconstructed image 338, 438 and / or the final output image 342, 442.

[0068] A halo can arise because of filtering and / or scaling, for example. In filtering, a kernel redistributes colours. A halo is generally, though not exclusively, produced by a large kernel.

[0069] With reference again to Figures 3 and 4, it is believed that halos are not caused by problems with the residuals 330, 430 themselves. For example, as explained above, halos can be visible in the decoded base 318, 418, prior to the residuals 330, 430 being generated and used. It can, however, sometimes be difficult to see halos in the decoded base 318, 418 because the resolution is relatively low, compared to that of the input image 302, 402, and because the amount of quantisation used in base encoding 312, 412 can be high.

[0070] In the context of Low Complexity Enhancement Video Codec (LCEVC), an upsampler may largely or entirely be set as part of the LCEVC standard. LCEVC is described in WO 2020 / 188273, WO 2020 / 188273, WO 2019 / 111010, WO 2019 / 111010, and US 2021 / 0211752, the entire contents of all of which are incorporated herein by reference. For example, it may only be possible to select an upsampler from a predetermined a list of permissible upsamplers. Although a custom upsampler may be selectable from within such a list, there may still be limits on the upsampler in terms of flexibility. However, the downsampler specification may not be part of the LCEVC (or other) standard. There may, therefore, be significantly more flexibility in terms of the downsampler. The filter may also be changeable while maintaining compliance with the LCEVC standard. Examples described herein, which primarily concern pre-processing and downsampling of an image, can reduce visible artefacts while still enabling standards-compliance, for example with LCEVC.

[0071] Different stages of an image and / or video processing pipeline may cause and / or amplify the effect of a halo. For example, pre-processing 304, 404 (for example, filtering such as m-filtering) might cause some halo. Downscaling 308, 408 might amplify a halo resulting from the filtering and / or might introduce further halo. Similarly, upscaling 320, 420 might amplify a halo and / or might introduce further halo. Base encoding 312, 412, residual addition 336, 436 and / or post-processing 340, 440 may also introduce and / or amplify halos. Thus, one or more stages may introduce halo. It might not always be readily apparent which stage(s) introduce(s) halo.

[0072] Standard low-resolution images and video (for example, images and video that are natively low-resolution) do not generally generate similar phenomena. Such phenomena may, however, be present as a result of a more complex image processing pipeline such as that described above with reference to Figures 3 and 4.

[0073] Halos may be particularly noticeable in XR, such as VR and AR. This is because XR can magnify and amplify even very small artefacts. Correcting this type of artefact via residuals may be prohibitively costly in terms of enhancement data rate. In particular, ringing artefacts, such as halos, are generally made of pixel value differences of ±1, ±2 or ±3. These are very small but are still visible. It would therefore take near lossless stepwidth to nullify such artefacts via residuals because the pixel value differences are so small.

[0074] Without loss of generality, examples describe herein provide measures to reduce such undesirable artefacts using edge-driven downsampling. Edge-driven downsampling may also be referred to as “edge-based downsampling”. An edge map is extracted from an input image. The input image is processed at least twice using different processing each time. For example, different pre-processing (for example, filtering) and / or downsampling may be used. This produces at least two different processed versions of the input image. The edge map, or a post-processed (for example, downscaled) version of the edge map, is used to select which processed version(s) of the image a pixel of an output image is to be based on. This may be performed on a pixel-by-pixel or block-by-block basis for the entire output image. One processed version may be used where the output image pixel corresponds to an edge and another processed version may be used where the output image pixel corresponds to a non-edge. A combination of both of the multiple processed versions may be used in some examples, for example where the edge map indicates a range of edge strength values rather than solely edge or non-edge. Thus, edge and non-edge pixels can be treated differently. Measures may be taken to reduce halos on edge-related pixels, for example by using a type of downsampler that reduces such halos. Non-edge-related pixels may be processed differently, where halos are not expected. Such measures can reduce halos where they are to be expected, while maintaining visual quality in other image regions and overall.

[0075] Referring to Figure 5, there is shown a schematic diagram of an example of edge-driven image generation 500.

[0076] In some examples, edge-driven image generation is performed by a central processing unit, CPU. In some examples, edge-driven image generation is performed by a GPU. The CPU implementation may enable more sophisticated edge maps than the GPU implementation. However, a GPU is typically faster than a CPU since many image processing algorithms can be parallelised on GPUs. A GPU may be faster in the sense of higher throughput and / or lower latency.

[0077] A first set of image elements 502 is obtained. An example of an image element is a pixel. Numerical identifiers used herein in relation to sets of image elements (such as “first”, “second”, “third”, and so on) do not imply, of themselves, a temporal order. For example, a third set of image elements may be obtained before a second set of image elements is obtained, and so on.

[0078] The first set of image elements 502 may correspond to all or part of an input image. The first set of image elements 502 is labelled in Figure 5 as “input image” accordingly. Part of an image may be referred to as an “image region”.

[0079] A subset of the first set of image elements 502 is depicted in Figure 5 using reference sign 504. In this specific example, the subset 504 comprises four image elements in a square, 2 x 2 array. If the set of image elements 502 corresponds to an image, the subset 504 may correspond to an image region of the image.

[0080] In some examples, the first set of image elements 502 is in a sequence of sets of image elements. The sequence of sets of image elements may, for example, correspond to video. Edge-driven image generation 500 may be performed for each set of image elements in the sequence.

[0081] An edge map 506 is obtained. The edge map 506 comprises a set of edge map elements. The edge map 506 may be obtained in various ways. In some examples, the edge map is obtained by using the OpenCV library to perform edge detection on the first set of image elements 502. Such edge detection identifies boundaries (in other words, edges) of objects represented by the first set of image elements 502.

[0082] An example edge-detection algorithm available in OpenCV is Canny edge detection. To perform Canny edge detection, a source luma plane is obtained. The source luma plane may correspond to the first set of image elements 502 in this example. Gaussian blur is applied to the source luma plane. A Sobel filter is applied both horizontally and vertically as follows, where A is the Gaussian-blurred image:

[0083]

[0084] Intensity gradients are determined as follows:

[0085] |G | — , -G — arctan(Gx / Gy).

[0086]

[0087] Non-maximum suppression is performed. Double-thresholding and edgetracking are also performed to calculate the edge map.

[0088] In this example, the first set of image elements 502 and the edge map 506 have the same spatial resolution as each other. For example, the first set of image elements 502 and the edge map 506 may both have width ‘IV’ and height ‘H’. By way of a specific example, the first set of image elements 502 and the edge map 506 may each have width W = 1920 and height H = 1080.

[0089] In other examples, the first set of image elements 502 and an associated edge map have different spatial resolutions from each other. For example, an edge map may have the same spatial resolution as the subset 504, i.e. 2 x 2. Multiple edge maps may be used to represent the first set of image elements 502. In other examples, as will be described in more detail below, an initial edge map having the same spatial resolution as the first set of image elements 502 may be post-processed, for example downscaled. Such post-processing may result in a post-processed edge map having a lower spatial resolution than that of the first set of image elements 502. A subset of the edge map elements of the edge map 506 is depicted in Figure 5 using reference sign 508. In this specific example, the subset 508 comprises four edge map elements in a 2 x 2 array. In this example, the four edge map elements are the first and second edge map elements of both the first and second rows of the edge map 506.

[0090] An edge map element in the edge map 506 indicates an edge characteristic of one or more corresponding image elements in the first set of image elements 502. Edge map elements and image elements in the first set of image elements 502 may correspond in that they are in the same positions of the edge map 506 and the first set of image elements 502 respectively. In this example, the first edge map element of the edge map 506 indicates an edge characteristic of the first image element in the first set of image elements 502.

[0091] An example of an edge characteristic is an extent to which the corresponding image element represents an edge. The extent to which the corresponding image element represents an edge may be expressed in various different ways. For example, an edge strength score may be used. An edge strength may also be referred to as an “edge weight”. In a specific example, the edge strength score is a score from 0 to 1 inclusive, with a score of 0 indicating a non-edge and a score of 1 indicating an edge. Any number of scores from 0 to 1 inclusive may be used. For example, scores at increments of 0.1 from 0 to 1 inclusive may be used.

[0092] Another example of an edge characteristic is whether or not the corresponding image element represents an edge. This may be represented by a binary edge strength score. For example, an edge strength score of 0 may be used for a non-edge and an edge strength score of 1 may be used for an edge. An edge strength score of 0 may be expressed as “NE” (non-edge) and an edge strength score of 1 may be expressed as “E” (edge).

[0093] Another example of an edge characteristic is an edge type of an edge represented by the corresponding image element. An example of an edge type is an edge direction of an edge represented by the corresponding image element. This may be represented in a manner that distinguishes between horizontal, vertical and / or diagonal edges. For example, “H”, “V” and “D” may be used.

[0094] Thus, the edge map 506 may define edges and / or outline boundary pixels of objects. The edge map 506 may be based on a preceding edge map of a preceding set of image elements in a sequence of image elements. Where the first set of image elements 502 is part of a sequence of sets of image elements, the preceding edge map may correspond to a preceding set of image elements. This can exploit temporal stability and correlation between sets of image elements and, hence, sets of edge maps. Similarly, one or more following edge maps may be based on the edge map 506.

[0095] A second set of image elements 510 is generated. The second set of image elements 510 may correspond to all or part of an output image. The second set of image elements 510 is labelled in Figure 5 as “output image” accordingly.

[0096] A subset of the second set of image elements 510 is depicted in Figure 5 using reference sign 512. In this specific example, the subset 512 comprises one image element in a 1 x 1 array.

[0097] An example method to generate the second set of image elements 510 will now be described with reference only to the first image element 512 of the second set of image elements 510.

[0098] For the image element 512 of the second set of image elements 510, one or more corresponding edge map elements of the edge map 506 are identified.

[0099] In this specific example, the four edge map elements in the subset 508 are identified. Thus, in this example, for a 1 x 1 array of image elements in the second set of image elements 512, a corresponding 2 x 2 array of edge map elements 508 of the edge map 506 is identified.

[0100] Figure 5 shows two further sets of image elements, namely a third set of image elements 514 and a fourth set of image elements 516. In this specific example, the third and fourth sets of image elements 514, 516 are obtained before the second set of image elements 510 is generated.

[0101] The third and fourth sets of image elements 514, 516 may correspond to all or part of first and second processed versions of the first set of image elements 502 respectively. Additionally, as will be described in more detail below, the third and fourth sets of image elements 514, 516 are associated with edges and non-edges respectively. The third and fourth sets of image elements 514, 516 are labelled inFigure 5 as “first processed version (edge)” and “second processed version (non-edge)” accordingly. Subsets of the third and fourth sets of image elements 514, 516 are depicted in Figure 5 using reference signs 518 and 520 respectively. In this specific example, each subset 518, 520 comprises a single image element in a 1 x 1 array.

[0102] In this example, the third and fourth sets of image elements 514, 516 are generated as a result of some or all of the first set of image elements 502 having been processed using first and second processing respectively. The first processing is different from the second processing.

[0103] Thus, the first set of image elements 502 may be processed once (for example using a first filtering kernel and / or a first downscaler) and may be processed again (for example using a second filtering kernel and / or a second downscaler) to produce the first and second processed versions of the first set of image elements 502. Downscaling may therefore be considered to be adaptive in that different downscalers may be used. In particular, different downscalers may be used in relation to a single input image.

[0104] The first processing may comprise filtering using a first kernel. The second processing may comprise filtering using a second, different kernel.

[0105] The first kernel may have a first size, and the second kernel may have a second, different size. The first kernel may have a first shape, and the second kernel may have a second, different shape. The size and / or shape of a kernel may be defined based on the number of coefficients vertically in the kernel and / or the number of coefficients horizontally in the kernel. The first kernel may have a first set of coefficients, and the second kernel may have a second, different set of coefficients. In examples, the first and second kernels are both square and are both symmetric. Different kernels may be used for different purposes. For example, a kernel that reduces halo compared to another kernel may be used for edge pixels, where halos are more likely to occur.

[0106] The first processing may comprise downsampling using a first downsampler. The second processing may comprise downsampling using a second, different downsampler. Examples of downsamplers include, but are not limited to, Lanczos3 and area downsamplers. Area downsamplers may also be known as “box filters”. Area downsamplers average the pixel values in an image box, i.e. in a square.

[0107] For the first image element 512 of the second set of image elements 510, one or more corresponding image elements 518, 520 of the third and / or fourth sets of image elements 514, 516 are selected. Such selecting may be implemented by a selector 522. In this specific example, the image element 518 in the 1 x 1 array of the third set of image elements 514, the image element 520 in the 1 x 1 array of the fourth set of image elements 516, or both may be selected.

[0108] Thus, in some examples, the selecting selects only one corresponding image element from only one of the third and fourth sets of image elements 514, 516.

[0109] In other examples, the selecting selects corresponding image elements from both the third and fourth sets of image elements 514, 516. In such other examples, generating the first image element 512 of the second set of image elements 510 may comprise determining first and second weightings {w1(w2] to be applied to the selected image elements 518, 520 from the third and fourth sets of image elements respectively. The first image element 512 of the second set of image elements 510 may therefore be generated based on the determined first and second weightings {w1(w2). The first and second weightings {w1(w2] may be based on the identified one or more corresponding edge map elements 508, for example the edge characteristic(s) of the one or more corresponding edge map elements 508. For example, the first and second weightings {w1(w2} may be based on the direction of any edges and / or the strength of any edges, rather than merely based on the presence or non-presence of an edge.

[0110] Where corresponding image elements from both the third and fourth sets of image elements 514, 516 are selected, generating the first image element 512 of the second set of image elements 510 may comprise combining a value of the selected image element 518 from the third set of image elements 514 and a value of the selected image element 520 from the fourth set of image elements 516. The combination may depend on the first and second weightings {w1(w2}. Thus, the selected image elements 518, 520 from the third and fourth sets of image elements 514, 516 may each make a weighted contribution to the value of the first image element 512. The weighted contribution may be from 0% to 100% inclusive.

[0111] The image elements 518, 520 in the 1 x 1 arrays of the third and fourth sets of image elements 514, 516 correspond to the first image element 512 of the second set of image elements 510 in that they are also the first image element of the third and fourth sets of image elements 514, 516 respectively.

[0112] The selection of the one or more corresponding image elements 518, 520 of the third and / or fourth sets of image elements 514, 516 is based on the one or more corresponding edge map elements of the edge map 506. In this example, the one or more corresponding edge map elements of the edge map 506 are the four edge map elements of the edge map 506 as described above.

[0113] In this example, the third and fourth sets of image elements 514, 516 are associated with edges and non-edges respectively. Thus, when an edge map element indicates that a corresponding image element represents an edge, an image element from the third set of image elements 514 is selected. Similarly, when an edge map element indicates that a corresponding image element represents a non-edge, an image element from the fourth set of image elements 516 is selected.

[0114] In this example, the second, third, and fourth sets of image elements 510, 514, 516 have lower spatial resolutions than the first set of image elements 502 and, therefore in this example, the edge map 506. Additionally, in this example, the second, third, and fourth sets of image elements 510, 514, 516 all have the same spatial resolution as each other. For example, the second, third, and fourth sets of image elements 510, 514, 516 may all have width ‘V / / 2’ and height ‘H / 2’, where the first set of image elements 502 has width ‘VF’ and height ‘H For example, where W = 1920 and H = 1080, V / / 2 = 960 and H / 2 = 540. As explained above, the edge map 506 may be downsampled. The downsampled edge map may have the same spatial resolution as the second, third, and fourth sets of image elements 510, 514, 516.

[0115] In some examples, the first, second, third and fourth sets of image elements 502, 510, 514, 516 comprise extended reality, XR, content. As explained above, visible artefacts such as halos are particularly noticeable in XR content. The edge-driven image processing 500 is, therefore, particularly effective in relation to XR content.

[0116] In some examples, the second set of image elements 510 is output for display on an XR display device. The XR display device may comprise an XR headset. Visible artefacts such as halos would be particularly noticeable on an XR display device, such an XR headset. Edge-driven image processing 500 may, however, reduce such artefacts. The second set of image elements 510 may be subject to processing prior to such output. For example, the second set of image elements 510 may be encoded prior to being output.

[0117] In some examples, a predicted reconstruction of the first set of image elements 502 is generated using the second set of image elements 510. A set of residuals may be generated using the predicted reconstruction of the first set of image elements 502. The set of residuals may be provided to an LCEVC encoder. This may correspond to the processing described above in connection with Figures 2A, 3 and 4.

[0118] Generating the set of residuals may use the first set of image elements 502, such as described above with reference to Figure 3.

[0119] Generating the set of residuals may use a pre-processed (for example, filtered) version of the first set of image elements 502, such as described above with reference to Figure 4.

[0120] Thus, an input image may be analysed. Such analysis may result in an edge map. Such analysis may be on a per-image (which may also be referred to as a “per-frame”) basis. One or more filtering kernels and / or one or more downsamplers may be selected based on the analysis and edge map. Such selection may be on a per-pixel basis.

[0121] An encoder may have access to the edge map 506. The encoder can therefore make edge-aware decisions using the edge map 506.

[0122] Referring to Figures 6A and 6B, there is shown a schematic diagram of another example of edge-driven image generation 600.

[0123] The example edge-driven image generation 600 shown in Figures 6A and 6B corresponds closely to the example edge-driven image generation 500 shown in Figure 5. Reference signs used in Figures 6A and 6B are the same as those used in Figure 5 for the same or similar features, but incremented by 100.

[0124] In this example, the edge map 606 of the first set of image elements 602 is obtained. Each edge map element in the edge map 606 indicates an edge characteristic of a corresponding image element in the first set of image elements 602. In this specific example, the edge characteristic is whether the corresponding image element in the first set of image elements 602 is an edge (“E”) or is a non-edge (“NE”). Thus, in this specific example, the first edge map element (“E”) indicates that the first image element in the first set of image elements 602 (“1”) is an edge, the second edge map element (“NE”) indicates that the second image element in the first set of image elements 602 (“2”) is a non-edge and so on.

[0125] Generation of the second set of image elements 610 will now be described. For the first image element 612-1 in the second set of image elements 610, one or more corresponding edge map elements of the edge map 606 are identified. In this specific example, the identifying identifies a single corresponding edge map element 608-1 in the edge map 606, namely the first edge map element 608-1 in the edge map 606.

[0126] One or more corresponding image elements 618-1, 620-1 of the third and / or fourth sets of image elements 614, 616 are selected based on the identified one or more corresponding edge map elements 608-1. The selection may be based on one or more selection criteria. In this specific example, the selecting selects only one corresponding image element 618-1, 620-1 from only one of the third and fourth sets of image elements 614, 616, namely the first image element 618-1 from the third set of image elements 614. This is because the selecting is based on the identified one or more corresponding edge map elements which, in this example, is the first edge map element 608-1 (“E”). Since the first edge map element 608-1 (“E”) indicates that the corresponding image element 604-1 of the first set of image elements 602 is an edge, an image element from the third set of image elements 614 is selected. This is because, in this example, the third set of image elements 614 is selected in the case of an edge and the fourth set of image elements 616 is selected in the case of a non-edge. In this specific example, no image elements from the fourth set of image elements 616 are therefore selected. Thus, in this specific example, the selected one or more corresponding image elements of the third and / or fourth sets of image elements 614, 616 comprise only the corresponding image element 618-1 of the third set of image elements 614 and no corresponding image elements 620 of the fourth set of image elements 616.

[0127] The first image element 612-1 in the second set of image elements 610 is generated based on the selected one or more corresponding image elements which, in this specific example, is only the corresponding image element 618-1 from the third set of image elements 614. The first image element 612-1 in the second set of image elements 610 may, for example, take the value of the corresponding image element 618-1 from the third set of image elements 614.

[0128] Thus, and as shown in Figure 6A, the first image element 612-1 in the second set of image elements 610 has a 100% contribution from the corresponding image element 618-1 from the third set of image elements 614 and a 0% contribution from the corresponding image element 620-1 from the fourth set of image elements 616. In this example, the selector 622 performs a binary selection in that the contribution of an image element from either the third or fourth sets of image elements 614, 616 is either 0% or 100%.

[0129] Referring now to Figure 6B, for the second image element 612-2 in the second set of image elements 610, one or more corresponding edge map elements of the edge map 606 are identified. In this specific example, the identifying identifies a single corresponding edge map element 608-2 in the edge map 606, namely the third edge map element 608-2 in the edge map 606. In this example, the third edge map element 608-2 indicates an edge characteristic of the third image element 604-2 in the first set of image elements 602.

[0130] In this example, the third edge map element 608-2 in the edge map 606, rather than the second edge map element in the edge map 606, is identified. This is because, in this specific example, the second, third and fourth sets of image elements 610, 614, 616 have half the width and height of the first set of image elements 602 and the edge map 606. Thus, in this specific example, the second edge map element in the edge map 606 would not be representative of an image element in the first set of image elements 602 that contributes to the second image element 612-2 in the second set of image elements 610.

[0131] One or more corresponding image elements 618, 620 of the third and / or fourth sets of image elements 614, 616 are selected based on the identified one or more corresponding edge map elements 608-2. In this specific example, the selecting selects only one corresponding image element 618, 620 from only one of the third and fourth sets of image elements 614, 616, namely the corresponding image element 620-2 from the fourth set of image elements 616. This is because the selecting is based on the identified one or more corresponding edge map elements 608-2 which, in this example, is the third edge map element 608-2 (“NE”). Since the third edge map element 608-2 (“NE”) indicates that the corresponding image element 604-2 of the first set of image elements 602 is a non-edge, an image element 620-2 from the fourth set of image elements 616 is selected. In this specific example, no image elements from the third set of image elements 614 are selected. Therefore, in this specific example, the selected one or more corresponding image elements 618, 620 of the third and / or fourth sets of image elements 614, 616 comprise only the corresponding image element 620-2 of the fourth set of image elements 616 and no corresponding image elements 618 of the third set of image elements 614.

[0132] The second image element 612-2 in the second set of image elements 610 is generated based on the selected one or more corresponding image elements which, in this specific example, is only the corresponding image element 620-2 from the fourth set of image elements 616. The second image element 612-2 in the second set of image elements 610 may, for example, take the value of the corresponding image element 618-2 from the fourth set of image elements 616.

[0133] Thus, and as shown in Figure 6B, the second image element 612-2 in the second set of image elements 610 has a 0% contribution from the corresponding image element 618-2 from the third set of image elements 614 and a 100% contribution from the corresponding image element 620-2 from the fourth set of image elements 616.

[0134] In this example, the selector 622 also performs a binary selection.

[0135] By way of a summary, the first image element 612-1 in the second set of image elements 610 may be generated by identifying one or more corresponding edge map elements in the edge map 606 and then selecting one or both of the third and fourth sets of image elements 614, 616 for the first image element 612-1.

[0136] In some such examples, the third and fourth sets of image elements 614, 616 have been fully generated before the selection occurs. Thus, all image elements of both the third and fourth sets of image elements 614, 616 may be available for selection when the selection occurs. Some of the image elements of the third and / or fourth sets of image elements 614, 616 may not, however, ultimately be selected and used. This may result in an inefficiently in that some processing performed may not be used.

[0137] In other such examples, only part of the third and fourth sets of image elements 614, 616 have been generated before the selection occurs. For example, only the first image element of both the third and fourth sets of image elements 614, 616 may be available for selection when the selection occurs in relation to the first image element 612-1 of the second set of image elements 610. In particular, the second and subsequent image elements of the third and fourth sets of image elements 614, 616 may not be available for selection at that time but may be available for selection subsequently.

[0138] In further such example, the corresponding image elements of the third and fourth sets of image elements 614, 616 have not been generated before the selection occurs. For example, an edge map element may indicate that only an image element from the third set of image elements 614 will be used for a given image element in the second set of image elements 610. All or part of the third set of image elements 614 including that image element may then be generated. Thus, image elements that are not, ultimately, used may not be generated.

[0139] Thus, instead of preparing two different processed versions of an input image, the input image may be processed on-the-fly. This may be particularly effective for non-live video. This may also enable parallel computing to be leveraged. The input image may be split into blocks. A single processed version, or multiple processed versions, may be generated on-the-fly and on a block-by-block basis, for example based on an edge characteristic of the block. The most effective version(s) may therefore be generated and used for the block.

[0140] Generation of image elements for the second set of image elements 610, based on selection of image elements from the third and / or fourth sets of image elements 614, 616 may then be repeated for each image element of the second set of image elements 610.

[0141] By way of a specific example, an input image having a 1920 x 1080 spatial resolution may represent an object with an edge. The input image may be downscaled once using a first downscaler and again using a second downscaler. Such downscaling may produce two half-resolution images, being 960 x 540. The half-resolution images may then be used (for example, binary-selected or proportion-blended) based on a fullresolution edge map having spatial resolution 1920 x 1080, or based on a half-resolution edge map having resolution 960 x 540. An output image may be created pixel-by-pixel by referring to the edge map and selecting one or both of the downsampled images to use for that pixel.

[0142] Thus, the selection of which processed version of the first set of image elements 602 is used can change across the whole second set of image elements 610 multiple times. For example, a filtering kernel and / or downsampler used can change across the whole second set of image elements 610 multiple times. Filtering and / or downsampling may be edge-dependent and / or may be content-dependent.

[0143] The edge-driven image generation 600 described with reference to Figures 6A and 6B differs from using a single downsampler that has a fixed-sized kernel and where the coefficients of the kernel can change on a per-frame basis at runtime. In particular, the edge-driven image generation 600 can use two or more downsamplers per frame, can use multiple kernels with different sizes per frame, and can blend two or more different versions of an input together per frame.

[0144] Since edge-awareness impacts pre-processing, the computation of downsampling and priority mapping may be integrated.

[0145] Referring to Figures 7A and 7B, there is shown a schematic diagram of another example of edge-driven image generation 700.

[0146] The example edge-driven image generation 700 shown in Figures 7A and 7B corresponds closely to the example edge-driven image generation 600 shown in Figures 6 A and 6B. Reference signs used in Figures 7 A and 7B are the same as those used in Figures 6A and 6B for the same or similar features, but incremented by 100.

[0147] In this example, the edge map 706 of the first set of image elements 702 is obtained. Each edge map element in the edge map 706 indicates an edge characteristic of a corresponding image element in the first set of image elements 702. In this specific example, the edge characteristic is whether the corresponding image element in the first set of image elements 702 is an edge (“E”) or is a non-edge (“NE”).

[0148] Generation of the second set of image elements 710 will now be described. For the first image element 712-1 in the second set of image elements 710, one or more corresponding edge map elements of the edge map 706 are identified. In this example, the identifying identifies a group of corresponding edge map elements 708-1 of the edge map 706. In this specific example, the identifying identifies four corresponding edge map elements 708-1 in the edge map 706, namely the first and second edge map elements in both the first and second rows of the edge map 706.

[0149] One or more corresponding image elements 718-1, 720-1 of the third and / or fourth sets of image elements 714, 716 are selected based on the identified one or more corresponding edge map elements 708-1.

[0150] In this example, generating the first image element 712-1 in the second set of image elements 710 comprises using a group edge characteristic of the group of corresponding edge map elements 708-1. In this specific example, the selecting selects a corresponding image element 718-1 from the third set of image elements 714 and also a corresponding image element 720-1 from the fourth set of image elements 716, namely the first image elements 718-1, 720-1 from the third and fourth sets of image elements 714, 716. This is because the selecting is based on the identified one or more corresponding edge map elements which, in this example, is the four corresponding edge map elements 708-1 in the edge map 706. Since the four edge map elements 708-1 (“E”, “NE”, “E”, “E”) indicate that the corresponding image elements 704-1 of the first set of image elements 702 are three edges and a non-edge, image elements from the third and fourth sets of image elements 714, 716 are selected. Thus, in this specific example, the group edge characteristic is that 75% of the edge map elements indicate edge and 25% of the edge map elements indicate non-edge.

[0151] In this example, the selector 722 performs a proportional blend in that the contribution of an image element from either the third or fourth sets of image elements 714, 716 can be a different percentage in addition to 0% and 100%. A proportional blend may also be referred to as a “weighted blend”. The third and fourth sets of image elements 714, 716 may be blended based on the strength of the edges.

[0152] Therefore, in this specific example, the selected one or more corresponding image elements 718-1, 720-1 of the third and / or fourth sets of image elements 714, 716 comprise the corresponding image element 718-1 of the third set of image elements 714 and the corresponding image element 720-1 of the fourth set of image elements 716.

[0153] The first image element 712-1 in the second set of image elements 710 is generated based on the selected one or more corresponding image elements which, in this specific example, are the corresponding image elements 718-1, 710-1 from the third and fourth sets of image elements 714. 716.

[0154] Thus, and as shown in Figure 7A, the first image element 712-1 in the second set of image elements 710 has a 75% contribution from the corresponding image element 718-1 from the third set of image elements 714 and a 25% contribution from the corresponding image element 720-1 from the fourth set of image elements 716.

[0155] Referring now to Figure 7B, for the second image element 712-2 in the second set of image elements 710, one or more corresponding edge map elements of the edge map 706 are identified. In this specific example, the identifying identifies a group of corresponding edge map elements 708-2 of the edge map 706. In this specific example, the identifying identifies four corresponding edge map elements 708-2 in the edge map 706, namely the third and fourth edge map elements in both the first and second rows of the edge map 706. In this example, the third and fourth edge map elements of the first and second rows 708-2 indicate edge characteristics of the third and fourth image elements in both the first and second rows of the first set of image elements 702 respectively.

[0156] One or more corresponding image elements 718-2, 720-2 of the third and / or fourth sets of image elements 714, 716 are selected based on the identified one or more corresponding edge map elements 708-2. In this specific example, the selecting selects both the corresponding image elements 718-2, 720-2 from both the third and fourth sets of image elements 714, 716. This is because the selecting is based on the identified one or more corresponding edge map elements 708-2 which, in this example, are the third and fourth edge map elements (“NE”, “E”, “E”, “E”) of the first and second rows of the edge map 706.

[0157] The second image element 712-2 in the second set of image elements 710 is generated based on the selected one or more corresponding image elements which, in this specific example, are the corresponding image elements 718-2, 720-2 from the third and fourth sets of image elements 714, 716.

[0158] Thus, and as shown in Figure 7B, the second image element 712-2 in the second set of image elements 710 has a 75% contribution from the corresponding image element 718-2 from the third set of image elements 714 and a 25% contribution from the corresponding image element 720-2 from the fourth set of image elements 716.

[0159] In this example, the selector 722 therefore performs a proportional blend. The example edge-driven image generation 600 described above with reference to Figures 6A and 6B will now be compared to the example edge-driven image generation 700 described above with reference to Figures 7A and 7B.

[0160] Firstly, Figure 6A shows a 100% contribution from the third set of image elements 614 and a 0% contribution from the fourth set of image elements 616. However, Figure 7 A shows a 75% contribution from the third set of image elements 714 and a 25% contribution from the fourth set of image elements 716. The contribution shown in Figure 7A is more representative of the first set of image elements 710 than the contribution shown in in Figure 6A. This is because the contribution shown in Figure 7A recognises that the second image element of the first set of image elements 710 is a non-edge. However, the edge-driven image generation 700 shown in Figure 7A may involve more processing resources (including time) than the edge-driven image generation 600 shown in Figure 6A, for example since there are more image elements to analyse.

[0161] Secondly, Figure 6B shows a 0% contribution from the third set of image elements 614 and a 100% contribution from the fourth set of image elements 616. However, Figure 7B shows a 75% contribution from the third set of image elements 714 and a 25% contribution from the fourth set of image elements 716. The contribution shown in Figure 7A is significantly more representative of the first set of image elements 710 than the contribution shown in in Figure 6A. This is because the contribution shown in Figure 7A recognises that three image elements of the second block of image elements of the first set of image elements 710 are edges, even though the first image element of the second block of image elements of the first set of image elements 710 is an edge. However, again, the edge-driven image generation 700 shown in Figure 7B may involve more processing resources (including time) than the edge-driven image generation 600 shown in Figure 6B, for example since there are more image elements to analyse.

[0162] Different scenarios may benefit from different ones of the edge-driven image generation techniques 600, 700.

[0163] Referring to Figure 8, there is shown a schematic diagram of another example of edge-driven image generation 800.

[0164] The example edge-driven image generation 800 shown in Figure 8 corresponds closely to the example edge-driven image generation 500 shown in Figure 5. Reference signs used in Figure 8 are the same as those used in Figure 5 for the same or similar features, but incremented by 300.

[0165] In this example, however, the edge map 806 is post-processed to generate a postprocessed edge map 824. In this specific example, such post-processing comprises downscaling the edge map 806. The post-processed edge map 814 therefore comprises a downscaled edge map 824. In this specific example, the edge map 806 is 2D-downsampled from a width ‘IV’ and a height ‘H’ to a width ‘VF / 2’ and a height ‘H / 2’ . If at least one edge map element within a given 2 x 2 block of the edge map 806 is an edge, then, after 2D-downsampling, the resulting and corresponding edge map element (1 x 1) is also an edge. In another example, the edge map 806 may be ID-downsampled from a width ‘V / ’ and a height ‘H’ to a width ‘1 / 7 / 2’ and a height ‘H’ . If at least one edge map element within a given 1 x2 block of the edge map 806 is an edge, then, after ID-downsampling, the resulting and corresponding edge map element (1 x 1) is also an edge.

[0166] In this specific example, the first set of image elements 802 and the edge map 806 have a width ‘IV’ and a height ‘H’. In this specific example, the second, third and fourth sets of image elements 810, 814, 816 and the downscaled edge map 824 have a width ‘1 / 7 / 2’ and a height ‘H / 2’ . For example, the first set of image elements 802 and the edge map 806 may have a width "W = 1920’ and a height ‘H = 1080, and the second, third and fourth sets of image elements 810, 814, 816 and the downscaled edge map 824 may have a width ‘77 / 2 = 960’ and a height ‘H / 2 = 540’.

[0167] Edge map post-processing may be performed as follows. The edge map 806 is calculated. The edge map 806 is downscaled to generate the downscaled edge map 824. The downscaled edge map 824 may be dilated, such that the edges are made thicker. This may use a 5 x 5 kernel, for example. An outline version of the downscaled edge map 824, namely ‘ edgemap jis mtlines’ may be calculated as edgemap_ds_outlines = dilated_edgemap_ds — edgemap_ds, where ‘dilated_edgemap_ds’ is the above-mentioned dilated edge map and where^edgemap_ds,is the downscaled edge map 824. Area downsampling may be used for pixels on the outlines of the ‘ e dg emap _ds_out line s’ and Lanczos3 downsampling may be used for pixels on the non-outlines of the ‘edgemap_ds_outlines’ .

[0168] Therefore, an edge map as used herein may be a downscaled version (for example, ‘edgemap_ds’) of another version of the edge map, where the other version of the edge map (for example, the edge map 806) has been calculated using the first set of image elements 802. Alternatively, the edge map may be an outline version (for example, 'edgemap_ds_outlines') of a downscaled version (for example, ‘edgemap_ds’) of another version of the edge map, where the other version of the edge map for example, the edge map 806) has been calculated using the first set of image elements 802.

[0169] Thus, an area downsampler might only be used where halos are expected, for example around an edge. An additional layer of detail may be added in relation to edges, such that the outline of what may be considered to be an edge is processed as an edge and such that the interior of what may be considered to be an edge is processed as a non-edge. In particular, all pixels of a thick, vertical line might be considered to be edge pixels based on an edge map. However, halos would only be expected at the outline of the line and not at the interior.

[0170] Results of various experiments conducted in relation to techniques described herein will now be provided.

[0171] A first such experiment used a 1920 x 1080 source luma plane image of the letter “e”. The source image was downsampled once using a Lanczos downsampler and once using an area downsampler. Both downsampled images were then upsampled using an animation upsampler. The area downsampler generated less halo than the Lanczos downsampler. However, the Lanczos downsampler generated more continuous diagonal lines than the area downsampler.

[0172] A second such experiment used a source luma plane image of the letters “JUV” and involved edge-driven downsampling. The source image was downsampled once using an area 2D downsampler and again using a Lanczos3 2D downsampler. An edge map was also extracted from the source image. In a first part of the second experiment, the edge map was downsampled using a first set of downsampling settings in which all image elements with even coordinates, in the vertical and horizontal dimensions, are retained. In a second part of the second experiment, the edge map was downsampled using a second set of downsampling settings in which each 2x 2 block of a binary edge map is converted to a single value, which is “1” if any of the four edge map elements also has a value of “1”, indicating an edge. In both the first and second parts of the second experiment, for a given pixel of the output image, if the corresponding edge map element of the edge map indicated an edge, then the Lanczos3 downsampled version of the source image was used. Otherwise, the area downsampled version of the source image was used.

[0173] In relation to the second experiment, using only a Lanczos3 2D downsampled version of the source, and then upsampling using modified cubic gave a peak signal-to-noise ratio (PSNR) of 36.241 decibel (dB). Using only an area downsampled version of the source, and then upsampling using modified cubic gave a PSNR of 36.221 dB. Using both the Lanczos32D downsampled version and the area downsampled versions of the source in the first part of the second experiment, and then upsampling using modified cubic gave a PSNR of 36.275 dB. Using both the Lanczos32D downsampled version and the area downsampled version of the source in the second part of the second experiment, and then upsampling using modified cubic gave a PSNR of 36.355 dB.

[0174] A third such experiment used a source luma plane image of the letters “Assetto” and involved edge-driven downsampling. The source image was downsampled once using an area 2D downsampler and again using a Lanczos3 2D downsampler. An SDK edge map was also extracted from the source image. In a first part of the third experiment, the edge map was downsampled using a first set of downsampling settings in which all image elements with even coordinates, in the vertical and horizontal dimensions, are retained. In a second part of the third experiment, the edge map was downsampled using a second set of downsampling settings in which each 2 x 2 block of a binary edge map is converted to a single value, which is “1” if any of the four edge map elements also has a value of “1”, indicating an edge. For a given pixel of the output image, if the corresponding edge map element of the edge map indicated an edge, then the Lanczos3 downsampled version of the source image was used. Otherwise, the area downsampled version of the source image was used.

[0175] In relation to the third experiment, using only a Lanczos3 2D downsampled version of the source, and then upsampling using modified cubic gave a PSNR of 34.465 dB. Using only an area downsampled version of the source, and then upsampling using modified cubic gave a PSNR of 34.457 dB. Using both the Lanczos3 2D downsampled version and the area downsampled version of the source in the first part of the third experiment, and then upsampling using modified cubic gave a PSNR of 34.470 dB. Using both the Lanczos3 2D downsampled version and the area downsampled version of the source in the second part of the third experiment, and then upsampling using modified cubic gave a PSNR of 34.498 dB.

[0176] A fourth such experiment used a source luma plane image of the letters “JUV” and involved outline-driven downsampling. The source image was downsampled once using an area ID downsampler and once using a Lanczos3 ID downsampler. An edge map was calculated and an outline version of a ID-downscaled version of the edge map was calculated. For a given pixel of the output image, if the corresponding element of the outline version of the edge map indicated an outline, then the area downsampled version of the source image was used. Otherwise, the Lanczos3 downsampled version of the source image was used. Since halos appear around edges, and since area downsampling generates less halo than Lanczos downsampling, area downsampling may be applied only at pixels around edges and not at pixels at edges.

[0177] Thus, by way of a summary, in some of these experiments, a luma plane source was area 2D downsampled and Lanczos32D downsampled. A full-resolution edge map was generated and 2D downsampled, for example using downsampling settings in which all image elements with even coordinates, in the vertical and horizontal dimensions, are retained. If a pixel for the output image was on an edge, the Lanczos 2D downsampled image was used. Otherwise, the area 2D downsampled image was used. The output was an edge-map-driven downsampled video.

[0178] Using edge-driven downsampling as described herein provides various benefits and considerations. Experiments indicate that there is less halo compared to using Lanczos3 downsampling alone. The halo is reduced to as close as possible to area downsampling alone. However, some of the output from Lanczos3 downsampling is still retained. Experiments also indicate that there is a higher PSNR than using either area or Lanczos3 downsampling, when using modified cubic upsampling. The edge map could potentially be reused in other parts of the image processing pipeline. There is flexibility on edge and non-edge downsampler pairs. In particular, other downsampler pairs could potentially be used for edge and non-edge. Experiments indicate more or a similar amount of halo compared to using area downsampling alone. Experiments also show a lower PSNR for area and Lanczos3 downsampling when using animation upsampling compared to when using modified cubic upsampling. Using edge-driven downsampling as described herein can, however, introduce additional complexities into the image processing pipeline, for example in terms of edge map calculation and edge map downsampling. Additionally, two downsamplings (for example areas and Lanczos3) were used, instead of only one. Further refinements may be made to the edge map calculation algorithm and / or the edge map downsampling algorithm.

[0179] Referring to Figure 9, there is shown a schematic block diagram of an example of an apparatus 900.

[0180] In an example, the apparatus 900 comprises an encoder. In another example, the apparatus 900 comprises a decoder. In other examples, the apparatus 900 comprises neither an encoder nor a decoder but is configured to communicate with an encoder and / or a decoder.

[0181] Examples of apparatus 900 include, but are not limited to, a mobile computer, a personal computer system, a wireless device, base station, phone device, desktop computer, laptop, notebook, netbook computer, mainframe computer system, handheld computer, workstation, network computer, application server, storage device, a consumer electronics device such as a camera, camcorder, mobile device, video game console, handheld video game device, an XR headset, or in general any type of computing or electronic device.

[0182] In this example, the apparatus 900 comprises one or more processors 901 configured to process information and / or instructions. The one or more processors 901 may comprise a CPU. The one or more processors 901 are coupled with a bus 902. Operations performed by the one or more processors 901 may be carried out by hardware and / or software. The one or more processors 901 may comprise multiple colocated processors or multiple disparately located processors.

[0183] In this example, the apparatus 900 comprises computer-useable volatile memory 903 configured to store information and / or instructions for the one or more processors 901. The computer-useable volatile memory 903 is coupled with the bus 902. The computer-useable volatile memory 903 may comprise random access memory (RAM).

[0184] In this example, the apparatus 900 comprises computer-useable non-volatile memory 904 configured to store information and / or instructions for the one or more processors 901. The computer-useable non-volatile memory 904 is coupled with the bus 902. The computer-useable non-volatile memory 904 may comprise read-only memory (ROM).

[0185] In this example, the apparatus 900 comprises one or more data-storage units 905 configured to store information and / or instructions. The one or more data-storage units 905 are coupled with the bus 902. The one or more data-storage units 905 may for example comprise a magnetic or optical disk and disk drive or a solid-state drive (SSD).

[0186] In this example, the apparatus 900 comprises one or more input / output (VO) devices 906 configured to communicate information to and / or from the one or more processors 901. The one or more VO devices 906 are coupled with the bus 902. The one or more I / O devices 906 may comprise at least one network interface. The at least one network interface may enable the apparatus 900 to communicate via one or more data communications networks. Examples of data communications networks include, but are not limited to, the Internet and a Local Area Network (LAN). The one or more I / O devices 906 may enable a user to provide input to the apparatus 900 via one or more input devices (not shown). The one or more input devices may include for example a remote control, one or more physical buttons etc. The one or more I / O devices 906 may enable information to be provided to a user via one or more output devices (not shown). The one or more output devices may for example include a display screen.

[0187] Various other entities are depicted for the apparatus 900. For example, when present, an operating system 907, image processing module 908, one or more further modules 909, and data 910 are shown as residing in one, or a combination, of the computer-usable volatile memory 903, computer-usable non-volatile memory 904 and the one or more data-storage units 905. The data signal processing module 908 may be implemented by way of computer program code stored in memory locations within the computer-usable non-volatile memory 904, computer-readable storage media within the one or more data-storage units 905 and / or other tangible computer-readable storage media. Examples of tangible computer-readable storage media include, but are not limited to, an optical medium (e.g., CD-ROM, DVD-ROM or Blu-ray), flash memory card, floppy or hard disk or any other medium capable of storing computer-readable instructions such as firmware or microcode in at least one ROM or RAM or Programmable ROM (PROM) chips or as an Application Specific Integrated Circuit (ASIC).

[0188] The apparatus 900 may therefore comprise a data signal processing module 908 which can be executed by the one or more processors 901. The data signal processing module 908 can be configured to include instructions to implement at least some of the operations described herein. During operation, the one or more processors 901 launch, run, execute, interpret or otherwise perform the instructions in the signal processing module 908.

[0189] Although at least some aspects of the examples described herein with reference to the drawings comprise computer processes performed in processing systems or processors, examples described herein also extend to computer programs, for example computer programs on or in a carrier, adapted for putting the examples into practice. The carrier may be any entity or device capable of carrying the program.

[0190] It will be appreciated that the apparatus 900 may comprise more, fewer and / or different components from those depicted in Figure 9.

[0191] The apparatus 900 may be located in a single location or may be distributed in multiple locations. Such locations may be local or remote.

[0192] The techniques described herein may be implemented in software or hardware, or may be implemented using a combination of software and hardware. They may include configuring an apparatus to carry out and / or support any or all of techniques described herein.

[0193] In examples described above, two different versions of all or part of an input image are generated. In other examples, more than two different version of all or part of an input image are generated. Thus, two or more filtering kernels and / or two or more downscalers may be used in accordance with examples.

[0194] In examples described above, a group of edge map elements may represent edge characteristics of the same number of image elements. For example, a group of four edge map elements may represent edge characteristics of a group of four image elements. In some examples, a group of edge map elements having a first number of edge map elements may represent edge characteristics of a second number of image elements, where the second number is larger than the first number. For example, a single edge map element may represent edge characteristics of a group of four image elements. A group of edge map elements may be “collapsed” into fewer edge map elements. For example, a group of four edge map elements being “E”, “NE”, “E”, “E” may be collapsed into a single edge map element being “75% E; 25% NE”. Thus, the resolution of an edge map may be different from the resolution of a set of image elements with which the edge map is associated.

[0195] It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

Claims

CLAIMS1. A method of performing edge-driven image generation, the method comprising:obtaining an edge map of a first set of image elements, wherein an edge map element in the edge map indicates an edge characteristic of one or more corresponding image elements in the first set of image elements; andgenerating a second set of image elements, wherein generating the second set of image elements comprises, for an image element in the second set of image elements:identifying one or more corresponding edge map elements of the edge map;selecting, based on the identified one or more corresponding edge map elements, one or more corresponding image elements of a third and / or fourth set of image elements, the selected one or more corresponding image elements comprising:a corresponding image element of the third set of image elements, the third set of image elements having been generated as a result of some or all of the first set of image elements having been processed using first processing; and / ora corresponding image element of the fourth set of image elements, the fourth set of image elements having been generated as a result of some or all of the first set of image elements having been processed using second, different processing; andgenerating the image element in the second set of image elements based on the selected one or more corresponding image elements.

2. A method according to claim 1, wherein the first processing comprises filtering using a first kernel, and wherein the second processing comprises filtering using a second, different kernel.

3. A method according to claim 2, wherein:the first kernel has a first size and the second kernel has a second, different size;the first kernel has a first shape and the second kernel has a second, different shape; and / orthe first kernel has a first set of coefficients and the second kernel has a second, different set of coefficients.

4. A method according to any of claims 1 to 3, wherein the first processing comprises downsampling using a first downsampler, and wherein the second processing comprises downsampling using a second, different downsampler.

5. A method according to claim 4, wherein one of the first and second downsamplers comprises an area downsampler.

6. A method according to claim 4 or 5, wherein one of the first and second downsamplers comprises a Lanczos downsampler7. A method according to any of claims 1 to 6, wherein:the first set of image elements and the edge map have the same spatial resolution as each other;the first set of image elements and the edge map have different spatial resolutions from each other;the second, third, and fourth sets of image elements have lower spatial resolutions than the first set of image elements;the second, third, and fourth sets of image elements have the same spatial resolution as each other; and / orthe second, third, and fourth sets of image elements have the same spatial resolution as the edge map.

8. A method according to any of claims 1 to 7, wherein the edge characteristic indicates:an extent to which the corresponding image element represents an edge; whether or not the corresponding image element represents an edge; and / or an edge type of an edge represented by the corresponding image element.

9. A method according to claim 8, wherein the edge type represents an edge direction of an edge represented by the corresponding image element.

10. A method according to any of claims 1 to 9, wherein the identifying of one or more corresponding image elements of the edge map identifies a single corresponding edge map element of the edge map.

11. A method according to any of claims 1 to 9, wherein the identifying of one or more corresponding image elements of the edge map identifies a group of corresponding edge map elements of the edge map.

12. A method according to claim 11, wherein generating the image element of the second set of image elements comprises using a group edge characteristic of the group of corresponding edge map elements.

13. A method according to any of claims 1 to 12, wherein the selecting selects only one corresponding image element from only one of the third and fourth sets of image elements.

14. A method according to any of claims 1 to 12, wherein the selecting selects: a corresponding image element from the third set of image elements; and a corresponding image element from the fourth set of image elements.

15. A method according to claim 14, wherein generating the image element of the second set of image elements comprises determining a first weighting to be applied to the selected image element from the third set of image elements and determining a second weighting to be applied to the selected image element from the fourth set of image elements, and wherein generating the image element of the second set of image elements is based on the determined first and second weightings.

16. A method according to claim 15, wherein the first and second weightings are based on the identified one or more corresponding edge map elements.

17. A method according to any of claims 14 to 16, wherein generating the image element of the second set of image elements comprises combining a value of the selected image element from the third set of image elements and a value of the selected image element from the fourth set of image elements.

18. A method according to any of claims 1 to 17, comprising:generating a predicted reconstruction of the first set of image elements using the second set of image elements;generating a set of residuals using the predicted reconstruction of the first set of image elements; andproviding the set of residuals to a Low Complexity Enhancement Video Coding, LCEVC, encoder.

19. A method according to claim 18, wherein generating the set of residuals uses:the first set of image elements; ora filtered version of the first set of image elements.

20. A method according to any of claims 1 to 19, wherein the first set of image elements is a set of image elements in a sequence of sets of image elements and wherein the method is performed for each set of image elements in the sequence.

21. A method according to claim 20, wherein the edge map is based on a preceding edge map of a preceding set of image elements in the sequence.

22. A method according to any of claims 1 to 21, comprising:obtaining the first set of image elements.

23. A method according to any of claims 1 to 22, wherein the third set of image elements has been generated as a result of a subset of the first set of image elementshaving been processed using the first processing, and / or wherein the fourth set of image elements has been generated as a result of a subset of the first set of image elements having been processed using the second processing.

23. A method according to any of claims 1 to 22, wherein the third set of image elements has been generated as a result of the first set of image elements having been processed using the first processing, and / or wherein the fourth set of image elements has been generated as a result of the first set of image elements having been processed using the second processing.

24. A method according to any of claims 1 to 23, wherein the edge map is a downscaled version of another version of the edge map, the other version of the edge map having been calculated using the first set of image elements.

25. A method according to any of claims 1 to 23, wherein the edge map is an outline version of a downscaled version of another version of the edge map, the other version of the edge map having been calculated using the first set of image elements.

26. A method according to any of claims 1 to 25, wherein the first, second, third and fourth sets of image elements comprise extended reality, XR, content.

27. A method according to any of claims 1 to 26, comprising:outputting the second set of image elements for display on an XR display device.

28. A method according to claim 27, wherein the XR display device comprises an XR headset.

29. A method according to any of claims 1 to 28, wherein the method is performed by a central processing unit, CPU.

30. A method according to any of claims 1 to 28, wherein the method is performed by a graphics processing unit, GPU.

31. Apparatus configured to perform a method according to any of claims 1 to 30.

32. A computer program configured to perform a method according to any of claims 1 to 30.