Enhancement layer processing of extended reality (XR) content

Enhancement layer processing for XR content addresses the challenges of limited bandwidth and latency by selectively enhancing visual quality within the viewer's FOV, resulting in improved XR experiences.

WO2025125816A1PCT designated stage expired Publication Date: 2025-06-19V NOVA INT LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2024/053106
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-12
Filing Date
2024-12-12
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current technologies face challenges in achieving high-quality, low-latency video transmission for extended reality (XR) applications, particularly due to limited wireless bandwidth and end-to-end latency issues.

Method used

The proposed solution involves enhancement layer processing, where a base layer encoding and an enhancement layer encoding are used to reconstruct XR content at higher quality levels. This approach selectively enhances visual quality in the field of view (FOV) of the viewer, reducing the need for high bandwidth and latency.

Benefits of technology

This method effectively enhances visual quality within the viewer's FOV while optimizing bandwidth usage and reducing latency, thereby improving the overall XR experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024053106_19062025_PF_FP_ABST
    Figure GB2024053106_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A base layer encoding of an image (500) is obtained. The image (500) comprises XR content and is for an XR display device. First enhancement layer processing is performed for a first image region (506) of the image (500). Performing first enhancement layer processing comprises obtaining an enhancement layer (700) encoding of the first image region (506). The enhancement layer (700) encoding of the first image region (506) is in accordance with a first encoding priority. Second enhancement layer processing for a second image region (508) of the image (500) is performed. Performing second enhancement layer processing may comprise inhibiting an enhancement layer (700) encoding of the second image region (508) from being obtained. Alternatively, performing second enhancement layer processing comprises obtaining an enhancement layer (700) encoding of the second image region (508). The enhancement layer (700) encoding of the second image region (508) is in accordance with a second, lower encoding priority.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ENHANCEMENT LAYER PROCESSING OF EXTENDED REALITY (XR)

[0002] CONTENT

[0003] Technical Field

[0004] The present disclosure relates to enhancement layer processing of extended reality, XR, content.

[0005] Background

[0006] An image may be encoded to provide a base layer encoding and an enhancement layer encoding. A signal (e.g., by way of non-limiting examples, an image, a video, a stereoscopic video, a depth map, a point cloud, etc.) may be encoded to provide a base layer encoding, corresponding to a rendition of the signal at a second level of quality, and an enhancement layer encoding, providing data to process along with the base layer encoding so as to reconstruct the signal at a first (higher) level of quality. The base layer encoding may be decoded and displayed without being enhanced by the enhancement layer. Alternatively, the enhancement layer encoding may be decoded and used to enhance the visual quality of the decoded base layer.

[0007] Enhanced visual quality is particularly, but not exclusively, relevant in the context of XR, which includes augmented reality (AR) and virtual reality (VR). In particular, XR content is often viewed on a screen (often a high-resolution screen) that is very close to the viewer. Low visual quality can significantly detract from the experience of the viewer.

[0008] However, in addition to visual quality, other factors can also have a significant impact on performance and viewer experience in the context of XR.

[0009] For example, wireless communication technology, which may be used for XR pixel streaming, has a limited bandwidth. For example, the bandwidth may be 30 megabits per second (Mbps). For example, wireless communication technology, which may be used for XR pixel streaming (e.g., wherein the XR display receives a video produced in real-time by a remote device based on the viewer motion and actions), has a limited bandwidth. In addition, even when more bandwidth is available, end-to-end latency and dropped packets, which both detract from the quality of experience, increase disproportionately with bandwidth. For example, in real-world Wi-Fi and / or 5G scenarios, the bandwidth limits to guarantee a reliable low-latency transmission may be as low as 30-40 megabits per second (Mbps). The display resolution of XR headsets is, however, growing much more rapidly than wireless communication technology bandwidths. These wireless communication technology limitations mean that low bitrates are still to be used during XR pixel streaming, even with growing display resolutions. State-of-the-art video encoding approaches cannot guarantee high quality ultra-low-latency 8K video transmission within the real-world constraints of wireless bandwidth.

[0010] Summary

[0011] Various aspects of the present disclosure are set out in the appended claims.

[0012] Further features and advantages will become apparent from the following description of preferred embodiments, given by way of example only, which is made with reference to the accompanying drawings.

[0013] Brief Description of the Drawings

[0014] Figure 1 shows a schematic block diagram of an example of an image processing system;

[0015] Figures 2A and 2B show a schematic block diagram of another example of an image processing system;

[0016] Figure 3 shows a schematic diagram of an example of an image;

[0017] Figure 4 shows a schematic diagram of an example of an enhancement layer;

[0018] Figure 5 shows a schematic diagram of another example of an image;

[0019] Figure 6 shows a schematic diagram of another example of an enhancement layer;

[0020] Figure 7 shows a schematic diagram of another example of an enhancement layer;

[0021] Figure 8 shows a schematic diagram of another example of an enhancement layer;

[0022] Figure 9 shows a schematic diagram of another example of an enhancement layer;

[0023] Figure 10 shows a schematic diagram of another example of an image; Figure 11 shows a schematic diagram of an example of a sequence of images; and

[0024] Figure 12 shows a schematic block diagram of an example of an apparatus.

[0025] Detailed Description

[0026] Referring to Figure 1, there is shown an example of a signal processing system 100. The signal processing system 100 is used to process signals. Examples of types of signal include, but are not limited to, video signals, multi-view video signals, image signals, audio signals, volumetric signals such as those used in medical, scientific or holographic imaging, or other multidimensional signals.

[0027] The signal processing system 100 includes a first apparatus 102 and a second apparatus 104. The first apparatus 102 and second apparatus 104 may have a clientserver relationship, with the first apparatus 102 performing the functions of a server device and the second apparatus 104 performing the functions of a client device. The signal processing system 100 may include at least one additional apparatus (not shown). The first apparatus 102 and / or second apparatus 104 may comprise one or more components. The one or more components may be implemented in hardware and / or software. The one or more components may be co-located or may be located remotely from each other in the signal processing system 100. Examples of types of apparatus include, but are not limited to, computerised devices, handheld or laptop computers, tablets, mobile devices, games consoles, smart televisions, set-top boxes, XR headsets (including AR and / or VR headsets) etc.

[0028] The first apparatus 102 is communicatively coupled to the second apparatus 104 via a data communications network 106. Examples of the data communications network 106 include, but are not limited to, the Internet, a Local Area Network (LAN) and a Wide Area Network (WAN). The first and / or second apparatus 102, 104 may have a wired and / or wireless connection to the data communications network 106.

[0029] In this example, the first apparatus 102 comprises an encoder 108. The encoder 108 is configured to encode data comprised in and / or derived based on the signal, which is referred to hereinafter as “signal data”. For example, where the signal is a video signal, the encoder 108 is configured to encode video data. Video data comprises a sequence of multiple images or frames. The encoder 108 may perform one or more further functions in addition to encoding signal data. The encoder 108 may be embodied in various different ways. For example, the encoder 108 may be embodied in hardware and / or software. The encoder 108 may encode metadata associated with the signal. The first apparatus 102 may use one or more than one encoder 108.

[0030] Although in this example the first apparatus 102 comprises the encoder 108, in other examples the first apparatus 102 is separate from the encoder 108. In such examples, the first apparatus 102 is communicatively coupled to the encoder 108. The first apparatus 102 may be embodied as one or more software functions and / or hardware modules.

[0031] In this example, the second apparatus 104 comprises a decoder 110. The decoder 110 is configured to decode signal data. The decoder 110 may perform one or more further functions in addition to decoding signal data. The decoder 110 may be embodied in various different ways. For example, the decoder 110 may be embodied in hardware and / or software. The decoder 110 may decoder metadata associated with the signal. The second apparatus 104 may use one or more than one decoder 110.

[0032] Although in this example the second apparatus 104 comprises the decoder 110, in other examples, the second apparatus 104 is separate from the decoder 110. In such examples, the second apparatus 104 is communicatively coupled to the decoder 110. The second apparatus 104 may be embodied as one or more software functions and / or hardware modules.

[0033] The encoder 108 encodes signal data and transmits the encoded signal data to the decoder 110 via the data communications network 106. The decoder 110 decodes the received, encoded signal data and generates decoded signal data. The decoder 110 may output the decoded signal data, or data derived using the decoded signal data. For example, the decoder 110 may output such data for display on one or more display devices associated with the second apparatus 104. The one or more display devices may be components of the second apparatus 104 or may otherwise be associated with the second apparatus 104. The one or more display devices may be operable to display XR content and may, therefore, be referred to as XR display devices.

[0034] In some examples described herein, the encoder 108 transmits to the decoder 110 a representation of a signal at a given level of quality and information the decoder 110 can use to reconstruct a representation of some or all of the signal at one or more higher levels of quality. Such information may be referred to as “reconstruction data”. In some examples, “reconstruction” of a representation involves obtaining a representation that is not an exact replica of an original representation. The extent to which the representation is the same as the original representation may depend on various factors including, but not limited to, quantisation levels. A representation of a signal at a given level of quality may be considered to be a rendition, version or depiction of data comprised in the signal at the given level of quality. In some examples, the reconstruction data is included in the signal data that is encoded by the encoder 108 and transmitted to the decoder 110. For example, the reconstruction data may be in the form of metadata. In some examples, the reconstruction data is encoded and transmitted separately from the signal data.

[0035] The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may comprise residual data, as described in more detail below. Residual data is an example of reconstruction data. The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may also comprise configuration data relating to processing of the residual data. The configuration data may indicate how the residual data has been processed by the encoder 108 and / or how the residual data is to be processed by the decoder 110. The configuration data may be signalled to the decoder 110, for example in the form of metadata.

[0036] The first and / or second apparatuses 102, 104 may be configured to perform some or all of the techniques described herein. A computer program may be configured to perform some or all of the techniques described herein.

[0037] Referring to Figures 2A and 2B, there is shown schematically an example of a signal processing system 200. The signal processing system 200 includes a first apparatus 202 and a second apparatus 204. In this example, the first apparatus 202 comprises an encoder and the second apparatus 204 comprises a decoder. However, as explained above, in other examples, the encoder is not comprised in the first apparatus 202 and / or the decoder is not comprised in the second apparatus 204. In each of the first apparatus 202 and the second apparatus 204, items are shown on two logical levels. The two levels are separated by a dashed line. Items on the first, highest level relate to data at a first level of quality. Items on the second, lowest level relate to data at a second level of quality. The first level of quality is higher than the second level of quality. The first and second levels of quality relate to a tiered hierarchy having multiple levels of quality. In some examples, the tiered hierarchy comprises more than two levels of quality. In such examples, the first apparatus 202 and the second apparatus 204 may include more than two different levels. There may be one or more other levels above and / or below those depicted in Figures 2A and 2B. As described herein, in certain cases, the levels of quality may correspond to different spatial resolutions.

[0038] Referring first to Figure 2A, the first apparatus 202 obtains a first representation of an image at the first level of quality 206. A representation of a given image is a representation of data comprised in the image. The image may be a given frame of a video. The first representation of the image at the first level of quality 206 will be referred to as “input data” hereinafter as, in this example, it is data provided as an input to the encoder in the first apparatus 202. The first apparatus 202 may receive the input data 206. For example, the first apparatus 202 may receive the input data 206 from at least one other apparatus. The first apparatus 202 may be configured to receive successive portions of input data 206, e.g. successive frames of a video, and to perform the operations described herein to each successive frame. For example, a video may comprise frames Fi, F2, ... FT and the first apparatus 202 may process each of these in turn.

[0039] The first apparatus 202 derives data 212 based on the input data 206. In this example, the data 212 based on the input data 206 is a representation 212 of the image at the second, lower level of quality. In this example, the data 212 is derived by performing a downsampling operation on the input data 206 and will therefore be referred to as “downsampled data” hereinafter. In other examples, the data 212 is derived by performing an operation other than a downsampling operation on the input data 206, or the data 212 is the same as the input data 206 (i.e. the input data 206 is not processed, e.g. downsampled).

[0040] In this example, the downsampled data 212 is processed to generate processed data 213 at the second level of quality. In other examples, the downsampled data 212 is not processed at the second level of quality. As such, the first apparatus 202 may generate data at the second level of quality, where the data at the second level of quality comprises the downsampled data 212 or the processed data 213. In some examples, generating the processed data 213 involves the downsampled data 212 being encoded. Such encoding may occur within the first apparatus 202, or the first apparatus 202 may output the processed data 213 to an external encoder. Encoding the downsampled data 212 produces an encoded image at the second level of quality. The first apparatus 202 may output the encoded image, for example for transmission to the second apparatus 204. A series of encoded images, e.g. forming an encoded video, as output for transmission to the second apparatus 204 may be referred to as a “base” stream or “base” layer. As explained above, instead of being produced in the first apparatus 202, the encoded image may be produced by an encoder that is separate from the first apparatus 202. The encoded image may be part of an HEVC encoded video, H264 (AVC) video, or otherwise. Generating the processed data 213 may, for example, comprise generating successive frames of video as output by a separate encoder such as an HEVC video encoder or a H264 (AVC) video encoder. An intermediate set of data for the generation of the processed data 213 may comprise the output of such an encoder, as opposed to any intermediate data generated by the separate encoder.

[0041] Generating the processed data 213 at the second level of quality may further involve decoding the encoded image at the second level of quality. The decoding operation may be performed to emulate a decoding operation at the second apparatus 204, as will become apparent below. Decoding the encoded image produces a decoded image at the second level of quality. In some examples, the first apparatus 202 decodes the encoded image at the second level of quality to produce the decoded image at the second level of quality. In other examples, the first apparatus 202 receives the decoded image at the second level of quality, for example from an encoder and / or decoder that is separate from the first apparatus 202. The encoded image may be decoded using an HEVC decoder (or for example a H264, AVC, decoder). The decoding by a separate decoder may comprise inputting encoded video, such as an encoded data stream configured for transmission to a remote decoder, into a separate black-box decoder implemented together with the first apparatus 202 to generate successive decoded frames of video. Processed data 213 may thus comprise a frame of video data that is generated via a complex non-linear encoding and decoding process, where the encoding and decoding process may involve modelling spatio-temporal correlations as per a particular encoding standard such as HEVC (or H264, AVC). However, because the output of any encoder is fed into a corresponding decoder, this complexity is effectively hidden from the first apparatus 202.

[0042] In an example, generating the processed data 213 at the second level of quality further involves obtaining correction data based on a comparison between the downsampled data 212 and the decoded image obtained by the first apparatus 202, for example based on the difference between the downsampled data 212 and the decoded image. The correction data can be used to correct for encoder-decoder errors (which may also be referred to as “encode-decode errors”), namely errors introduced in encoding and decoding the downsampled data 212. In some examples, the first apparatus 202 outputs the correction data, for example for transmission to the second apparatus 204, as well as the encoded signal. This allows the recipient to correct for the encoder-decoder errors introduced in encoding and decoding the downsampled data 212. This correction data may also be referred to as a “first enhancement” stream. As the correction data may be based on the difference between the downsampled data 212 and the decoded image it may be seen as a form of residual data (e.g. that is different from the other set of residual data described later below). An item of residual data may be referred to as “a residual”.

[0043] In some examples, generating the processed data 213 at the second level of quality further involves correcting the decoded image using the correction data. For example, the correction data as output for transmission may be placed into a form suitable for combination with the decoded image, and then added to the decoded image. This may be performed on a frame-by-frame basis. In other examples, rather than correcting the decoded image using the correction data, the first apparatus 202 uses the downsampled data 212. For example, in certain cases, just the encoded then decoded data may be used and in other cases, encoding and decoding may be replaced by other processing.

[0044] In some examples, generating the processed data 213 involves performing one or more operations other than the encoding, decoding, obtaining, and correcting acts described above.

[0045] The first apparatus 202 obtains data 214 based on the data at the second level of quality. As indicated above, the data at the second level of quality may comprise the processed data 213, or the downsampled data 212 where the downsampled data 212 is not processed at the lower level. As described above, in certain cases, the processed data 213 may comprise a reconstructed video stream (e.g. from an encoding-decoding operation) that is corrected using correction data. In the example of Figures 2A and 2B, the data 214 is a second representation of the image at the first level of quality, the first representation of the image at the first level of quality being the input data 206. The second representation at the first level of quality may be considered to be a preliminary or predicted representation of the image at the first level of quality. In this example, the first apparatus 202 derives the data 214 by performing an upsampling operation on the data at the second level of quality. The data 214 will be referred to hereinafter as “upsampled data”. However, in other examples one or more other operations could be used to derive the data 214, for example where data 212 is not derived by downsampling the input data 206.

[0046] The input data 206 and the upsampled data 214 are used to obtain residual data 216. The residual data 216 is associated with the image. The residual data 216 may be in the form of a set of residual elements, which may be referred to as a “residual frame” or a “residual image”. A residual element may be referred to as “a residual”. A residual element in the set of residual elements 216 may be associated with a respective image element in the input data 206. An example of an image element is a pixel.

[0047] In this example, a given residual element is obtained by subtracting a value of an image element in the upsampled data 214 from a value of a corresponding image element in the input data 206. As such, the residual data 216 is useable in combination with the upsampled data 214 to reconstruct the input data 206. The residual data 216 may also be referred to as “reconstruction data” or “enhancement data”. In one case, the residual data 216 may form part of a “second enhancement” stream. The residual data 216 may therefore result from upsampler-downsampler asymmetry. Upsampler- downsampler asymmetry may also be referred to as “upsamping-downsamping asymmetry”, “upsample-downsample” asymmetry or the like. In some examples, enhancement data (e.g., residual data 216) may include information to guide postprocessing operations to be performed on the reconstructed image to obtain the rendition at the first level of quality.

[0048] The first apparatus 202 obtains configuration data relating to processing of the residual data 216. The configuration data indicates how the residual data 216 has been processed and / or generated by the first apparatus 202 and / or how the residual data 216 is to be processed by the second apparatus 204. The configuration data may comprise a set of configuration parameters. The configuration data may be useable to control how the second apparatus 204 processes data and / or reconstructs the input data 206 using the residual data 216. The configuration data may relate to one or more characteristics of the residual data 216. The configuration data may relate to one or more characteristics of the input data 206. Different configuration data may result in different processing being performed on and / or using the residual data 216. The configuration data is therefore useable to reconstruct the input data 206 using the residual data 216. As described below, in certain cases, configuration data may also relate to the correction data described herein.

[0049] In this example, the first apparatus 202 transmits to the second apparatus 204 data based on the downsampled data 212, data based on the residual data 216, and the configuration data (or data based on the configuration data), to enable the second apparatus 204 to reconstruct the input data 206.

[0050] Turning now to Figure 2B, the second apparatus 204 receives data 220 based on (e.g. derived from) the downsampled data 212. The second apparatus 204 also receives data based on the residual data 216. For example, the second apparatus 204 may receive a “base” stream (data 220), a “first enhancement stream” (any correction data) and a “second enhancement stream” (residual data 216). The base stream may be referred to as a “base layer”. The first and / or second enhancement stream may be referred to as, and / or may be comprised in, an “enhancement layer”. The second apparatus 204 also receives the configuration data relating to processing of the residual data 216. The data 220 based on the downsampled data 212 may be the downsampled data 212 itself, the processed data 213, or data derived from the downsampled data 212 or the processed data 213. The data based on the residual data 216 may be the residual data 216 itself, or data derived from the residual data 216.

[0051] In some examples, the received data 220 comprises the processed data 213, which may comprise the encoded image at the second level of quality and / or the correction data. In some examples, for example where the first apparatus 202 has processed the downsampled data 212 to generate the processed data 213, the second apparatus 204 processes the received data 220 to generate processed data 222. Such processing by the second apparatus 204 may comprise decoding an encoded image (e.g. that forms part of a “base” encoded video stream) to produce a decoded image at the second level of quality. In some examples, the processing by the second apparatus 204 comprises correcting the decoded image using obtained correction data. Hence, the processed data 222 may comprise a frame of corrected data at the second level of quality. In some examples, the encoded image at the second level of quality is decoded by a decoder that is separate from the second apparatus 204. The encoded image at the second level of quality may be decoded using an HEVC decoder (or for example a H264, AVC, decoder).

[0052] In other examples, the received data 220 comprises the downsampled data 212 and does not comprise the processed data 213. In some such examples, the second apparatus 204 does not process the received data 220 to generate processed data 222.

[0053] The second apparatus 204 uses data at the second level of quality to derive the upsampled data 214. As indicated above, the data at the second level of quality may comprise the processed data 222, or the received data 220 where the second apparatus 204 does not process the received data 220 at the second level of quality. The upsampled data 214 is a preliminary representation of the image at the first level of quality. The upsampled data 214 may be derived by performing an upsampling operation on the data at the second level of quality.

[0054] The second apparatus 204 obtains the residual data 216. The residual data 216 is useable with the upsampled data 214 to reconstruct the input data 206. The residual data 216 is indicative of a comparison between the input data 206 and the upsampled data 214.

[0055] The second apparatus 204 also obtains the configuration data related to processing of the residual data 216. The configuration data is useable by the second apparatus 204 to reconstruct the input data 206. For example, the configuration data may indicate a characteristic or property relating to the residual data 216 that affects how the residual data 216 is to be used and / or processed, or whether the residual data 216 is to be used at all. In some examples, the configuration data comprises the residual data 216.

[0056] There are several considerations relating to such processing. One such consideration is the amount of information that is generated, stored, transmitted and / or processed. The more information that is used, the greater the amount of resources that may be involved in handling such information. Examples of such resources include transmission resources, storage resources and processing resources. Some signal processing techniques allow a relatively small amount of information to be used. This may reduce the amount of data transmitted via the data communications network 106. The savings may be particularly relevant where the data relates to high quality video data, where the amount of information transmitted can be especially high.

[0057] Another consideration is latency. Complex image processing may introduce latency, which may negatively impact performance.

[0058] Another consideration is wireless communication technology limitations.

[0059] Other considerations include the ability of the decoder to perform image reconstruction accurately, reliably, and / or efficiently. Performing image reconstruction accurately and reliably may affect the ultimate visual quality of the displayed image and consequently may affect a viewer’s engagement with the image and / or with a video comprising the image. This can be especially relevant to XR. Efficient reconstruction is especially effective for mobile computing devices, which may readily be used in XR applications.

[0060] Referring to Figure 3, there is shown a schematic diagram of an example of an image 300. The image 300 may be a standalone image or may be a frame in video content. The image 300 may, for example, represent a frame of a movie, a television programme, a sporting event, and so on.

[0061] Base layer processing may be performed in relation to the image 300. The base layer processing results in a base layer encoding of the image 300. The base layer encoding may correspond to the processed data 213 described above with reference to Figure 2A, for example.

[0062] The base layer encoding may be obtained in various different ways. In some examples, the base layer encoding is obtained by generating the base layer encoding. In some examples, the base layer encoding is obtained by instructing generation of the base layer encoding. In some examples, the base layer encoding is obtained by configuring an encoder such that the encoder generates the base layer encoding. In some examples, the base layer encoding is obtained by receiving the base layer encoding. In some examples, the base layer encoding is obtained by retrieving the base layer encoding from storage. The base layer encoding may be obtained in another manner in other examples.

[0063] Enhancement layer processing may be performed in relation to the image 300. The enhancement layer processing results in an enhancement layer encoding of the image 300. The enhancement layer encoding may correspond to the correction data and / or the residual data 216 described above with reference to Figure 2A, for example.

[0064] The enhancement layer encoding may be obtained in various different ways. In some examples, the enhancement layer encoding is obtained by generating the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by instructing generation of the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by configuring an encoder such that the encoder generates the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by receiving the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by retrieving the enhancement layer encoding from storage. The enhancement layer encoding may be obtained in another manner in other examples.

[0065] More generally, references herein to obtaining an encoding encompass generating the encoding, instructing generation of the encoding, configuring an encoder to enable the encoder to generate the encoding, receiving the encoding, retrieving the encoding from storage, and obtaining the encoding in any other manner.

[0066] A base layer encoding may be generated by a base layer encoder. An enhancement layer encoding may be generated by an enhancement layer encoder. The base layer encoder may be different from the enhancement layer encoder. In particular, the base and enhancement layer encoders may use different codecs.

[0067] In some examples, the enhancement layer residuals result from upsampler- downsampler asymmetry. Such residuals may correspond to the residual data 216 described above with reference to Figure 2A.

[0068] In some examples, the enhancement layer residuals result from encoder-decoder errors. Such residuals may correspond to the correction data described above with reference to Figure 2A.

[0069] In some examples, the base and enhancement layer encodings are output. In some examples, the base and enhancement layer encodings are stored. The base and enhancement layer encodings may be decoded, combined and displayed, such as described above with reference to Figure 2B.

[0070] Referring to Figure 4, there is shown a schematic diagram of an example of an enhancement layer 400. In this example, the enhancement layer 400 is generated based on the example image 300 shown in Figure 3.

[0071] In this example, the enhancement layer 400 comprises a set of residuals 402.

[0072] The residuals 402 may take various different forms. Residuals are more generally referred to as “correction data” or “enhancement data”.

[0073] In some examples, the residuals 402 comprise MPEG-5 Part 2 Low Complexity Enhancement Video Coding (LCEVC) residuals, which will be described in more detail below. In some examples, the residuals 402 comprise residuals used in scalable codecs, such as SMPTE ST 2117-1 (often known as “VC-6”). In some examples, the scalable codec is a multi-layer, scalable codec. Thus, encodings and / or decodings described herein may be multi-layer-scalable-codec-derived. In some examples, the residuals 402 comprise residuals used in scalable MPEG codecs such as scalable AVC (SVC), scalable HEVC (SHVC), scalable AVI and scalable VVC.

[0074] There are no restrictions on where the residuals 402 may be positioned in the enhancement layer 400. As explained above, the enhancement layer 400 may be used to enhance the visual quality of a base layer encoding of the example image 300. In particular, the residuals 402 enhance the visual quality at their respective positions. As such, visual quality enhancement can readily be provided across the whole of the image 300.

[0075] In many practical situations, the entire image 300 is visible to a viewer. For example, where the image 300 is a frame from a movie, a sporting event, a television programme, a video game, or the like, the entire image 300 is visible to the viewer. The viewer may be looking at any part of the image 300. In addition, the viewer may focus on different parts of a video at different times. For example, the viewer may focus on an object in the centre of the image 300 during one part of the video, may focus on an object in the top right corner of the image 300 during a later part of the video, and so on. Foveation of human vision makes it impossible to view the entire field of view at the maximum retina resolution at any one time. Different viewers may focus on different parts of any given image 300. For example, some viewers may focus on an object at the bottom left comer of the image 300 and other users may focus on a different object at the top right corner of the image 300. High visual quality across the whole image can therefore improve viewer experience and perception of visual quality. For example, by default, residuals may be applied anywhere and everywhere across the image. By way of further example, by default, to cater to all possible areas of focus, residuals may be applied anywhere and everywhere across the image.

[0076] Referring to Figure 5, there is shown a schematic diagram of another example of an image 500.

[0077] Various items are shown, schematically, as being part of the image 500. However, this is solely to facilitate an understanding of the present disclosure. In practice, the items may not be visible in the image 500 itself and the image 500 may represent other content.

[0078] In this example, the image 500 comprises XR content and is for an XR display device. As indicated above, XR content may comprise AR content and / or VR content. The XR display device may correspond to the second apparatus 104 described above with reference to Figure 1, for example. As will be described above, base and enhancement layer encodings generated using the image 500 may be output for display on the XR display device.

[0079] In this example, the image 500 is a stereoscopic image. In this example, the image comprises a left-side image portion 502 and a right-side image portion 504.

[0080] In this specific example, the image 500 is rectangular, with width, w, and height, h. In this specific example, the left-side and right-side image portions 502, 504 are also rectangular with width, w / 2, and height, h. However, the image 500 may have a different shape in other examples. Additionally, in other examples, the left-side and right-side image portions 502, 504 may have different proportions and / or shapes and / or sizes. In some examples, the image 500 comprises a different number of image portions.

[0081] The example image 500 shown in Figure 5 comprises a first image region 506. In this example, the first image region 502 comprises first and second sub-regions 506- 1, 506-2. In this example, the first sub-region 506-1 is in the left-side image portion 502 and the second sub-region 506-2 is in the right-side image portion 504.

[0082] As will be described in more detail below, in some examples, the first image region 506 is a foveated area in that the first image region 506 corresponds to an area that is, or at least is expected to be, gazed by fovea of eyes of a viewer. An area outside of the foveated area may be a peripheral area, namely an area corresponding to peripheral vision of the viewer, or at least expected peripheral vision of the viewer.

[0083] In this example, the first sub-region 506-1 is centred in the centre of the leftside image portion 502 and the second sub-region 506-2 is centred in the centre of the right-side image portion 504. In particular, neither the first sub-region 506-1 nor the second sub-region 506-2 is centred in the centre of the image 500 itself.

[0084] In this example, the first sub-region 506-1 is rectangular with width, and height, H1;1, where VF1;1< w / 2 and < h. In this example, the second sub-region 506-2 is rectangular with width, IV1 2, and height, H1 2, where VF1;2< w / 2 and H1 2< h. In this example, VF1;1= VF1;2and = H1 2. The first and / or second sub-region 506-1, 506-2 may have different shapes and / or dimensions in other examples. In some examples, the first and / or second sub-region 506-1 is rectangular and has a height less than its width. In other examples, the shape of sub-regions is more rounded, approximating the typical shape of small eye-glasses. In yet other examples, the shape of sub-regions may be non-contiguous, to allow for details to be always present in certain areas of the view-port that are looked at frequently via rapid eye movements (e.g., HUD overlays with key statistics of a game).

[0085] In this example, the first and second sub-regions 506-1, 506-2 are nonoverlapping. The first region 506 may therefore be said to be non-contiguous.

[0086] The example image 500 shown in Figure 5 also comprises a second image region 508. In this example, the second image region 508 surrounds the first image region 506. In this example, the second image region 508 is the complement of the first image region 506 in the image 500.

[0087] In this example, the first image region 506 corresponds to an expected field of view (FOV) of a viewer of the image 500. An expected FOV is the FOV the viewer is expected to have. The FOV described in this example and elsewhere (for example the first image region 506), may correspond to an (e.g. expected) ‘foveated’ view area of a viewer of the image 500, which may also be referred to as a foveated (e.g. expected) FOV. An foveated FOV is the area of the image comprising the region that the viewer can see with maximum detail. An expected foveated FOV is the area of the image comprising the region that the viewer is expected to see with maximum detail. For example, a FOV may include areas at the edge of a viewer’s view / vision (which consequently cannot be seen in high detail because they are at the edge of view). A foveated FOV may not comprise such edge areas (‘edge areas being areas at the edge of a viewer’s view / vision). The foveated FOV may (e.g. only) comprise area(s) of the image that the viewer can see in high detail (because such areas are closer than the edge areas to the viewer’s gaze or expected gaze). The expected FOV may differ from an actual FOV of the viewer. For example, last-minute (e.g. rapid) eye movements may change the actual FOV of the viewer compared to what is expected, for example based on a recent FOV of the viewer, eye-tracking data, or otherwise. The expected FOV may also be referred to as a “probable”, “predicted”, “central” or “focal” FOV, and for that reason may be generally bigger than the actual region that a viewer is able to see with maximum detail.

[0088] In this example, the viewer of the image 500 is a user of the XR device.

[0089] More specifically, in this example, the first sub-region 506-1 corresponds to a left-eye view of the viewer of the image 500 and the second sub-region 506-2 corresponds to a right-eye view of the viewer of the image 500. In this example, some or all of the second image region 508 is outside the FOV of the viewer of the image 500.

[0090] A boundary 510 exists between the left-side image portion 502 and the rightside image portion 504. In this example, the second image region 508 includes the boundary 510.

[0091] The boundary 510 may appear in the image 500 as a dark line. For example, the left-eye view (corresponding to the left-side image portion 502) may transition from a relatively light colour at its left side to a relatively dark colour at its right side. The right-eye view (corresponding to the right-side image portion 504) may similarly transition from a relatively light colour at its left side to a relatively dark colour at its right side. The boundary 510 between the left-side and right-side image portions 502, 504 would therefore appear as a dark vertical line. In some examples (e.g. including but not limited to the example of figure 5) multi-view video encoding may be used for the base layer (e.g., MV-HEVC instead of HEVC), so that left eye view and right eye view are encoded with separate image layers. In some examples, the second image region 508 includes at least part of an edge 512 of the image 500. The edge 512 may be referred to as a “boundary”, “perimeter”, or the like. The edge may similarly include a dark line or region, especially in the case of XR content.

[0092] In some examples, the positions of the first and second image regions 506, 508 are independent of pixel values of pixels of the image 500 and so may be said to be pixel-value-independent. Such examples therefore differ from examples in which pixel analysis is performed to determine the positions of the first and second image regions 506, 508.

[0093] In some examples, the positions of the first and second image regions 506, 508 are independent of a genre of the XR content of the image 500 and so may be said to be genre-independent. Such examples therefore differ from examples in which knowledge of the genre of the XR content influences the positions of the first and second image regions 506, 508.

[0094] Therefore, in some examples, the positions of the first and second image regions 506, 508 are content-independent.

[0095] Instead of the positions of the first and second image regions 506, 508 being independent of pixel values and / or XR content genre, pre-analysis of the image 500 could be performed to determine the positions of the first and / or second image regions 506, 508. However, such pre-analysis introduces latency. For XR content, reduced latency generally improves performance and user experience.

[0096] Referring to Figure 6, there is shown a schematic diagram of another example of an enhancement layer 600. In this example, the enhancement layer 600 is generated based on the example image 500 shown in Figure 5. Reference signs used in Figure 6 are the same as those used in Figure 5 for the same or similar features, but incremented by 100.

[0097] The example enhancement layer 600 shown in Figure 6 corresponds generally to the example enhancement layer 400 shown in Figure 4 in that the residuals in the set of residuals are in corresponding positions in Figure 6.

[0098] However, Figure 6 illustrates that a first subset 614 of the residuals (represented by residual 614) are in a first region 606 of the enhancement layer 600 and that a second subset 616 of the residuals (represented by residual 616) are in a second region 608 of the enhancement layer 600. The first and second regions 606, 608 of the enhancement layer 600 correspond to the first and second image regions 506, 508 of the image 500 in terms of their positions. Residuals in the first region 606 of the enhancement layer 600 may therefore be said to correspond to the first image region 506 of the image 500, and residuals in the second region 608 of the enhancement layer 600 may therefore be said to correspond to the second image region 508 of the image 500.

[0099] As explained above, for many types of image, the whole image may be visible to the viewer. However, an entire image may not be visible, for example for an XR image. Moreover, an entire image may not be visible with maximum detail, for example for an XR image.

[0100] Examples that will now be described can improve rate control, performance, and / or visual quality in XR.

[0101] Such examples selectively enhance XR content. For example, part of the XR content may be enhanced and another part of the XR content may be enhanced to a lesser degree, or even not enhanced at all. Such selective enhancement is particularly, but not exclusively, effective for XR content where not all of the XR content provided to the XR display device is within the FOV of the user of the XR display device. XR content outside the FOV of the user can be enhanced less than XR content within the FOV of the user or may not even be enhanced at all. In some examples, the XR display device can still display XR content outside of the FOV of the user based on a base layer encoding of such content, even if a corresponding enhancement layer encoding of such content is not available. Lower visual quality of XR content within the FOV impacts user experience more significantly than outside the FOV. Visual quality outside the FOV may therefore be sacrificed, at least to some degree, relative to visual quality inside the FOV, without significantly negatively affecting performance and user experience. This can improve rate control as less enhancement layer data can be communicated to the XR display device. It is sometimes preferable to sacrifice the size of the FOV rather than the quality of the enhancement within the FOV, letting the rate controller dynamically adjusting the size of the FOV from frame to frame. In some examples, some or all of the saving made by reducing the enhancement outside the FOV and / or reducing the size of the FOV is used to enhance visual quality within the FOV even further.

[0102] In more detail, in a possible implementation, residuals are applied to an object for some frames but not for others. This can improve rate control but can impact visual quality in relation to the object in question.

[0103] Further, as explained above, a stereoscopic image can have an artificially created sharp border. Significant amounts of residuals may be spent trying to make this border a perfect line. However, a precisely defined border line does not improve visual quality when the stereoscopic image is viewed on an XR display device because the border is outside the FOV of the viewer. This applies correspondingly to dark lines or regions at the edge of an image.

[0104] Examples described below address this by not generating residuals, or by generating lower-priority residuals, outside the FOV of the viewer. The border would fall outside of the FOV of the viewer and so no, or smaller, residuals would be spent on the border.

[0105] In some examples, only residuals for regions of an XR image in the FOV of the viewer are generated, encoded, quantised to non-zero values and / or sent. Such regions may be in the centre of each half of the stereoscopic image. In other examples, the location of such regions may vary from frame to frame based on receiving real-time information on eye movements of the viewer, while the size of such regions may similarly vary from frame to frame based on rate control decisions, informed by data comprising the available bitrate and a metric of the complexity of the image.

[0106] Where residuals (e.g. such residuals) are quantised, the quantisation step width may be based on a size of the region(s) in the FOV of the viewer, a target bitrate, and / or one or more other factors.

[0107] Referring to Figure 7, there is shown a schematic diagram of another example of an enhancement layer 700.

[0108] The example enhancement layer 700 shown in Figure 7 corresponds generally to the example enhancement layer 600 shown in Figure 6. Reference signs used in Figure 7 are the same as those used in Figure 6 for the same or similar features, but incremented by 100. However, Figure 7 illustrates that the first subset 714 of residuals is processed differently from the second subset 716 of residuals. This is depicted in Figure 7 by the different shading of the first and second subsets 714, 716 of residuals.

[0109] In more detail, a base layer encoding of the example image 500 shown in Figure 5 is obtained.

[0110] In some examples, the base layer encoding of the image 500 has a consistent encoding quality across the image.

[0111] In other examples, the base layer encoding of the image 500 comprises a first base layer encoding portion corresponding to the first image region 506, and a second base layer encoding portion corresponding to the second image region 508. The first base layer encoding portion has a first encoding quality, and the second base layer encoding portion has a second, lower encoding quality. A lower encoding quality may result in a lower visual quality when the encoding is decoded.

[0112] First enhancement layer processing may be performed for the first image region 506 of the image 500. Performing first enhancement layer processing comprises obtaining an enhancement layer 700 encoding of the first image region 506. The enhancement layer 700 encoding of the first image region 506 is in accordance with a first encoding priority. In this example, the first subset 714 of residuals, which are in the first region 706 of the enhancement layer 700, are encoded in accordance with the first encoding priority.

[0113] An encoding priority may correspond to a visual importance. For example, a higher encoding priority may correspond to a higher visual importance and a lower encoding priority correspond to a lower visual importance. A higher encoding priority may correspond to higher prioritisation of an encoding relative to another encoding having a lower encoding priority. A higher encoding priority may correspond to a lower quantisation step width of an encoding relative to another encoding having a lower encoding priority. A higher encoding priority may correspond to use of a codec having a higher visual quality for an encoding relative to another encoding having a lower encoding priority.

[0114] Second enhancement layer processing is performed for the second image region 508 of the image 500. The first and second enhancement layer processing may be part of a single enhancement encoding process.

[0115] In some examples, performing second enhancement layer processing comprises inhibiting an enhancement layer 700 encoding of the second image region 508 from being obtained. As such, although the second subset 716 of residuals, which are in the second region 708 of the enhancement layer 700, are depicted in Figure 7, this is solely to facilitate understanding. The second subset 716 of residuals may not have been generated at all.

[0116] Inhibiting the enhancement layer 700 encoding of the second image region 508 from being obtained may be implemented as an encoding condition. The encoding condition may be that residual encoding logic is not run for residuals with specific positions.

[0117] This differs from a priority map implementation, which will be described in more detail below, in which residuals are generated for all positions, but in which residuals in specific positions are de-prioritised and quantised down to zero.

[0118] Reference is made, in this connection to WO 2020 / 188229, the contents of which are incorporated herein by reference. WO 2020 / 188229 describes, for example, an encoder analysing an input video and preparing a set of residual masks for each frame of the video. WO 2020 / 188229 also describes that, instead of the encoder performing such pre-analysis of an input video, a central server can propose a residual weighted mask according to a type of input video. The central server may provide a set of residual weighted masks covering different genres, for example sports, movies, news, etc. The residual masks may thereby prioritise the area of the picture in which detail is required. For example, the residual masks may thereby prioritise the area of the picture in which detail is required, responsive to areas of the field of view that may be characterized as not influenced by head movements (e.g. graphics overlays). However, such pre-analysis and / or genre-dependence may introduce processing latency.

[0119] The techniques described herein also differ from WO 2018 / 015764, the contents of which are incorporated herein by reference. In WO 2018 / 015764, one or more enhancement layers may be provided to a decoder for an entire image and the decoder may select which portions of the enhancement layer(s) are to be used. In some examples, inhibiting the enhancement layer 700 encoding of the second image region 508 from being obtained comprises causing an enhancement layer mask to be applied to the second image region 508.

[0120] In such examples, the enhancement layer mask may correspond, in shape, to the second image region 508. In such examples, the enhancement layer mask is applied to one or more lower-priority image regions such that one or more higher-priority image regions are unmasked and outside the enhancement layer mask. Residuals corresponding to the masked, lower-priority image region(s) and the unmasked, higher- priority image region(s) can be processed accordingly.

[0121] In other examples, inhibiting the enhancement layer 700 encoding of the second image region 508 from being obtained comprises causing an enhancement layer mask to be applied to the first image region 506. In such examples, the enhancement layer mask may correspond, in shape, to the first image region 506. In such examples, the enhancement layer mask is applied to one or more higher-priority image regions such that one or more lower-priority image regions are unmasked and outside the enhancement layer mask. Residuals corresponding to the masked, higher-priority image region(s) and the unmasked, lower-priority image region(s) can be processed accordingly.

[0122] In some examples, the enhancement layer mask has a geometric shape. For example, the enhancement layer mask may be rectangular (such as square), circular, oval, or may have a more complex shape. Less complex shapes may have improved computation performance relative to more complex shapes. However, less complex shapes may resemble the FOV less accurately than more complex shapes.

[0123] The geometry and / or size of the shape may be determined based on one or more factors. Examples of such factors include, but are not limited to, available bandwidth, target bitrate, complexity of the scene, and visual quality target. As such, rate control may be provided by changing the geometry and / or size of the shape. As such, rate control may be provided by (e.g. dynamically) changing the geometry and / or size of the shape along with the quantisation of the enhancement data, effectively trading off- for a given bitrate target - the size of the FOV with the quality level within the FOV

[0124] More specifically, in response to a predetermined change in a condition, a predetermined change may be made to the shape. The condition may be a communications network condition. For example, in response to a predetermined decrease in bandwidth, the size of the shape may be decreased. The predetermined decrease in bandwidth may correspond to the bandwidth dropping below a threshold bandwidth value, or otherwise. Advantageously, this allows the rate controller to rapidly reduce the bitrate allocation to the enhancement data without necessarily producing visible degradation in the central area of the FOV. By way of another example, in response to a predetermined increase in bandwidth, the size of the shape may be increased. The predetermined increase in bandwidth may correspond to the bandwidth rising above a threshold bandwidth value, or otherwise.

[0125] The size, shape and / or position of the enhancement layer mask may be fixed. For example, the size, shape and / or position of the enhancement layer mask may be the same in relation to different images in a sequence of images.

[0126] The size, shape and / or position of the enhancement layer mask may, instead, be dynamic. For example, the size, shape and / or position of the enhancement layer mask may differ in relation to different images in a sequence of images.

[0127] The enhancement layer mask may be considered to be a foveated priority map. foveated priority map provide a map of different priorities (in this example, encoding priorities) in relation to foveated imaging. In foveated imaging, visual quality differs across an image as explained herein. A foveated priority map may also be referred to as a foveated priority map area.

[0128] In a specific example, the enhancement layer mask (in other words, a foveated priority map area) is a rectangular shape with its height less than its width. A foveated priority map area may therefore be shaped in a more ‘horizontal’ or ‘landscape’ orientation than a more ‘vertical’ or ‘portrait’ orientation, and may therefore be akin to typical eyeglasses. The foveated priority map area may therefore be considered to be rectangular and in a landscape orientation. In such examples, the long edges of the rectangle may be aligned with (for example, parallel to) top and bottom edges of an image frame and / or view. In such examples, the short edges of the rectangle may be aligned with (for example, parallel to) left and right edges of an image frame and / or view. This can be particularly effective in the context of XR content. This is because, when a viewer is not moving their head, their eyes typically cover more distance over a left-right (horizontal) axis than an up-down (vertical) axis. This is because viewers typically read a scene like they would read a book, namely moving their eyes left to right, whereas viewers typically move their head up and down instead of straining their pupils to view up or down by a large amount.

[0129] Thus, the enhancement layer mask may be considered to be a foveated priority map that gives priority to one or more areas of the FOV. For example, the foveated priority map may correspond to a rectangle, circle or any other shape in the centre of the field of view for each of the left-eye and right-eye views. In other examples where the XR device can provide eye tracking data, the encoder dynamically adapts the location of the foveated priority map based on the specific locations that the viewer is focusing on from time to time.

[0130] In some examples, performing second enhancement layer processing comprises obtaining an enhancement layer 700 encoding of the second image region 508. In such examples, the enhancement layer 700 encoding of the second image region 508 is in accordance with a second encoding priority. The second encoding priority is lower than the first encoding priority.

[0131] In some examples, obtaining the enhancement layer 700 encoding of the second image region 508 comprises obtaining a set of residuals 716 corresponding to the second image region 508, where the set of residuals 716 corresponding to the second image region 508 are quantised to zero. In this example, the set of residuals 716 corresponding to the second image region 508 is the second subset 716 of residuals. As such, the second subset 716 of residuals may be generated, and their values may be set to zero during quantisation.

[0132] In some examples, obtaining the enhancement layer 700 encoding of the first image region 506 comprises obtaining a set of residuals 714 corresponding to the first image region 506, where the set of residuals 714 corresponding to the first image region 506 are quantised with a first quantisation step width. In this example, the set of residuals 714 corresponding to the first image region 506 is the first subset 714 of residuals. In such examples, generating the enhancement layer 700 encoding of the second image region 508 comprises obtaining a set of residuals 716 corresponding to the second image region 508, where the set of residuals 716 corresponding to the second image region 508 are quantised with a second, higher quantisation step width. In this example, the set of residuals 716 corresponding to the second image region 508 is the second subset 716 of residuals.

[0133] In some examples, a rate controller is configured to determine one or more encoding properties for residuals in a given image region and to determine one or more geometric properties of a region in which the residuals will be used. For example, the rate controller may be configured to determine the quantisation step width (an example of an encoding parameter) for residuals in a given image region and the size (an example of a geometric property) of the region in which the residuals will be used. In other words, the rate controller may be configured to trade off the size and / or shape of a foveated area with one or more encoding parameters. An example of such an encoding parameter is the quantisation step width of residuals in that foveated area. Another example of such an encoding parameter is the quantisation step width of residuals outside the foveated area. Another example of such an encoding parameter is a bitrate allocation to a base layer. Another example of such an encoding parameter is a bitrate allocation to an enhancement layer. Another example of such an encoding parameter is a metric of complexity for the image.

[0134] Thus, a set of residuals corresponding to a first image region may be quantised with a given quantisation step width.

[0135] The given quantisation step width may be based on a size of the first image region and / or a target bitrate. Some examples comprise determining the given quantisation step width based on the size of the first image region and / or the target bitrate. Some examples comprise quantising the set of residuals corresponding to the first image region in accordance with the determined given quantisation step width.

[0136] The size of the first image region may be based on the given quantisation step width and / or the target bitrate. Some examples comprise determining the size of the first image region based on the given quantisation step width and / or the target bitrate. Some examples comprise selecting, setting and / or adjusting the size of the first image region based on the given quantisation step width and / or the target bitrate.

[0137] Thus, measures (such as methods, systems and computer programs) are provided for processing an image comprising XR content and being for an XR display device. A base layer encoding of the image is obtained. First enhancement layer processing may be performed for a first image region of the image. Performing first enhancement layer processing may comprise a rate controller determining one or more encoding properties for a first set of residuals in a first image region of the image and determining one or more properties (for example, geometric properties) of a first region in which the first set of residuals is to be used. Such determining may comprise the rate controller trading off the one or more encoding properties with the one or more geometric properties. Second enhancement layer processing may be performed for a second image region of the image. Performing second enhancement layer processing may comprise the rate controller determining one or more encoding properties for a second set of residuals in a second image region of the image and determining one or more geometric properties of a second region in which the second set of residuals is to be used.

[0138] Thus, the same high-quality residuals may be used for a relatively small area as opposed to worsening visual quality by using (lower-quality) residuals everywhere in an image. The quality of residuals may correspond to their level of quantisation such that higher quantisation corresponds to lower quality, and lower quantisation corresponds to higher quality. In particular, lower-quality residuals may be more heavily quantised than higher-quality residuals. If the foveated area reaches a given (small) size and if the rate controller determines that the size of the residuals is still to be decreased (for example by increasing the quantisation step width), the rate controller can adjust the trade-off, such that the trade-off seeks to maintain a constant size of the foveated area and increase the quantization step width to achieve a target bit rate. In this way, the rate controller can dynamically trade off, on a frame-by-frame basis, the visual quality provided in the foveated area with the actual size of the foveated area.

[0139] Without performing the above-described second enhancement layer processing, the second subset 716 of residuals would, in effect, be wasted. This is because they are not important, or at least have limited importance, in terms of visual quality as perceived by the viewer.

[0140] The first and second enhancement layer processing may be performed as part of a single enhancement layer processing procedure. For example, the first and second enhancement layer encodings may be generated by the same enhancement layer encoder as each other. In some examples, the encodings are output for display on the XR display device. In some examples, the encodings are stored.

[0141] The encodings comprise the base layer encoding and the enhancement layer 700 encoding of the first image region 506. The encodings may comprise the enhancement layer 700 encoding of the second image region 508 when this is obtained. The encodings may be output to the XR display device for display on the XR display device or may be output to another device for display on the XR display device. Such outputting may comprise streaming. An example of such another device is a computing device that is paired with an XR display device in the form of an XR headset.

[0142] In some examples, the base layer encoding of the image 500 is decodable and / or displayable independently of the enhancement layer 700 encoding(s).

[0143] In terms of decoder-side processing, a base layer encoding of an image 500 may be obtained. An enhancement layer 700 encoding of a first image region 506 (the encoding being in accordance with the first encoding priority) may also be obtained. An enhancement layer 700 encoding of a second image region 508 (the encoding being in accordance with the second encoding priority) may be obtained, or an enhancement layer 700 encoding of the second image region 508 may not be obtained. Such obtaining, or not obtaining, may be by the XR display device. The obtained encodings may be used to reconstruct the image 500. The reconstructed image may be caused to be displayed on the XR display device.

[0144] In some examples, the residuals 714, 716 have intra-independency and so may be said to be “intra-independent”. Thus, there may be no “intra” dependency for such residuals 714, 716. This enables the use of no residuals, or residuals quantised to zero, for a portion of the enhancement layer 700. A codec, such as LCEVC, that does not perform intra coding (e.g. intra prediction coding) may be used to generate such intra- independent residuals, since it is suitable to foveated priority areas that vary from frame to frame. LCEVC is described in WO 2020 / 188273 (PCT / GB2020 / 050695) and WO 2019 / 111010 (PCT / GB2018 / 053552), the entire contents of which are incorporated herein by reference. Additionally, LCEVC may also allow XR content to be streamed with ultra-low-latency within wireless communication technology limitations, such as those described above. If the residuals 714, 716 had intra-dependency, then decoding values of some of the residuals may require values of other residuals. For example, decoding values of some of the residuals 714 in the first region 706 of the enhancement layer 700 may require values of some of the residuals 716 in the second region 708 of the enhancement layer 700. If the residuals 716 in the second region 708 of the enhancement layer 700 were not provided at all, or were quantised heavily, the enhancement layer 700 might not be useable or at least might not be as effective. For example, if the residuals 716 in the second region 708 of the enhancement layer 700 were heavily quantised, then intra prediction would be less accurate but may still provide a useable prediction. If, however, the residuals 716 in the second region 708 of the enhancement layer 700 were completely omitted or all quantised to zero, then they would not be useable for intra prediction.

[0145] Referring to Figure 8, there is shown a schematic diagram of another example of an enhancement layer 800.

[0146] The example enhancement layer 800 shown in Figure 8 corresponds generally to the example enhancement layer 700 shown in Figure 7. Reference signs used in Figure 8 are the same as those used in Figure 7 for the same or similar features, but incremented by 100.

[0147] However, Figure 8 only depicts the first subset 814 of residuals.

[0148] It can be seen from Figure 8 that fewer residuals can be used than in the example enhancement layer 700 shown in Figure 7, without negatively impacting visual quality, either significantly or at all.

[0149] Referring to Figure 9, there is shown a schematic diagram of another example of an enhancement layer 900.

[0150] The example enhancement layer 900 shown in Figure 9 corresponds generally to the example enhancement layers 700, 800 shown in Figures 7 and 8. Reference signs used in Figure 9 are the same as those used in Figures 7 and 8 for the same or similar features, but incremented by 200 and 100 respectively.

[0151] However, in this example, residuals that would have been generated outside of the FOV of the viewer are, in effect, redistributed to be inside the FOV of the viewer.

[0152] In more detail, in some examples, a data-saving measure is determined. The data-saving measure represents an amount of data saved by performing the second enhancement layer processing for the second image region 508, for example compared to performing the first enhancement layer processing for the second image region 508. Some or all of the data saved in this manner may be allocated to the first enhancement layer processing for the first image region 506.

[0153] In particular, Figure 7 shows six residuals 714 in the first region 706 of the enhancement layer 700 and seven residuals 716 in the second region 708 of the enhancement layer 700, i.e. thirteen residuals in total.

[0154] In contrast, Figure 9 shows thirteen residuals 918 in the first region 906 of the enhancement layer 900 and no residuals in the second region 908 of the enhancement layer 900, i.e. still thirteen residuals in total.

[0155] By limiting the region(s) of the enhancement layer in which residuals are allowed, or at least residuals with limited quantisation are allowed, visual quality can be increased for the most important region(s) of an image, potentially without increasing the amount of data transmitted compared to where residuals are allowed to be in any part of the enhancement layer.

[0156] By placing residuals mostly, or solely, in a foveated area, quality can be increased significantly with a relatively low bitrate. This can be particularly effective (in particular for very high-resolution signals) in bandwidth-constrained systems, such as where low-bandwidth wireless communication technologies are used. For example, if residuals are limited to an area that is one fifth of the total screen area, the equivalent quality of a five-times-larger enhancement may be used without increasing the bitrate.

[0157] Referring to Figure 10, there is shown a schematic diagram of another example of an image 1000.

[0158] The example image 1000 shown in Figure 10 corresponds generally to the example image 500 shown in Figure 5. Reference signs used in Figure 10 are the same as those used in Figure 5 for the same or similar features, but incremented by 500.

[0159] However, in this example, the image 1000 comprises a third image region 1014.

[0160] In this example, the third image region 1014 comprises first and second subregions 1014-1, 1014-2. In this example, the first sub-region 1014-1 is in the left-side image portion 1002 and the second sub-region 1014-2 is in the right-side image portion 1004.

[0161] In this example, the first sub-region 1014-1 is centred in the centre of the leftside image portion 1002 and the second sub-region 1014-2 is centred in the centre of the right-side image portion 1004. The first sub-region 1014-1 may therefore be said to be central to the left-side image portion 1002 and the second sub-region 1014-2 may be said to be central to the right-side image portion 1004.

[0162] In this example, the first sub-region 1014-1 is rectangular with width, V3;1, and height, H3 1, where Wi r< W3 1< w / 2 and < H3 1< h. In this example, the second sub-region 1014-2 is rectangular with width V / 32, and height, H3 2, where this example, W3 1= W3 2and H3 1= H3 2. The first and / or second sub-region 1014-1, 1014-2 may have a different shape and / or dimensions in other examples.

[0163] In this example, the first and second sub-regions 1014-1, 1014-2 of the third image region 1014 are non-overlapping. The third image region 1014 may therefore be said to be non-contiguous.

[0164] In some examples, third enhancement layer processing is performed for the third image region 1014. Performing third enhancement layer processing may comprise obtaining an enhancement layer encoding of the third image region 1014 in accordance with a third encoding priority. The third encoding priority may be different from the first and second encoding priorities described above with reference to Figure 7.

[0165] In some examples, the third image region 1014 is between the first and second image regions 1006, 1008. In some examples, the third encoding priority is between the first and second encoding priorities.

[0166] The third image region 1014 may correspond to a peripheral region of the FOV of the user of the XR display device.

[0167] Therefore, more than two encoding priorities may be used. Some residuals may not be used at all. In other words, such residuals may not be generated at all or may be generated but quantised to zero. These would be the lowest priority residuals. Residuals in the peripheral region may have intermediate priority and may be encoded, but heavily quantised. Residuals in the central field of view region may have the highest priority. In some examples, residuals are therefore also used beyond a foveated priority map corresponding to the first image region 1006. However, in some examples, such residuals are restricted to being only in one or more static areas of a screen. An example of a static area of a screen is an area of which the visual content does not change between at least some successive images in a sequence of images. This allows static graphics overlays, for example, to have continuously good quality, with very limited bitrate being used.

[0168] Referring to Figure 11, there is shown a schematic diagram of an example of a sequence of images 1100. Some reference signs used in Figure 11 are the same as those used in Figure 5 for the same or similar features.

[0169] In this example, base layer and enhancement layer encodings of additional images are created, for example for output for display on the XR display device. In this example, there are three additional images 1102, 1104, 1106. However, in practice, there may be significantly more than three additional images in the sequence of images 1100. The additional images 1102, 1104, 1106 are additional to the example image 500 shown in Figure 5 and reproduced in Figure 11. The image 500 and the additional images 1102, 1104, 1106 are part of the sequence of images 1100 comprising XR content for the XR display device. The sequence of images 1100 may correspond to an XR stream. The XR stream may comprise a stereoscopic XR stream.

[0170] In some examples, the first and second image regions (506, 508 for the image 500) have fixed positions across the sequence of images 1110. This can reduce, or even avoid, using resources to determine their positions across the sequence of images 1110. Fixed positions may also be referred to as “static”, “set”, or “set frame” positions. Using a fixed priority map may therefore reduce or avoid positioning-related calculation costs. The first and second image regions (506, 508 for the image 500) may have fixed shapes and / or sizes across the sequence of images 1110.

[0171] In some examples, the first and second image regions (506, 508 for the image 500) have dynamic positions, shapes and / or sizes across the sequence of images 1110. The dynamic positions, shapes and / or sizes may be determined in various different ways.

[0172] In some examples, eye-tracking data is received from the XR display device, and the positions, shapes and / or sizes of the first and second image regions (506, 508 for the image 500) are determined based on the eye-tracking data. The eye-tracking data may also be referred to as “eye-tracking information”. As explained above, the XR display device may be or may comprise an XR headset. In such examples, the first and second image regions (506, 508 for the image 500) may be centred or may not be centred. Thus, an encoder may receive feedback from a XR display device such that the first and second image regions (506, 508 for the image 500) have positions, shapes and / or sizes linked to eye-tracking.

[0173] Thus, residuals may be placed in selected parts of an overall screen. This may focus detail on where detail matters most, for example in accordance with the natural foveation of the human vision system. By using eye-tracking data, real-time feedback may be provided to an encoder. This may introduce a delay. However, the delay may be a tolerable delay, such as a 50-100ms delay. The encoder may therefore put detail precisely around the area at which the viewer is actually looking, rather than assuming that the viewer is looking at roughly the centre of the FOV.

[0174] LCEVC is particularly effective in this regard. For example, the characteristics of LCEVC enable this level of control in a manner that is not possible, or at least that is not as effective, with at least some single-layer codecs. Thus, the specific combination of using eye-tracking data to inform an LCEVC encoder to select a particular set of residuals in the context of XR pixel streaming is particularly effective.

[0175] In the absence of eye-tracking and eye-tracking data, the foveated area may be in a fixed, centred location across a sequence of frames. The foveated area may operate like a pair of (variable-size) eyeglasses on the face of the viewer. When eye-tracking is available, an encoder may receive real-time information on where the viewer is looking. The encoder may, consequently, centre the foveated area there.

[0176] In some examples, the shape of the foveated area is based at least in part on the location(s) at which the viewer is looking. In some examples, the shape of the foveated area is, alternatively or additionally, based at least in part on the location(s) at which the viewer is likely to look next. Such prediction may be based on eye motion. For example, the prediction may be based on recent eye motion. This may allow to several tens of milliseconds to be gained, balancing end-to-end system latency.

[0177] The examples described herein may be used in accordance with various different use cases of XR pixel streaming. Examples of such use cases include, but are not limited to, game streaming, six degrees of freedom (6DoF) storytelling with a prerendered zone of view, and design review of digital twins and / or digital doubles.

[0178] Referring to Figure 12, there is shown a schematic block diagram of an example of an apparatus 1200. In an example, the apparatus 1200 comprises an encoder. In another example, the apparatus 1200 comprises a decoder. In other examples, the apparatus 1200 comprises neither an encoder nor a decoder but is configured to communicate with an encoder and / or a decoder.

[0179] Examples of apparatus 1200 include, but are not limited to, a mobile computer, a personal computer system, a wireless device, base station, phone device, desktop computer, laptop, notebook, netbook computer, mainframe computer system, handheld computer, workstation, network computer, application server, storage device, a consumer electronics device such as a camera, camcorder, mobile device, video game console, handheld video game device, an XR headset, or in general any type of computing or electronic device.

[0180] In this example, the apparatus 1200 comprises one or more processors 1201 configured to process information and / or instructions. The one or more processors 1201 may comprise a central processing unit (CPU). The one or more processors 1201 are coupled with a bus 1202. Operations performed by the one or more processors 1201 may be carried out by hardware and / or software. The one or more processors 1201 may comprise multiple co-located processors or multiple disparately located processors.

[0181] In this example, the apparatus 1200 comprises computer-useable volatile memory 1203 configured to store information and / or instructions for the one or more processors 1201. The computer-useable volatile memory 1203 is coupled with the bus 1202. The computer-useable volatile memory 1203 may comprise random access memory (RAM).

[0182] In this example, the apparatus 1200 comprises computer-useable non-volatile memory 1204 configured to store information and / or instructions for the one or more processors 1201. The computer-useable non-volatile memory 1204 is coupled with the bus 1202. The computer-useable non-volatile memory 1204 may comprise read-only memory (ROM).

[0183] In this example, the apparatus 1200 comprises one or more data-storage units 1205 configured to store information and / or instructions. The one or more data-storage units 1205 are coupled with the bus 1202. The one or more data-storage units 1205 may for example comprise a magnetic or optical disk and disk drive or a solid-state drive (SSD). In this example, the apparatus 1200 comprises one or more input / output (I / O) devices 1206 configured to communicate information to and / or from the one or more processors 1201. The one or more I / O devices 1206 are coupled with the bus 1202. The one or more I / O devices 1206 may comprise at least one network interface. The at least one network interface may enable the apparatus 1200 to communicate via one or more data communications networks. Examples of data communications networks include, but are not limited to, the Internet and a Local Area Network (LAN). The one or more I / O devices 1206 may enable a user to provide input to the apparatus 1200 via one or more input devices (not shown). The one or more input devices may include for example a remote control, one or more physical buttons etc. The one or more I / O devices 1206 may enable information to be provided to a user via one or more output devices (not shown). The one or more output devices may for example include a display screen.

[0184] Various other entities are depicted for the apparatus 1200. For example, when present, an operating system 1207, image processing module 1208, one or more further modules 1209, and data 1210 are shown as residing in one, or a combination, of the computer-usable volatile memory 1203, computer-usable non-volatile memory 1204 and the one or more data-storage units 1205. The data signal processing module 1208 may be implemented by way of computer program code stored in memory locations within the computer-usable non-volatile memory 1204, computer-readable storage media within the one or more data-storage units 1205 and / or other tangible computer- readable storage media. Examples of tangible computer-readable storage media include, but are not limited to, an optical medium (e.g., CD-ROM, DVD-ROM or Blu- ray), flash memory card, floppy or hard disk or any other medium capable of storing computer-readable instructions such as firmware or microcode in at least one ROM or RAM or Programmable ROM (PROM) chips or as an Application Specific Integrated Circuit (ASIC).

[0185] The apparatus 1200 may therefore comprise a data signal processing module 1208 which can be executed by the one or more processors 1201. The data signal processing module 1208 can be configured to include instructions to implement at least some of the operations described herein. During operation, the one or more processors 1201 launch, run, execute, interpret or otherwise perform the instructions in the signal processing module 1208.

[0186] Although at least some aspects of the examples described herein with reference to the drawings comprise computer processes performed in processing systems or processors, examples described herein also extend to computer programs, for example computer programs on or in a carrier, adapted for putting the examples into practice. The carrier may be any entity or device capable of carrying the program.

[0187] It will be appreciated that the apparatus 1200 may comprise more, fewer and / or different components from those depicted in Figure 12.

[0188] The apparatus 1200 may be located in a single location or may be distributed in multiple locations. Such locations may be local or remote.

[0189] The techniques described herein may be implemented in software or hardware, or may be implemented using a combination of software and hardware. They may include configuring an apparatus to carry out and / or support any or all of techniques described herein.

[0190] A bitstream may be provided. The bitstream may comprise a base layer encoding of an image comprising XR content and being for an XR display device. The bitstream may comprise an enhancement layer encoding of a first image region of the image. The enhancement layer encoding of the first image region may have been generated in accordance with a first encoding priority. In some examples, the bitstream comprises an enhancement layer encoding of the second image region, where enhancement layer encoding of the second image region has been generated in accordance with a second, lower encoding priority. In other examples, the bitstream does not comprise an enhancement layer encoding of the second image region.

[0191] In a specific example, a codec feature of carrying local user data is used for post-processing operations. The codec may be a multi-layer, scalable codec. LCEVC is one example of such a codec that has such a feature. For example, the user data may specify one or more specific areas to which a decoder is to apply one or more given post-processing operations. For instance, a de-ringing and / or an Al-based postprocessing operation may be costly for an entire 8K-pixel display. However, it may be readily feasible to apply a post-processing operation (e.g. a de-ringing and / or Al-based post-processing operation), at decoding, to a given foveated priority area. A foveated priority area signalled to the decoder for post-processing may be different from a foveated priority area used by an encoder to prioritise residuals. Also, as explained above, the foveated priority area used to prioritise residuals may or may not be dependent on real-time eye-tracking data received from a headset.

[0192] Thus, measures (such as methods, systems and computer programs) are provided for processing an image comprising XR content and being for an XR display device. A base layer encoding of the image is obtained. Enhancement layer processing may be performed for a first image region of the image. Performing such enhancement layer processing may comprise obtaining an enhancement layer encoding of the first image region. The enhancement layer encoding of the first image region may be prioritised. Such prioritising may be with respect to an enhancement layer encoding of one or more further image regions of the image. For example, the first image region may correspond to a first foveated priority area. User data may be generated. The user data may specify a second image region of the image to which a decoder is to apply one or more given post-processing operations. The first and second image regions of the image may be different from each other. The user data may be signalled to the decoder. The techniques described herein may be applied to a depth map. The depth map may be sent with video data for cases in which the combination of video and depth are used for depth-based reprojection and, thus, frame rate adaptation to a display frame rate. For the depth map, and in a similar manner to that described above for video data, residuals may be prioritised for a foveated area. Such prioritisation may involve sending residuals only for the foveated area, using a lower quantisation step width for residuals for the foveated area relative to a quantisation step width for residuals outside the foveated area, or otherwise. The foveated area of the depth map may or may not depend on eye-tracking data. In this manner, precise depth processing may be provided where it is most effective and without wasting large amounts of bitrate. In a specific example, the depth map may be sent at a low resolution. The depth map may be border-snapped to objects according to luma or chroma information, using any available method. The depth map may then be corrected with residual data for the foveated area at which the user is looking. This makes the depth map more precise in the foveated area.

[0193] Thus, a base layer encoding of a depth map may be obtained. The depth map may indicate depth information associated with an image. First enhancement layer processing may be performed for a first region of the depth map. Performing first enhancement layer processing for the first region of the depth map may comprise obtaining an enhancement layer encoding of the first region of the depth map. The enhancement layer encoding of the first region of the depth map may be in accordance with a first depth map encoding priority. Second enhancement layer processing may be performed for a second region of the depth map. Performing second enhancement layer processing for the second region of the depth map may comprise: inhibiting an enhancement layer encoding of the second region of the depth map from being obtained; or obtaining an enhancement layer encoding of the second region of the depth map, the enhancement layer encoding of the second region of the depth map being in accordance with a second, lower depth map encoding priority. The first region of the depth map may correspond to a first image region of the image. The second region of the depth map may correspond to a second image region of the image.

[0194] In examples described above, the image is a stereoscopic image. However, other types of image having left and right image portions may be used.

[0195] For example, one of the left and right image portions may comprise a view of a scene and the other of the left and right image portions may comprise a corresponding depth map. It may be acceptable to sacrifice some quality of the depth map in favour of the view of the scene, for instance.

[0196] In another example, one of the left and right image portions may comprise a view of a sporting event from a first angle, and the other of the left and right image portions may comprise a view of the sporting event from a second, different angle. It may be acceptable to sacrifice some quality of one of the views of the sporting event (for example a zoomed-out view) in favour of the other view of the sporting event.

[0197] Instead of the example techniques described above, a base layer encoding of an image may be obtained in which a first image region of the image is not distorted and in which a second image region of the image is distorted. Such distortion may involve removing some pixel values in the second image region. A decoder may then interpolate the remaining pixel values in the second image region to attempt to reconstruct the removed pixel values. Imperfections, resulting from the interpolation, may be tolerable in the second image region where the second image region is outside the FOV of a viewer. However, such an interpolation-based technique can introduce latency, which can negatively impact performance in the context of XR.

[0198] The example techniques described above may be combined with methods described in (PCT / GB2017 / 05214, WO2018 / 015764). For example, a device (e.g. user device, XR display device, XR headset and so forth) may be configured to receive one or more of the base layer and the enhancement layer. The enhancement layer may be obtained as described above, e.g. by performing first enhancement layer processing for a first image region of the image, wherein performing first enhancement layer processing comprises obtaining an enhancement layer encoding of the first image region, the enhancement layer encoding of the first image region being in accordance with a first encoding priority; and performing second enhancement layer processing for a second image region of the image, wherein performing second enhancement layer processing comprises: inhibiting an enhancement layer encoding of the second image region from being obtained; or obtaining an enhancement layer encoding of the second image region, the enhancement layer encoding of the second image region being in accordance with a second, lower encoding priority. The device may be configured to obtain a decoding of the base encoded layer. The device may be configured to obtain a decoding of the enhancement encoded layer. The device may be configured to combine the decoded of the base encoded layer and the decoded enhancement encoded layer to obtain an enhanced signal (e.g. image). In particular, the device may be configured to decode and / or apply (e.g. only) a portion of the received enhancement layer. The portion may be less than a whole. The device may be configured to decode only a portion of the received enhancement layer. The device may be configured to apply only a portion of the decoded enhancement layer. The portion may correspond to a region of the image at which the viewer is looking at and / or is expected to look at. The portion may correspond to a region of the image at which the viewer is looking at and / or is expected to look at, at a time of decoding and or displaying. In other words the portion may by a FOV and / or a foveated FOV and / or a expected FOV and / or an expected foveated FOV.

[0199] In this way, efficient decoding based on real-time eye-tracking information is synergistically combined with efficient encoding (by way of encoding only enhancements within a FOV, foveated FOV, expected FOV, expected foveated FOV, and so forth).

[0200] This synergistic combination may be considered as a "dual pass": in the first pass we describe encoding residuals only for a (generally central) portion (say X%) as described above, this may be performed using independent tiles of residuals. Then, as a second pass, we describe, e.g. based on eye-tracking info, sending and / or decoding and / or applying only the residuals (e.g. tile(s)) that are in focus (e.g. a FOV, foveated FOV, expected FOV, expected foveated FOV, and so forth), at a given time. In this way, there is only encoded enhancement data (i.e. enhancement layer) for X% of the image and only Y% of the image has enhancement data decoded and / or applied to it (Y<X). This can be used to achieve full resolution of the display only where a user is looking, without having to encode and send full resolution data across the entire image. As there is some error with eye tracking prediction, enhancement data may be generated for X% of the image (centered around the predicted eye direction), however, only a portion of this enhancement data may be decoded and / or applied, where that portion may be centered around an updated predicted eye direction (i.e. gaze). Wherein the updated predicted eye direction is calculated at a time closer to the display time than when the enhancement data was encoded.

[0201] We described the following clauses. The described clauses may be combined with any of the description and any of the claims. In particular, the target region of the clauses may correspond to a subset of the first image region. The enhancement data of the clauses may correspond to the enhancement layer encoding.

[0202] 1. A decoder device configured to: receive data useable to generate data for representing a data signal at a first level of quality; receive enhancement data useable to generate data for representing the data signal at a second level of quality based on the data for representing the data signal at the first level of quality, the second level of quality being higher than the first level of quality; generate data for representing a target region of the data signal at a target level of quality using a selected portion of the received enhancement data, the selected portion of the received enhancement data being associated with the target region of the data signal, the target level of quality being higher than the first level of quality; and generate data for representing a further region of the data signal at a level of quality lower than the target level of quality.

[0203] 2. A decoder device according to clause 1, wherein the target level of quality is the second level of quality.

[0204] 3. A decoder device according to clause 1, wherein the target level of quality is between the first level of quality and the second level of quality.

[0205] 4. A decoder device according to any of clauses 1 to 3, wherein the level of quality of the further region of the data signal is between the first level of quality and the target level of quality.

[0206] 5. A decoder device according to any of clauses 1 to 4, the decoder device being configured to generate the data for representing the further region of the data signal using a selected further portion of the enhancement data, the selected further portion of the enhancement data being associated with the further region of the data signal.

[0207] 6. A decoder device according to any of clauses 1 to 3, wherein the level of quality of the further region of the data signal is the first level of quality.

[0208] 7. A decoder device according to any of clauses 1 to 6, wherein the further region of the data signal at least partly surrounds the target region of the data signal.

[0209] 8. A decoder device according to any of clauses 1 to 7, the decoder device being configured to use data associated with a field of view and / or data associated with one or more gaze positions to identify the target region of the data signal. 9. A decoder device according to clause 8, the decoder device being configured to select the target region of the data signal so as to be in the field of view.

[0210] 10. A decoder device according to clause 8 or 9, the decoder device being configured to select at least part of the further region of the data signal so as to be in the field of view.

[0211] 11. A decoder device according to any of clauses 8 to 10, the decoder device being configured to: monitor the field of view and / or the one or more gaze positions at multiple points in time; and determine a position of a target region in a subsequent data signal dependent on the field of view and / or the one or more gaze positions at a point in time associated with the subsequent data signal.

[0212] 12. A decoder device according to clause 11, wherein the target region of the data signal is associated with one or more data signal tiles and the further region of the data signal is associated with one or more data signal tiles, wherein at least one of the data signal tiles associated with the further region of the data signal is within the target region in the subsequent data signal.

[0213] 13. A decoder device according to any of clauses 1 to 12, the decoder device being configured to: generate data for representing the target region of the data signal at the first level of quality; and use the generated data for representing the target region of the data signal at the first level of quality to generate the data for representing the target region of the data signal at the target level of quality.

[0214] 14. A decoder device according to any of clauses 1 to 13, the decoder device being configured to operate in accordance with a hierarchical data signal processing arrangement, the hierarchical data signal processing arrangement comprising at least one layer having a set of sub-layers, each sub-layer being associated with a respective level of quality.

[0215] 15. A decoder device according to clause 14, the decoder device being configured not to use enhancement data associated with at least one of the sub-layers in a first operating mode of the decoder device.

[0216] 16. A decoder device according to clause 14, the decoder device being configured use enhancement data associated with all of the sub-layers to generate the data for representing the target region of the data signal at a target level of quality in a second operating mode of the decoder device.

[0217] 17. A decoder device according to any of clauses 1 to 16, the decoder device being configured to operate in accordance with a hierarchical data signal processing arrangement, the hierarchical data signal processing arrangement comprising a first layer having a first set of sub-layers and a second layer having a set of sub-layers, each sub-layer being associated with a respective level of quality.

[0218] 18. A decoder device according to clause 17, the decoder device being configured to use enhancement data associated with at least one sub-layer of the first and second layers to generate the data for representing the target region of the data signal at the target level of quality in a third operating mode of the decoder device.

[0219] 19. A decoder device according to any of clauses 14 to 18, wherein the first level of quality corresponds to a level of quality associated with the lowest sub-layer in the hierarchical data signal processing arrangement.

[0220] 20. A decoder device according to any of clauses 14 to 19, wherein the second level of quality corresponds to a level of quality associated with the highest sub-layer in the hierarchical data signal processing arrangement. 21. A decoder device according to any of clauses 14 to 20, wherein the target level of quality corresponds to a level of quality associated with a sub-layer between the highest sub-layer and the lowest sub-layer in the hierarchical data signal processing arrangement.

[0221] 22. A decoder device according to any of clauses 1 to 21, wherein the data signal comprises image data.

[0222] 23. A decoder device according to any of clauses 1 to 22, wherein the data signal comprises video data.

[0223] 24. A decoder device according to any of clauses 1 to 23, the decoder device being configured to identify the target region of the data signal.

[0224] 25. A decoder device according to any of clauses 1 to 24, the decoder device being configured to select the portion of the enhancement data associated with the target region of the data signal.

[0225] 26. A decoder device according to any of clauses 1 to 25, the decoder device being comprised in virtual reality equipment, medical imaging equipment, machine vision equipment and / or or a mobile communications device.

[0226] 27. Virtual reality equipment comprising a decoder device according to any of clauses 1 to 25.

[0227] 28. Medical imaging equipment comprising a decoder device according to any of clauses 1 to 25.

[0228] 29. Machine vision equipment comprising a decoder device according to any of clauses 1 to 25. 30. A mobile communications device comprising a decoder device according to any of clauses 1 to 25.

[0229] 31. A method comprising, at a decoder device: receiving data useable to generate data for representing a data signal at a first level of quality; receiving enhancement data useable to generate data for representing the data signal at a second level of quality based on data for representing the data signal at the first level of quality, the second level of quality being higher than the first level of quality; generating data for representing a target region of the data signal at a target level of quality using a selected portion of the received enhancement data, the selected portion of the received enhancement data being associated with the target region of the data signal, the target level of quality being higher than the first level of quality; and generating data for representing a further region of the data signal at a level of quality lower than the target level of quality.

[0230] 32. A method according to clause 31, wherein the target level of quality is the second level of quality.

[0231] 33. A method according to clause 31, wherein the target level of quality is between the first level of quality and the second level of quality.

[0232] 34. A method according to any of clauses 31 to 33, wherein the level of quality of the further region of the data signal is between the first level of quality and the target level of quality.

[0233] 35. A method according to any of clauses 31 to 34, wherein generating the data for representing the further region of the data signal comprises using a selected further portion of the enhancement data, the selected further portion of the enhancement data being associated with the further region of the data signal. 36. A method according to any of clauses 31 to 33, wherein the level of quality of the further region of the data signal is the first level of quality.

[0234] 37. A method according to any of clauses 31 to 36, wherein the further region of the data signal at least partly surrounds the target region.

[0235] 38. A method according to any of clauses 31 to 37, the method comprising using data associated with a field of view and / or data associated with one or more gaze positions to identify the target region of the data signal.

[0236] 39. A method according to clause 38, the method comprising selecting the target region of the data signal so as to be in the field of view.

[0237] 40. A method according to clause 38 or 39, the method comprising selecting at least part of the further region of the data signal so as to be in the field of view.

[0238] 41. A method according to any of clauses 38 to 40, the method comprising: monitoring the field of view and / or the one or more gaze positions at multiple points in time; and determining a position of a target region in a subsequent data signal dependent on the field of view and / or the one or more gaze positions at a point in time associated with the subsequent data signal.

[0239] 42. A method according to clause 41, wherein the target region of the data signal is associated with one or more data signal tiles and the further region of the data signal is associated with one or more data signal tiles, wherein at least one of the data signal tiles associated with the further region of the data signal is within the target region in the subsequent data signal.

[0240] 43. A method according to any of clauses 31 to 42, the method comprising: generating data for representing the target region of the data signal at the first level of quality; and using the generated data for representing the target region of the data signal at the first level of quality to generate the data for representing the target region of the data signal at the target level of quality.

[0241] 44. A method according to any of clauses 31 to 43, the method comprising operating in accordance with a hierarchical data signal processing arrangement, the hierarchical data signal processing arrangement comprising at least one layer having a set of sublayers, each sub-layer being associated with a respective level of quality.

[0242] 45. A method according to clause 44, the method comprising not using enhancement data associated with at least one of the sub-layers in a first operating mode of the decoder device.

[0243] 46. A method according to clause 44, the method comprising using enhancement data associated with all of the sub-layers to generate the data for representing the target region of the data signal at a target level of quality in a second operating mode of the decoder device.

[0244] 47. A method according to any of clauses 31 to 46, the method comprising operating in accordance with a hierarchical data signal processing arrangement, the hierarchical data signal processing arrangement comprising a first layer having a first set of sublayers and a second layer having a set of sub-layers, each sub-layer being associated with a respective level of quality.

[0245] 48. A method according to clause 47, the method comprising using enhancement data associated with at least one sub-layer of the first and second layers to generate the data for representing the target region of the data signal at the target level of quality in a third operating mode of the decoder device.

[0246] 49. A method according to any of clauses 44 to 48, wherein the first level of quality corresponds to a level of quality associated with the lowest sub-layer in the hierarchical data signal processing arrangement. 50. A method according to any of clauses 44 to 49, wherein the second level of quality corresponds to a level of quality associated with the highest sub-layer in the hierarchical data signal processing arrangement.

[0247] 51. A method according to any of clauses 44 to 50, wherein the target level of quality corresponds to a level of quality associated with a sub-layer between the highest sub-layer and the lowest sub-layer in the hierarchical data signal processing arrangement.

[0248] 52. A method according to any of clauses 31 to 51, wherein the data signal comprises image data.

[0249] 53. A method according to any of clauses 31 to 52, wherein the data signal comprises video data.

[0250] 54. A method according to any of clauses 31 to 53, the method comprising identifying the target region of the data signal.

[0251] 55. A method according to any of clauses 31 to 54, the method comprising selecting the portion of the enhancement data associated with the target region of the data signal.

[0252] 56. A computer program comprising instructions which, when executed, cause a decoder device to perform a method according to any of clauses 31 to 55.

[0253] 57. A computer-readable medium comprising a computer program according to clause 56.

[0254] It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

Claims

CLAIMS1. A method of processing an image, the image comprising extended reality, XR, content and being for an XR display device, the method comprising: obtaining a base layer encoding of the image; performing first enhancement layer processing for a first image region of the image, wherein performing first enhancement layer processing comprises obtaining an enhancement layer encoding of the first image region, the enhancement layer encoding of the first image region being in accordance with a first encoding priority; and performing second enhancement layer processing for a second image region of the image, wherein performing second enhancement layer processing comprises: inhibiting an enhancement layer encoding of the second image region from being obtained; or obtaining an enhancement layer encoding of the second image region, the enhancement layer encoding of the second image region being in accordance with a second, lower encoding priority.

2. A method according to claim 1, wherein positions of the first and second image regions are independent of pixel values of pixels of the image.

3. A method according to claim 1 or 2, comprising: outputting the encodings for display on the XR display device.

4. A method according to any of claims 1 to 3, comprising: outputting, for display on the XR display device, base layer and enhancement layer encodings of additional images, the image and the additional images being part of a sequence of images comprising XR content for the XR display device, wherein the first and second image regions have fixed positions, shapes and / or sizes across the sequence of images.

5. A method according to any of claims 1 to 3, comprising:outputting, for display on the XR display device, base layer and enhancement layer encodings of additional images, the image and the additional images being part of a sequence of images comprising XR content for the XR display device, wherein the first and second image regions have dynamic positions, shapes and / or sizes across the sequence of images.

6. A method according to claim 5, comprising: receiving eye-tracking data from the XR display device; and determining the positions, shapes and / or sizes of the first and second image regions based on the eye-tracking data.

7. A method according to any of claims 1 to 6, wherein the first image region has a rectangular shape, and wherein a height of the rectangular shape of the first image region is less than a width of the rectangular shape.

8. A method according to claim 7, wherein the first image region has a landscape orientation.

9. A method according to any of claims 1 to 8, wherein performing the second enhancement layer processing comprises: inhibiting the enhancement layer encoding of the second image region from being obtained.

10. A method according to claim 9, wherein obtaining the enhancement layer encoding of the first image region comprises: obtaining a set of residuals corresponding to the first image region, the set of residuals corresponding to the first image region being quantised with a given quantisation step width, and wherein: the given quantisation step width is based on a size of the first image region and / or a target bitrate; and / orthe size of the first image region is based on the given quantisation step width and / or the target bitrate.

11. A method according to any of claims 1 to 8, wherein performing the second enhancement layer processing comprises: obtaining the enhancement layer encoding of the second image region.

12. A method according to any of claims 1 to 11, wherein obtaining the enhancement layer encoding of the second image region comprises: obtaining a set of residuals corresponding to the second image region, the set of residuals corresponding to the second image region being quantised to zero.

13. A method according to any of claims 1 to 11, wherein obtaining the enhancement layer encoding of the first image region comprises: obtaining a set of residuals corresponding to the first image region, the set of residuals corresponding to the first image region being quantised with a first quantisation step width, and wherein generating the enhancement layer encoding of the second image region comprises: obtaining a set of residuals corresponding to the second image region, the set of residuals corresponding to the second image region being quantised with a second, higher quantisation step width.

14. A method according to claim 13, wherein: the first quantisation step width is based on a size of the first image region and / or a target bitrate; and / or the size of the first image region is determined based on the first quantisation step width and / or the target bitrate.

15. A method according to any of claims 1 to 14, wherein the second image region surrounds the first image region.

16. A method according to any of claims 1 to 15, wherein the first image region corresponds to an expected field of view of a user of the XR display device.

17. A method according to any of claims 1 to 16, wherein the image is a stereoscopic image having a left-side image portion and a right-side image portion.

18. A method according to claim 17, wherein the second image region includes a boundary between the left-side and right-side image portions.

19. A method according to claim 17 or 18, wherein the first image region comprises first and second sub-regions, and wherein the first and second sub-regions are nonoverlapping.

20. A method according to claim 19, wherein the first sub-region is central to one of the left-side and right-side image portions, and wherein the second sub-region is central to the other of the left-side and right-side image portions.

21. A method according to claim 19 or 20, wherein the first and / or second subregion has a rectangular shape, and wherein a height of the rectangular shape of the first and / or second sub-region is less than a width of the rectangular shape of the first and / or second sub-region.

22. A method according to any of claims 1 to 21, wherein inhibiting the enhancement layer encoding of the second image region from being obtained comprises: causing an enhancement layer mask to be applied to the second image region.

23. A method according to any of claims 1 to 22, wherein the second image region includes at least part of an edge of the image.

24. A method according to any of claims 1 to 23, comprising:performing third enhancement layer processing for a third image region of the image, wherein performing third enhancement layer processing comprises obtaining an enhancement layer encoding of the third image region in accordance with a third, different encoding priority.

25. A method according to claim 24, wherein the third image region is between the first and second image regions, and the third encoding priority is between the first and second encoding priorities.

26. A method according to any of claims 1 to 25, wherein the enhancement layer comprises residuals resulting from upsampler-downsampler asymmetry and / or residuals resulting from encoder-decoder errors.

27. A method according to any of claims 1 to 26, wherein the base layer encoding of the image comprises: a first base layer encoding portion corresponding to the first image region; and a second base layer encoding portion corresponding to the second image region, wherein the first base layer encoding portion has a first encoding quality and wherein the second base layer encoding portion has a second, lower encoding quality.

28. A method according to any of claims 1 to 27, wherein the XR content comprises XR pixel streaming content.

29. A method according to any of claims 1 to 28, wherein the enhancement layer encoding of the first image region and / or the enhancement layer encoding of the second image region is a multi-layer-scalable-codec-derived enhancement layer encoding.

30. A method according to any of claims 1 to 29, comprising: outputting user data specifying one or more areas of the image in which one or more post-processing operations are to be applied.

31. A method according to claim 30, wherein the one or more areas of the image are different from the first and / or second image region of the image.

32. A method according to claim 30 or 31, wherein the one or more post-processing operations comprise a de-ringing post-processing operation.

33. A method according to any of claims 1 to 32, comprising obtaining a base layer encoding of a depth map, the depth map indicating depth information associated with the image; performing first enhancement layer processing for a first region of the depth map, wherein performing first enhancement layer processing for the first region of the depth map comprises obtaining an enhancement layer encoding of the first region of the depth map, the enhancement layer encoding of the first region of the depth map being in accordance with a first depth map encoding priority; and performing second enhancement layer processing for a second region of the depth map, wherein performing second enhancement layer processing for the second region of the depth map comprises: inhibiting an enhancement layer encoding of the second region of the depth map from being obtained; or obtaining an enhancement layer encoding of the second region of the depth map, the enhancement layer encoding of the second region of the depth map being in accordance with a second, lower depth map encoding priority.

34. A method according to claim 33, wherein the first region of the depth map corresponds to the first image region of the image.

35. A method of processing an image, the image comprising extended reality, XR, content and being for an XR display device, the method comprising: obtaining a base layer encoding of the image; obtaining an enhancement layer encoding of a first image region of the image, the enhancement layer encoding of the first image region being in accordance with a first encoding priority; andeither: obtaining an enhancement layer encoding of the second image region, the enhancement layer encoding of the second image region being in accordance with a second, lower encoding priority; or not obtaining an enhancement layer encoding of the second image region, wherein the method further comprises: using the obtained encodings to reconstruct the image; and causing the reconstructed image to be displayed on the XR display device.

36. A method according to claim 35, wherein the method is performed during video game streaming.

37. A method according to claim 35, wherein the method is performed during six degrees of freedom, 6DoF, storytelling with a pre-rendered zone of view.

38. A method according to claim 35, wherein the method is performed during design review of a digital twin and / or a digital double.

39. Apparatus configured to perform a method according to any of claims 1 to 38.

40. A bitstream comprising: a base layer encoding of an image comprising extended reality, XR, content and being for an XR display device; and an enhancement layer encoding of a first image region of the image, the enhancement layer encoding of a first image region having been generated in accordance with a first encoding priority, wherein either: the bitstream comprises an enhancement layer encoding of the second image region, the enhancement layer encoding of the second image region having been generated in accordance with a second, lower encoding priority; orthe bitstream does not comprise an enhancement layer encoding of the second image region.

Citation Information

Patent Citations

  • Decoder devices, methods and computer programs

    WO2018015764A1

  • Methods and apparatuses for encoding and decoding a bytestream

    WO2019111010A1

  • Processing of residuals in video coding

    WO2020188229A1

  • Low complexity enhancement video coding

    WO2020188273A1

  • Virtual reality panoramic video system using scalable video coding layers

    US20170347084A1