Enhancement layer processing of extended reality (XR) content

A tiered encoding and decoding system for XR content optimizes visual quality and resource usage by selectively applying enhancement layer processing based on the viewer's field of view, addressing performance and quality challenges in extended reality applications.

GB2636412APending Publication Date: 2025-06-18V NOVA INT LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2023019027
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing XR content delivery systems face challenges in providing high visual quality and efficient performance, particularly in augmented and virtual reality applications, due to factors like low visual quality and complex image processing that introduce latency and resource inefficiencies.

Method used

A tiered encoding and decoding system is employed, utilizing a base layer encoding and an enhancement layer encoding to enhance visual quality, where the enhancement layer processing is selectively applied based on the viewer's field of view, optimizing resource usage and reducing latency by minimizing unnecessary data transmission and processing.

Benefits of technology

This approach enhances visual quality within the viewer's field of view while reducing data transmission and processing overhead, improving user experience and performance in extended reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A base layer encoding of an image, 500, is obtained. The image comprises XR content and is for an XR display device. First enhancement layer processing is performed for a first image region, 506, of t
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field The present disclosure relates to enhancement layer processing of extended reality, XR, content. Background An image may be encoded to provide a base layer encoding and an enhancement layer encoding. The base layer encoding may be decoded and displayed without being enhanced by the enhancement layer. Alternatively, the enhancement layer encoding may be decoded and used to enhance the visual quality of the decoded base layer. Enhanced visual quality is particularly, but not exclusively, relevant in the context of XR, which includes augmented reality (AR) and virtual reality (VR). In particular, XR content is often viewed on a screen that is very close to the viewer. Low visual quality can significantly detract from the experience of the viewer. However, in addition to visual quality, other factors can also have a significant impact on performance and viewer experience in the context of XR. Summary Various aspects of the present disclosure are set out in the appended claims. Further features and advantages will become apparent from the following description of preferred embodiments, given by way of example only, which is made with reference to the accompanying drawings. Brief Description of the Drawings Figure 1 shows a schematic block diagram of an example of an image processing system; Figures 2A and 2B show a schematic block diagram of another example of an image processing system; Figure 3 shows a schematic diagram of an example of an image; Figure 4 shows a schematic diagram of an example of an enhancement layer; Figure 5 shows a schematic diagram of another example of an image; Figure 6 shows a schematic diagram of another example of an enhancement layer; Figure 7 shows a schematic diagram of another example of an enhancement layer; Figure 8 shows a schematic diagram of another example of an enhancement layer; Figure 9 shows a schematic diagram of another example of an enhancement layer; Figure 10 shows a schematic diagram of another example of an image; Figure 11 shows a schematic diagram of an example of a sequence of images; and Figure 12 shows a schematic block diagram of an example of an apparatus. Detailed Description Referring to Figure 1, there is shown an example of a signal processing system 100. The signal processing system 100 is used to process signals. Examples of types of signal include, but are not limited to, video signals, image signals, audio signals, volumetric signals such as those used in medical, scientific or holographic imaging, or other multidimensional signals. The signal processing system 100 includes a first apparatus 102 and a second apparatus 104. The first apparatus 102 and second apparatus 104 may have a clientserver relationship, with the first apparatus 102 performing the functions of a server device and the second apparatus 104 performing the functions of a client device. The signal processing system 100 may include at least one additional apparatus (not shown). The first apparatus 102 and / or second apparatus 104 may comprise one or more components. The one or more components may be implemented in hardware and / or software. The one or more components may be co-located or may be located remotely from each other in the signal processing system 100. Examples of types of apparatus include, but are not limited to, computerised devices, handheld or laptop computers, tablets, mobile devices, games consoles, smart televisions, set-top boxes, XR headsets (including AR and / or VR headsets) etc. The first apparatus 102 is communicatively coupled to the second apparatus 104 via a data communications network 106. Examples of the data communications network 106 include, but are not limited to, the Internet, a Local Area Network (LAN) and a Wide Area Network (WAN). The first and / or second apparatus 102, 104 may have a wired and / or wireless connection to the data communications network 106. In this example, the first apparatus 102 comprises an encoder 108. The encoder 108 is configured to encode data comprised in and / or derived based on the signal, which is referred to hereinafter as “signal data”. For example, where the signal is a video signal, the encoder 108 is configured to encode video data. Video data comprises a sequence of multiple images or frames. The encoder 108 may perform one or more further functions in addition to encoding signal data. The encoder 108 may be embodied in various different ways. For example, the encoder 108 may be embodied in hardware and / or software. The encoder 108 may encode metadata associated with the signal. The first apparatus 102 may use one or more than one encoder 108. Although in this example the first apparatus 102 comprises the encoder 108, in other examples the first apparatus 102 is separate from the encoder 108. In such examples, the first apparatus 102 is communicatively coupled to the encoder 108. The first apparatus 102 may be embodied as one or more software functions and / or hardware modules. In this example, the second apparatus 104 comprises a decoder 110. The decoder 110 is configured to decode signal data. The decoder 110 may perform one or more further functions in addition to decoding signal data. The decoder 110 may be embodied in various different ways. For example, the decoder 110 may be embodied in hardware and / or software. The decoder 110 may decoder metadata associated with the signal. The second apparatus 104 may use one or more than one decoder 110. Although in this example the second apparatus 104 comprises the decoder 110, in other examples, the second apparatus 104 is separate from the decoder 110. In such examples, the second apparatus 104 is communicatively coupled to the decoder 110. The second apparatus 104 may be embodied as one or more software functions and / or hardware modules. The encoder 108 encodes signal data and transmits the encoded signal data to the decoder 110 via the data communications network 106. The decoder 110 decodes the received, encoded signal data and generates decoded signal data. The decoder 110 may output the decoded signal data, or data derived using the decoded signal data. For example, the decoder 110 may output such data for display on one or more display devices associated with the second apparatus 104. The one or more display devices may be components of the second apparatus 104 or may otherwise be associated with the second apparatus 104. The one or more display devices may be operable to display XR content and may, therefore, be referred to as XR display devices. In some examples described herein, the encoder 108 transmits to the decoder 110 a representation of a signal at a given level of quality and information the decoder 110 can use to reconstruct a representation of some or all of the signal at one or more higher levels of quality. Such information may be referred to as “reconstruction data”. In some examples, “reconstruction” of a representation involves obtaining a representation that is not an exact replica of an original representation. The extent to which the representation is the same as the original representation may depend on various factors including, but not limited to, quantisation levels. A representation of a signal at a given level of quality may be considered to be a rendition, version or depiction of data comprised in the signal at the given level of quality. In some examples, the reconstruction data is included in the signal data that is encoded by the encoder 108 and transmitted to the decoder 110. For example, the reconstruction data may be in the form of metadata. In some examples, the reconstruction data is encoded and transmitted separately from the signal data. The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may comprise residual data, as described in more detail below. Residual data is an example of reconstruction data. The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may also comprise configuration data relating to processing of the residual data. The configuration data may indicate how the residual data has been processed by the encoder 108 and / or how the residual data is to be processed by the decoder 110. The configuration data may be signalled to the decoder 110, for example in the form of metadata. The first and / or second apparatuses 102, 104 may be configured to perform some or all of the techniques described herein. A computer program may be configured to perform some or all of the techniques described herein. Referring to Figures 2A and 2B, there is shown schematically an example of a signal processing system 200. The signal processing system 200 includes a first apparatus 202 and a second apparatus 204. In this example, the first apparatus 202 comprises an encoder and the second apparatus 204 comprises a decoder. However, as explained above, in other examples, the encoder is not comprised in the first apparatus 202 and / or the decoder is not comprised in the second apparatus 204. In each of the first apparatus 202 and the second apparatus 204, items are shown on two logical levels. The two levels are separated by a dashed line. Items on the first, highest level relate to data at a first level of quality. Items on the second, lowest level relate to data at a second level of quality. The first level of quality is higher than the second level of quality. The first and second levels of quality relate to a tiered hierarchy having multiple levels of quality. In some examples, the tiered hierarchy comprises more than two levels of quality. In such examples, the first apparatus 202 and the second apparatus 204 may include more than two different levels. There may be one or more other levels above and / or below those depicted in Figures 2A and 2B. As described herein, in certain cases, the levels of quality may correspond to different spatial resolutions. Referring first to Figure 2A, the first apparatus 202 obtains a first representation of an image at the first level of quality 206. A representation of a given image is a representation of data comprised in the image. The image may be a given frame of a video. The first representation of the image at the first level of quality 206 will be referred to as “input data” hereinafter as, in this example, it is data provided as an input to the encoder in the first apparatus 202. The first apparatus 202 may receive the input data 206. For example, the first apparatus 202 may receive the input data 206 from at least one other apparatus. The first apparatus 202 may be configured to receive successive portions of input data 206, e.g. successive frames of a video, and to perform the operations described herein to each successive frame. For example, a video may comprise frames Fi, F2, ... Ft and the first apparatus 202 may process each of these in turn. The first apparatus 202 derives data 212 based on the input data 206. In this example, the data 212 based on the input data 206 is a representation 212 of the image at the second, lower level of quality. In this example, the data 212 is derived by performing a downsampling operation on the input data 206 and will therefore be referred to as “downsampled data” hereinafter. In other examples, the data 212 is derived by performing an operation other than a downsampling operation on the input data 206, or the data 212 is the same as the input data 206 (i.e. the input data 206 is not processed, e g. downsampled). In this example, the downsampled data 212 is processed to generate processed data 213 at the second level of quality. In other examples, the downsampled data 212 is not processed at the second level of quality. As such, the first apparatus 202 may generate data at the second level of quality, where the data at the second level of quality comprises the downsampled data 212 or the processed data 213. In some examples, generating the processed data 213 involves the downsampled data 212 being encoded. Such encoding may occur within the first apparatus 202, or the first apparatus 202 may output the processed data 213 to an external encoder. Encoding the downsampled data 212 produces an encoded image at the second level of quality. The first apparatus 202 may output the encoded image, for example for transmission to the second apparatus 204. A series of encoded images, e.g. forming an encoded video, as output for transmission to the second apparatus 204 may be referred to as a “base” stream or “base” layer. As explained above, instead of being produced in the first apparatus 202, the encoded image may be produced by an encoder that is separate from the first apparatus 202. The encoded image may be part of an H.264 encoded video, or otherwise. Generating the processed data 213 may, for example, comprise generating successive frames of video as output by a separate encoder such as an H.264 video encoder. An intermediate set of data for the generation of the processed data 213 may comprise the output of such an encoder, as opposed to any intermediate data generated by the separate encoder. Generating the processed data 213 at the second level of quality may further involve decoding the encoded image at the second level of quality. The decoding operation may be performed to emulate a decoding operation at the second apparatus 204, as will become apparent below. Decoding the encoded image produces a decoded image at the second level of quality. In some examples, the first apparatus 202 decodes the encoded image at the second level of quality to produce the decoded image at the second level of quality. In other examples, the first apparatus 202 receives the decoded image at the second level of quality, for example from an encoder and / or decoder that is separate from the first apparatus 202. The encoded image may be decoded using an H.264 decoder. The decoding by a separate decoder may comprise inputting encoded video, such as an encoded data stream configured for transmission to a remote decoder, into a separate black-box decoder implemented together with the first apparatus 202 to generate successive decoded frames of video. Processed data 213 may thus comprise a frame of video data that is generated via a complex non-linear encoding and decoding process, where the encoding and decoding process may involve modelling spatiotemporal correlations as per a particular encoding standard such as H.264. However, because the output of any encoder is fed into a corresponding decoder, this complexity is effectively hidden from the first apparatus 202. In an example, generating the processed data 213 at the second level of quality further involves obtaining correction data based on a comparison between the downsampled data 212 and the decoded image obtained by the first apparatus 202, for example based on the difference between the downsampled data 212 and the decoded image. The correction data can be used to correct for encoder-decoder errors (which may also be referred to as “encode-decode errors”), namely errors introduced in encoding and decoding the downsampled data 212. In some examples, the first apparatus 202 outputs the correction data, for example for transmission to the second apparatus 204, as well as the encoded signal. This allows the recipient to correct for the encoder-decoder errors introduced in encoding and decoding the downsampled data 212. This correction data may also be referred to as a “first enhancement” stream. As the correction data may be based on the difference between the downsampled data 212 and the decoded image it may be seen as a form of residual data (e.g. that is different from the other set of residual data described later below). An item of residual data may be referred to as “a residual”. In some examples, generating the processed data 213 at the second level of quality further involves correcting the decoded image using the correction data. For example, the correction data as output for transmission may be placed into a form suitable for combination with the decoded image, and then added to the decoded image. This may be performed on a frame-by-frame basis. In other examples, rather than correcting the decoded image using the correction data, the first apparatus 202 uses the downsampled data 212. For example, in certain cases, just the encoded then decoded data may be used and in other cases, encoding and decoding may be replaced by other processing. In some examples, generating the processed data 213 involves performing one or more operations other than the encoding, decoding, obtaining, and correcting acts described above. The first apparatus 202 obtains data 214 based on the data at the second level of quality. As indicated above, the data at the second level of quality may comprise the processed data 213, or the downsampled data 212 where the downsampled data 212 is not processed at the lower level. As described above, in certain cases, the processed data 213 may comprise a reconstructed video stream (e.g. from an encoding-decoding operation) that is corrected using correction data. In the example of Figures 2A and 2B, the data 214 is a second representation of the image at the first level of quality, the first representation of the image at the first level of quality being the input data 206. The second representation at the first level of quality may be considered to be a preliminary or predicted representation of the image at the first level of quality. In this example, the first apparatus 202 derives the data 214 by performing an upsampling operation on the data at the second level of quality. The data 214 will be referred to hereinafter as “upsampled data”. However, in other examples one or more other operations could be used to derive the data 214, for example where data 212 is not derived by downsampling the input data 206. The input data 206 and the upsampled data 214 are used to obtain residual data 216. The residual data 216 is associated with the image. The residual data 216 may be in the form of a set of residual elements, which may be referred to as a “residual frame” or a “residual image”. A residual element may be referred to as “a residual”. A residual element in the set of residual elements 216 may be associated with a respective image element in the input data 206. An example of an image element is a pixel. In this example, a given residual element is obtained by subtracting a value of an image element in the upsampled data 214 from a value of a corresponding image element in the input data 206. As such, the residual data 216 is useable in combination with the upsampled data 214 to reconstruct the input data 206. The residual data 216 may also be referred to as “reconstruction data” or “enhancement data”. In one case, the residual data 216 may form part of a “second enhancement” stream. The residual data 216 may therefore result from upsampler-downsampler asymmetry. Upsampler-downsampler asymmetry may also be referred to as “upsamping-downsamping asymmetry”, “upsample-downsample” asymmetry or the like. The first apparatus 202 obtains configuration data relating to processing of the residual data 216. The configuration data indicates how the residual data 216 has been processed and / or generated by the first apparatus 202 and / or how the residual data 216 is to be processed by the second apparatus 204. The configuration data may comprise a set of configuration parameters. The configuration data may be useable to control how the second apparatus 204 processes data and / or reconstructs the input data 206 using the residual data 216. The configuration data may relate to one or more characteristics of the residual data 216. The configuration data may relate to one or more characteristics of the input data 206. Different configuration data may result in different processing being performed on and / or using the residual data 216. The configuration data is therefore useable to reconstruct the input data 206 using the residual data 216. As described below, in certain cases, configuration data may also relate to the correction data described herein. In this example, the first apparatus 202 transmits to the second apparatus 204 data based on the downsampled data 212, data based on the residual data 216, and the configuration data (or data based on the configuration data), to enable the second apparatus 204 to reconstruct the input data 206. Turning now to Figure 2B, the second apparatus 204 receives data 220 based on (e.g. derived from) the downsampled data 212. The second apparatus 204 also receives data based on the residual data 216. For example, the second apparatus 204 may receive a “base” stream (data 220), a “first enhancement stream” (any correction data) and a “second enhancement stream” (residual data 216). The base stream may be referred to as a “base layer”. The first and / or second enhancement stream may be referred to as, and / or may be comprised in, an “enhancement layer”. The second apparatus 204 also receives the configuration data relating to processing of the residual data 216. The data 220 based on the downsampled data 212 may be the downsampled data 212 itself, the processed data 213, or data derived from the downsampled data 212 or the processed data 213. The data based on the residual data 216 may be the residual data 216 itself, or data derived from the residual data 216. In some examples, the received data 220 comprises the processed data 213, which may comprise the encoded image at the second level of quality and / or the correction data. In some examples, for example where the first apparatus 202 has processed the downsampled data 212 to generate the processed data 213, the second apparatus 204 processes the received data 220 to generate processed data 222. Such processing by the second apparatus 204 may comprise decoding an encoded image (e.g. that forms part of a “base” encoded video stream) to produce a decoded image at the second level of quality. In some examples, the processing by the second apparatus 204 comprises correcting the decoded image using obtained correction data. Hence, the processed data 222 may comprise a frame of corrected data at the second level of quality. In some examples, the encoded image at the second level of quality is decoded by a decoder that is separate from the second apparatus 204. The encoded image at the second level of quality may be decoded using an H.264 decoder. In other examples, the received data 220 comprises the downsampled data 212 and does not comprise the processed data 213. In some such examples, the second apparatus 204 does not process the received data 220 to generate processed data 222. The second apparatus 204 uses data at the second level of quality to derive the upsampled data 214. As indicated above, the data at the second level of quality may comprise the processed data 222, or the received data 220 where the second apparatus 204 does not process the received data 220 at the second level of quality. The upsampled data 214 is a preliminary representation of the image at the first level of quality. The upsampled data 214 may be derived by performing an upsampling operation on the data at the second level of quality. The second apparatus 204 obtains the residual data 216. The residual data 216 is useable with the upsampled data 214 to reconstruct the input data 206. The residual data 216 is indicative of a comparison between the input data 206 and the upsampled data 214. The second apparatus 204 also obtains the configuration data related to processing of the residual data 216. The configuration data is useable by the second apparatus 204 to reconstruct the input data 206. For example, the configuration data may indicate a characteristic or property relating to the residual data 216 that affects how the residual data 216 is to be used and / or processed, or whether the residual data 216 is to be used at all. In some examples, the configuration data comprises the residual data 216. There are several considerations relating to such processing. One such consideration is the amount of information that is generated, stored, transmitted and / or processed. The more information that is used, the greater the amount of resources that may be involved in handling such information. Examples of such resources include transmission resources, storage resources and processing resources. Some signal processing techniques allow a relatively small amount of information to be used. This may reduce the amount of data transmitted via the data communications network 106. The savings may be particularly relevant where the data relates to high quality video data, where the amount of information transmitted can be especially high. Another consideration is latency. Complex image processing may introduce latency, which may negatively impact performance. Other considerations include the ability of the decoder to perform image reconstruction accurately, reliably, and / or efficiently. Performing image reconstruction accurately and reliably may affect the ultimate visual quality of the displayed image and consequently may affect a viewer’s engagement with the image and / or with a video comprising the image. This can be especially relevant to XR. Efficient reconstruction is especially effective for mobile computing devices, which may readily be used in XR applications. Referring to Figure 3, there is shown a schematic diagram of an example of an image 300. The image 300 may be a standalone image or may be a frame in video content. The image 300 may, for example, represent a frame of a movie, a television programme, a sporting event, and so on. Base layer processing may be performed in relation to the image 300. The base layer processing results in a base layer encoding of the image 300. The base layer encoding may correspond to the processed data 213 described above with reference to Figure 2A, for example. The base layer encoding may be obtained in various different ways. In some examples, the base layer encoding is obtained by generating the base layer encoding. In some examples, the base layer encoding is obtained by instructing generation of the base layer encoding. In some examples, the base layer encoding is obtained by configuring an encoder such that the encoder generates the base layer encoding. In some examples, the base layer encoding is obtained by receiving the base layer encoding. In some examples, the base layer encoding is obtained by retrieving the base layer encoding from storage. The base layer encoding may be obtained in another manner in other examples. Enhancement layer processing may be performed in relation to the image 300. The enhancement layer processing results in an enhancement layer encoding of the image 300. The enhancement layer encoding may correspond to the correction data and / or the residual data 216 described above with reference to Figure 2A, for example. The enhancement layer encoding may be obtained in various different ways. In some examples, the enhancement layer encoding is obtained by generating the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by instructing generation of the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by configuring an encoder such that the encoder generates the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by receiving the enhancement layer encoding. In some examples, the enhancement layer encoding is obtained by retrieving the enhancement layer encoding from storage. The enhancement layer encoding may be obtained in another manner in other examples. More generally, references herein to obtaining an encoding encompass generating the encoding, instructing generation of the encoding, configuring an encoder to enable the encoder to generate the encoding, receiving the encoding, retrieving the encoding from storage, and obtaining the encoding in any other manner. A base layer encoding may be generated by a base layer encoder. An enhancement layer encoding may be generated by an enhancement layer encoder. The base layer encoder may be different from the enhancement layer encoder. In particular, the base and enhancement layer encoders may use different codecs. In some examples, the enhancement layer residuals result from upsampler-downsampler asymmetry. Such residuals may correspond to the residual data 216 described above with reference to Figure 2A. In some examples, the enhancement layer residuals result from encoder-decoder errors. Such residuals may correspond to the correction data described above with reference to Figure 2A. In some examples, the base and enhancement layer encodings are output. In some examples, the base and enhancement layer encodings are stored. The base and enhancement layer encodings may be decoded, combined and displayed, such as described above with reference to Figure 2B. Referring to Figure 4, there is shown a schematic diagram of an example of an enhancement layer 400. In this example, the enhancement layer 400 is generated based on the example image 300 shown in Figure 3. In this example, the enhancement layer 400 comprises a set of residuals 402. The residuals 402 may take various different forms. Residuals are more generally referred to as “correction data” or “enhancement data”. In some examples, the residuals 402 comprise Low Complexity Enhancement Video Coding (LCEVC) residuals, which will be described in more detail below. In some examples, the residuals 402 comprise residuals used in scalable codecs, such as SMPTE ST 2117-1 (often known as “VC-6”). In some examples, the residuals 402 comprise residuals used in scalable MPEG codecs such as scalable AVC (SVC), scalable HEVC (SHVC) and scalable VVC. There are no restrictions on where the residuals 402 may be positioned in the enhancement layer 400. As explained above, the enhancement layer 400 may be used to enhance the visual quality of a base layer encoding of the example image 300. In particular, the residuals 402 enhance the visual quality at their respective positions. As such, visual quality enhancement can readily be provided across the whole of the image 300. In many practical situations, the entire image 300 is visible to a viewer. For example, where the image 300 is a frame from a movie, a sporting event, a television programme, or the like, the entire image 300 is visible to the viewer. The viewer may be looking at any part of the image 300. In addition, the viewer may focus on different parts of a video at different times. For example, the viewer may focus on an object in the centre of the image 300 during one part of the video, may focus on an object in the top right comer of the image 300 during a later part of the video, and so on. Different viewers may focus on different parts of any given image 300. For example, some viewers may focus on an object at the bottom left comer of the image 300 and other users may focus on a different object at the top right corner of the image 300. High visual quality across the whole image can therefore improve viewer experience and perception of visual quality. For example, by default, residuals may be applied anywhere and everywhere across the image. Referring to Figure 5, there is shown a schematic diagram of another example of an image 500. Various items are shown, schematically, as being part of the image 500. However, this is solely to facilitate an understanding of the present disclosure. In practice, the items may not be visible in the image 500 itself and the image 500 may represent other content. In this example, the image 500 comprises XR content and is for an XR display device. As indicated above, XR content may comprise AR content and / or VR content. The XR display device may correspond to the second apparatus 104 described above with reference to Figure 1, for example. As will be described above, base and enhancement layer encodings generated using the image 500 may be output for display on the XR display device. In this example, the image 500 is a stereoscopic image. In this example, the image comprises a left-side image portion 502 and a right-side image portion 504. In this specific example, the image 500 is rectangular, with width, w, and height, h. In this specific example, the left-side and right-side image portions 502, 504 are also rectangular with width, w / 2, and height, h. However, the image 500 may have a different shape in other examples. Additionally, in other examples, the left-side and right-side image portions 502, 504 may have different proportions. In some examples, the image 500 comprises a different number of image portions. The example image 500 shown in Figure 5 comprises a first image region 506. In this example, the first image region 502 comprises first and second sub-regions 506-1, 506-2. In this example, the first sub-region 506-1 is in the left-side image portion 502 and the second sub-region 506-2 is in the right-side image portion 504. In this example, the first sub-region 506-1 is centred in the centre of the leftside image portion 502 and the second sub-region 506-2 is centred in the centre of the right-side image portion 504. In particular, neither the first sub-region 506-1 nor the second sub-region 506-2 is centred in the centre of the image 500 itself. In this example, the first sub-region 506-1 is rectangular with width, 14^, and height, where W11 <w / 2 and <h. In this example, the second sub-region 506-2 is rectangular with width, W12, and height, H12, where 2 <w / 2 and H12 <h. In this example, W11 = l¥12 and H11 = H12. The first and / or second sub-region 506-1, 506-2 may have different shapes and / or dimensions in other examples. In this example, the first and second sub-regions 506-1, 506-2 are nonoverlapping. The first region 506 may therefore be said to be non-contiguous. The example image 500 shown in Figure 5 also comprises a second image region 508. In this example, the second image region 508 surrounds the first image region 506. In this example, the second image region 508 is the complement of the first image region 506 in the image 500. In this example, the first image region 506 corresponds to an expected field of view (FOV) of a viewer of the image 500. An expected FOV is the FOV the viewer is expected to have. The expected FOV may differ from an actual FOV of the viewer. For example, last-minute eye movements may change the actual FOV of the viewer compared to what is expected, for example based on a recent FOV of the viewer. The expected FOV may also be referred to as a “probable”, “predicted”, “central” or “focal” FOV. In this example, the viewer of the image 500 is a user of the XR device. More specifically, in this example, the first sub-region 506-1 corresponds to a left-eye view of the viewer of the image 500 and the second sub-region 506-2 corresponds to a right-eye view of the viewer of the image 500. In this example, some or all of the second image region 508 is outside the FOV of the viewer of the image 500. A boundary 510 exists between the left-side image portion 502 and the rightside image portion 504. In this example, the second image region 508 includes the boundary 510. The boundary 510 may appear in the image 500 as a dark line. For example, the left-eye view (corresponding to the left-side image portion 502) may transition from a relatively light colour at its left side to a relatively dark colour at its right side. The right-eye view (corresponding to the right-side image portion 504) may similarly transition from a relatively light colour at its left side to a relatively dark colour at its right side. The boundary 510 between the left-side and right-side image portions 502, 504 would therefore appear as a dark vertical line. In some examples, the second image region 508 includes at least part of an edge 512 of the image 500. The edge 512 may be referred to as a “boundary”, “perimeter”, or the like. The edge may similarly include a dark line or region, especially in the case of XR content. In some examples, the positions of the first and second image regions 506, 508 are independent of pixel values of pixels of the image 500 and so may be said to be pixel-value-independent. Such examples therefore differ from examples in which pixel analysis is performed to determine the positions of the first and second image regions 506, 508. In some examples, the positions of the first and second image regions 506, 508 are independent of a genre of the XR content of the image 500 and so may be said to be genre-independent. Such examples therefore differ from examples in which knowledge of the genre of the XR content influences the positions of the first and second image regions 506, 508. Therefore, in some examples, the positions of the first and second image regions 506, 508 are content-independent. Instead of the positions of the first and second image regions 506, 508 being independent of pixel values and / or XR content genre, pre-analysis of the image 500 could be performed to determine the positions of the first and / or second image regions 506, 508. However, such pre-analysis introduces latency. For XR content, reduced latency generally improves performance and user experience. Referring to Figure 6, there is shown a schematic diagram of another example of an enhancement layer 600. In this example, the enhancement layer 600 is generated based on the example image 500 shown in Figure 5. Reference signs used in Figure 6 are the same as those used in Figure 5 for the same or similar features, but incremented by 100. The example enhancement layer 600 shown in Figure 6 corresponds generally to the example enhancement layer 400 shown in Figure 4 in that the residuals in the set of residuals are in corresponding positions in Figure 6. However, Figure 6 illustrates that a first subset 614 of the residuals (represented by residual 614) are in a first region 606 of the enhancement layer 600 and that a second subset 616 of the residuals (represented by residual 616) are in a second region 608 of the enhancement layer 600. The first and second regions 606, 608 of the enhancement layer 600 correspond to the first and second image regions 506, 508 of the image 500 in terms of their positions. Residuals in the first region 606 of the enhancement layer 600 may therefore be said to correspond to the first image region 506 of the image 500, and residuals in the second region 608 of the enhancement layer 600 may therefore be said to correspond to the second image region 508 of the image 500. As explained above, for many types of image, the whole image may be visible to the viewer. However, an entire image may not be visible, for example for an XR image. Examples that will now be described can improve rate control, performance, and / or visual quality in XR. Such examples selectively enhance XR content. For example, part of the XR content may be enhanced and another part of the XR content may be enhanced to a lesser degree, or even not enhanced at all. Such selective enhancement is particularly, but not exclusively, effective for XR content where not all of the XR content provided to the XR display device is within the FOV of the user of the XR display device. XR content outside the FOV of the user can be enhanced less than XR content within the FOV of the user or may not even be enhanced at all. In some examples, the XR display device can still display XR content outside of the FOV of the user based on a base layer encoding of such content, even if a corresponding enhancement layer encoding of such content is not available. Lower visual quality of XR content within the FOV impacts user experience more significantly than outside the FOV. Visual quality outside the FOV may therefore be sacrificed, at least to some degree, relative to visual quality inside the FOV, without significantly negatively affecting performance and user experience. This can improve rate control as less enhancement layer data can be communicated to the XR display device. In some examples, some or all of the saving made by reducing the enhancement outside the FOV is used to enhance visual quality within the FOV even further. In more detail, in a possible implementation, residuals are applied to an object for some frames but not for others. This can improve rate control but can impact visual quality in relation to the object in question. Further, as explained above, a stereoscopic image can have an artificially created sharp border. Significant amounts of residuals may be spent trying to make this border a perfect line. However, a precisely defined border line does not improve visual quality when the stereoscopic image is viewed on an XR display device because the border is outside the FOV of the viewer. This applies correspondingly to dark lines or regions at the edge of an image. Examples described below address this by not generating residuals, or by generating lower-priority residuals, outside the FOV of the viewer. The border would fall outside of the FOV of the viewer and so no, or smaller, residuals would be spent on the border. In particular, in some examples, only residuals for regions of an XR image in the FOV of the viewer are generated, encoded, quantised to non-zero values and / or sent. Such regions may be in the centre of each half of the stereoscopic image. Referring to Figure 7, there is shown a schematic diagram of another example of an enhancement layer 700. The example enhancement layer 700 shown in Figure 7 corresponds generally to the example enhancement layer 600 shown in Figure 6. Reference signs used in Figure 7 are the same as those used in Figure 6 for the same or similar features, but incremented by 100. However, Figure 7 illustrates that the first subset 714 of residuals is processed differently from the second subset 716 of residuals. This is depicted in Figure 7 by the different shading of the first and second subsets 714, 716 of residuals. In more detail, a base layer encoding of the example image 500 shown in Figure 5 is obtained. In some examples, the base layer encoding of the image 500 has a consistent encoding quality across the image. In other examples, the base layer encoding of the image 500 comprises a first base layer encoding portion corresponding to the first image region 506, and a second base layer encoding portion corresponding to the second image region 508. The first base layer encoding portion has a first encoding quality, and the second base layer encoding portion has a second, lower encoding quality. A lower encoding quality may result in a lower visual quality when the encoding is decoded. First enhancement layer processing may be performed for the first image region 506 of the image 500. Performing first enhancement layer processing comprises obtaining an enhancement layer 700 encoding of the first image region 506. The enhancement layer 700 encoding of the first image region 506 is in accordance with a first encoding priority. In this example, the first subset 714 of residuals, which are in the first region 706 of the enhancement layer 700, are encoded in accordance with the first encoding priority. An encoding priority may correspond to a visual importance. For example, a higher encoding priority may correspond to a higher visual importance and a lower encoding priority correspond to a lower visual importance. A higher encoding priority may correspond to higher prioritisation of an encoding relative to another encoding having a lower encoding priority. A higher encoding priority may correspond to a lower quantisation step width of an encoding relative to another encoding having a lower encoding priority. A higher encoding priority may correspond to use of a codec having a higher visual quality for an encoding relative to another encoding having a lower encoding priority. Second enhancement layer processing is performed for the second image region 508 of the image 500. In some examples, performing second enhancement layer processing comprises inhibiting an enhancement layer 700 encoding of the second image region 508 from being obtained. As such, although the second subset 716 of residuals, which are in the second region 708 of the enhancement layer 700, are depicted in Figure 7, this is solely to facilitate understanding. The second subset 716 of residuals may not have been generated at all. Inhibiting the enhancement layer 700 encoding of the second image region 508 from being obtained may be implemented as an encoding condition. The encoding condition may be that residual encoding logic is not run for residuals with specific positions. This differs from a priority map implementation, which will be described in more detail below, in which residuals are generated for all positions, but in which residuals in specific positions are de-prioritised and quantised down to zero. Reference is made, in this connection to WO 2020 / 188229, the contents of which are incorporated herein by reference. WO 2020 / 188229 describes, for example, an encoder analysing an input video and preparing a set of residual masks for each frame of the video. WO 2020 / 188229 also describes that, instead of the encoder performing such pre-analysis of an input video, a central server can propose a residual weighted mask according to a type of input video. The central server may provide a set of residual weighted masks covering different genres, for example sports, movies, news, etc. The residual masks may thereby prioritise the area of the picture in which detail is required. However, such pre-analysis and / or genre-dependence can introduce processing latency. The techniques described herein also differ from WO 2018 / 015764, the contents of which are incorporated herein by reference. In WO 2018 / 015764, one or more enhancement layers may be provided to a decoder for an entire image and the decoder may select which portions of the enhancement layer(s) are to be used. In some examples, inhibiting the enhancement layer 700 encoding of the second image region 508 from being obtained comprises causing an enhancement layer mask to be applied to the second image region 508. In such examples, the enhancement layer mask may correspond, in shape, to the second image region 508. In such examples, the enhancement layer mask is applied to one or more lower-priority image regions such that one or more higher-priority image regions are unmasked and outside the enhancement layer mask. Residuals corresponding to the masked, lower-priority image region(s) and the unmasked, higher-priority image region(s) can be processed accordingly. In other examples, inhibiting the enhancement layer 700 encoding of the second image region 508 from being obtained comprises causing an enhancement layer mask to be applied to the first image region 506. In such examples, the enhancement layer mask may correspond, in shape, to the first image region 506. In such examples, the enhancement layer mask is applied to one or more higher-priority image regions such that one or more lower-priority image regions are unmasked and outside the enhancement layer mask. Residuals corresponding to the masked, higher-priority image region(s) and the unmasked, lower-priority image region(s) can be processed accordingly. In some examples, the enhancement layer mask has a geometric shape. For example, the enhancement layer mask may be square, circular, oval, or may have a more complex shape. Less complex shapes may have improved computation performance relative to more complex shapes. However, less complex shapes may resemble the FOV less accurately than more complex shapes. In other examples, performing second enhancement layer processing comprises obtaining an enhancement layer 700 encoding of the second image region 508. In such examples, the enhancement layer 700 encoding of the second image region 508 is in accordance with a second encoding priority. The second encoding priority is lower than the first encoding priority. In some examples, obtaining the enhancement layer 700 encoding of the second image region 508 comprises obtaining a set of residuals 716 corresponding to the second image region 508, where the set of residuals 716 corresponding to the second image region 508 are quantised to zero. In this example, the set of residuals 716 corresponding to the second image region 508 is the second subset 716 of residuals. As such, the second subset 716 of residuals may be generated, and their values may be set to zero during quantisation. In some examples, obtaining the enhancement layer 700 encoding of the first image region 506 comprises obtaining a set of residuals 714 corresponding to the first image region 506, where the set of residuals 714 corresponding to the first image region 506 are quantised with a first quantisation step width. In this example, the set of residuals 714 corresponding to the first image region 506 is the first subset 714 of residuals. In such examples, generating the enhancement layer 700 encoding of the second image region 508 comprises obtaining a set of residuals 716 corresponding to the second image region 508, where the set of residuals 716 corresponding to the second image region 508 are quantised with a second, higher quantisation step width. In this example, the set of residuals 716 corresponding to the second image region 508 is the second subset 716 of residuals. Without performing the above-described second enhancement layer processing, the second subset 716 of residuals would, in effect, be wasted. This is because they are not important, or at least have limited importance, in terms of visual quality as perceived by the viewer. The first and second enhancement layer processing may be performed as part of a single enhancement layer processing procedure. For example, the first and second enhancement layer encodings may be generated by the same enhancement layer encoder as each other. In some examples, the encodings are output for display on the XR display device. In some examples, the encodings are stored. The encodings comprise the base layer encoding and the enhancement layer 700 encoding of the first image region 506. The encodings may comprise the enhancement layer 700 encoding of the second image region 508 when this is obtained. The encodings may be output to the XR display device for display on the XR display device or may be output to another device for display on the XR display device. An example of such another device is a computing device that is paired with an XR display device in the form of an XR headset. In some examples, the base layer encoding of the image 500 is decodable and / or displayable independently of the enhancement layer 700 encoding(s). In terms of decoder-side processing, a base layer encoding of an image 500 may be obtained. An enhancement layer 700 encoding of a first image region 506 (the encoding being in accordance with the first encoding priority) may also be obtained. An enhancement layer 700 encoding of a second image region 508 (the encoding being in accordance with the second encoding priority) may be obtained, or an enhancement layer 700 encoding of the second image region 508 may not be obtained. Such obtaining, or not obtaining, may be by the XR display device. The obtained encodings may be used to reconstruct the image 500. The reconstructed image may be caused to be displayed on the XR display device. In some examples, the residuals 714, 716 have intra-independency and so may be said to be “intra-independent”. Thus, there may be no “intra” dependency for such residuals 714, 716. This enables the use of no residuals, or residuals quantised to zero, for a portion of the enhancement layer 700. A codec, such as LCEVC, that does not perform intra coding may be used to generate such intra-independent residuals. LCEVC is described in WO 2020 / 188273 (PCT / GB2020 / 050695) and WO 2019 / 111010 (PCT / GB2018 / 053552), the entire contents of which are incorporated herein by reference. If the residuals 714, 716 had intra-dependency, then decoding values of some of the residuals may require values of other residuals. For example, decoding values of some of the residuals 714 in the first region 706 of the enhancement layer 700 may require values of some of the residuals 716 in the second region 708 of the enhancement layer 700. If the residuals 716 in the second region 708 of the enhancement layer 700 were not provided at all, or were quantised heavily, the enhancement layer 700 might not be useable or at least might not be as effective. For example, if the residuals 716 in the second region 708 of the enhancement layer 700 were heavily quantised, then intra prediction would be less accurate but may still provide a useable prediction. If, however, the residuals 716 in the second region 708 of the enhancement layer 700 were completely omitted or all quantised to zero, then they would not be useable for intra prediction. Referring to Figure 8, there is shown a schematic diagram of another example of an enhancement layer 800. The example enhancement layer 800 shown in Figure 8 corresponds generally to the example enhancement layer 700 shown in Figure 7. Reference signs used in Figure 8 are the same as those used in Figure 7 for the same or similar features, but incremented by 100. However, Figure 8 only depicts the first subset 814 of residuals. It can be seen from Figure 8 that fewer residuals can be used than in the example enhancement layer 700 shown in Figure 7, without negatively impacting visual quality, either significantly or at all. Referring to Figure 9, there is shown a schematic diagram of another example of an enhancement layer 900. The example enhancement layer 900 shown in Figure 9 corresponds generally to the example enhancement layers 700, 800 shown in Figures 7 and 8. Reference signs used in Figure 9 are the same as those used in Figures 7 and 8 for the same or similar features, but incremented by 200 and 100 respectively. However, in this example, residuals that would have been generated outside of the FOV of the viewer are, in effect, redistributed to be inside the FOV of the viewer. In more detail, in some examples, a data-saving measure is determined. The data-saving measure represents an amount of data saved by performing the second enhancement layer processing for the second image region 508, for example compared to performing the first enhancement layer processing for the second image region 508. Some or all of the data saved in this manner may be allocated to the first enhancement layer processing for the first image region 506. In particular, Figure 7 shows six residuals 714 in the first region 706 of the enhancement layer 700 and seven residuals 716 in the second region 708 of the enhancement layer 700, i.e. thirteen residuals in total. In contrast, Figure 9 shows thirteen residuals 918 in the first region 906 of the enhancement layer 900 and no residuals in the second region 908 of the enhancement layer 900, i.e. still thirteen residuals in total. By limiting the region(s) of the enhancement layer in which residuals are allowed, or at least residuals with limited quantisation are allowed, visual quality can be increased for the most important region(s) of an image, potentially without increasing the amount of data transmitted compared to where residuals are allowed to be in any part of the enhancement layer. Referring to Figure 10, there is shown a schematic diagram of another example of an image 1000. The example image 1000 shown in Figure 10 corresponds generally to the example image 500 shown in Figure 5. Reference signs used in Figure 10 are the same as those used in Figure 5 for the same or similar features, but incremented by 500. However, in this example, the image 1000 comprises a third image region 1014. In this example, the third image region 1014 comprises first and second subregions 1014-1, 1014-2. In this example, the first sub-region 1014-1 is in the left-side image portion 1002 and the second sub-region 1014-2 is in the right-side image portion 1004. In this example, the first sub-region 1014-1 is centred in the centre of the leftside image portion 1002 and the second sub-region 1014-2 is centred in the centre of the right-side image portion 1004. The first sub-region 1014-1 may therefore be said to be central to the left-side image portion 1002 and the second sub-region 1014-2 may be said to be central to the right-side image portion 1004. In this example, the first sub-region 1014-1 is rectangular with width, 1¥3j1, and height, H31, where W±1 <W31 <w / 2 and <H31 <h. In this example, the second sub-region 1014-2 is rectangular with width W32, and height, H32, where V / 12 <VF3 2 <w / 2 and H±1 <H32 <h. In this example, W31 = J¥3 2 and H31 = H3 2. The first and / or second sub-region 1014-1, 1014-2 may have a different shape and / or dimensions in other examples. In this example, the first and second sub-regions 1014-1, 1014-2 of the third image region 1014 are non-overlapping. The third image region 1014 may therefore be said to be non-contiguous. In some examples, third enhancement layer processing is performed for the third image region 1014. Performing third enhancement layer processing may comprise obtaining an enhancement layer encoding of the third image region 1014 in accordance with a third encoding priority. The third encoding priority may be different from the first and second encoding priorities described above with reference to Figure 7. In some examples, the third image region 1014 is between the first and second image regions 1006, 1008. In some examples, the third encoding priority is between the first and second encoding priorities. The third image region 1014 may correspond to a peripheral region of the FOV of the user of the XR display device. Therefore, more than two encoding priorities may be used. Some residuals may not be used at all. In other words, such residuals may not be generated at all or may be generated but quantised to zero. These would be the lowest priority residuals. Residuals in the peripheral region may have intermediate priority and may be encoded, but heavily quantised. Residuals in the central field of view region may have the highest priority. Referring to Figure 11, there is shown a schematic diagram of an example of a sequence of images 1100. Some reference signs used in Figure 11 are the same as those used in Figure 5 for the same or similar features. In this example, base layer and enhancement layer encodings of additional images are created, for example for output for display on the XR display device. In this example, there are three additional images 1102, 1104, 1106. However, in practice, there may be significantly more than three additional images in the sequence of images 1100. The additional images 1102, 1104, 1106 are additional to the example image 500 shown in Figure 5 and reproduced in Figure 11. The image 500 and the additional images 1102, 1104, 1106 are part of the sequence of images 1100 comprising XR content for the XR display device. The sequence of images 1100 may correspond to an XR stream. The XR stream may comprise a stereoscopic XR stream. In some examples, the first and second image regions (506, 508 for the image 500) have fixed positions across the sequence of images 1110. This can reduce, or even avoid, using resources to determine their positions across the sequence of images 1110. Fixed positions may also be referred to as “static”, “set”, or “set frame” positions. Using a fixed priority map may therefore reduce or avoid positioning-related calculation costs. In some examples, the first and second image regions (506, 508 for the image 500) have dynamic positions across the sequence of images 1110. The dynamic positions may be determined in various different ways. In some examples, eye-tracking data is received from the XR display device, and the positions of the first and second image regions (506, 508 for the image 500) are determined based on the eye-tracking data. Thus, an encoder may receive feedback from a XR display device such that the first and second image regions (506, 508 for the image 500) have positions linked to eye-tracking. Referring to Figure 12, there is shown a schematic block diagram of an example of an apparatus 1200. In an example, the apparatus 1200 comprises an encoder. In another example, the apparatus 1200 comprises a decoder. In other examples, the apparatus 1200 comprises neither an encoder nor a decoder but is configured to communicate with an encoder and / or a decoder. Examples of apparatus 1200 include, but are not limited to, a mobile computer, a personal computer system, a wireless device, base station, phone device, desktop computer, laptop, notebook, netbook computer, mainframe computer system, handheld computer, workstation, network computer, application server, storage device, a consumer electronics device such as a camera, camcorder, mobile device, video game console, handheld video game device, an XR headset, or in general any type of computing or electronic device. In this example, the apparatus 1200 comprises one or more processors 1201 configured to process information and / or instructions. The one or more processors 1201 may comprise a central processing unit (CPU). The one or more processors 1201 are coupled with a bus 1202. Operations performed by the one or more processors 1201 may be carried out by hardware and / or software. The one or more processors 1201 may comprise multiple co-located processors or multiple disparately located processors. In this example, the apparatus 1200 comprises computer-useable volatile memory 1203 configured to store information and / or instructions for the one or more processors 1201. The computer-useable volatile memory 1203 is coupled with the bus 1202. The computer-useable volatile memory 1203 may comprise random access memory (RAM). In this example, the apparatus 1200 comprises computer-useable non-volatile memory 1204 configured to store information and / or instructions for the one or more processors 1201. The computer-useable non-volatile memory 1204 is coupled with the bus 1202. The computer-useable non-volatile memory 1204 may comprise read-only memory (ROM). In this example, the apparatus 1200 comprises one or more data-storage units 1205 configured to store information and / or instructions. The one or more data-storage units 1205 are coupled with the bus 1202. The one or more data-storage units 1205 may for example comprise a magnetic or optical disk and disk drive or a solid-state drive (SSD). In this example, the apparatus 1200 comprises one or more input / output (I / O) devices 1206 configured to communicate information to and / or from the one or more processors 1201. The one or more I / O devices 1206 are coupled with the bus 1202. The one or more I / O devices 1206 may comprise at least one network interface. The at least one network interface may enable the apparatus 1200 to communicate via one or more data communications networks. Examples of data communications networks include, but are not limited to, the Internet and a Local Area Network (LAN). The one or more EO devices 1206 may enable a user to provide input to the apparatus 1200 via one or more input devices (not shown). The one or more input devices may include for example a remote control, one or more physical buttons etc. The one or more I / O devices 1206 may enable information to be provided to a user via one or more output devices (not shown). The one or more output devices may for example include a display screen. Various other entities are depicted for the apparatus 1200. For example, when present, an operating system 1207, image processing module 1208, one or more further modules 1209, and data 1210 are shown as residing in one, or a combination, of the computer-usable volatile memory 1203, computer-usable non-volatile memory 1204 and the one or more data-storage units 1205. The data signal processing module 1208 may be implemented by way of computer program code stored in memory locations within the computer-usable non-volatile memory 1204, computer-readable storage media within the one or more data-storage units 1205 and / or other tangible computer-readable storage media. Examples of tangible computer-readable storage media include, but are not limited to, an optical medium (e g., CD-ROM, DVD-ROM or Blu-ray), flash memory card, floppy or hard disk or any other medium capable of storing computer-readable instructions such as firmware or microcode in at least one ROM or RAM or Programmable ROM (PROM) chips or as an Application Specific Integrated Circuit (ASIC). The apparatus 1200 may therefore comprise a data signal processing module 1208 which can be executed by the one or more processors 1201. The data signal processing module 1208 can be configured to include instructions to implement at least some of the operations described herein. During operation, the one or more processors 1201 launch, run, execute, interpret or otherwise perform the instructions in the signal processing module 1208. Although at least some aspects of the examples described herein with reference to the drawings comprise computer processes performed in processing systems or processors, examples described herein also extend to computer programs, for example computer programs on or in a carrier, adapted for putting the examples into practice. The carrier may be any entity or device capable of carrying the program. It will be appreciated that the apparatus 1200 may comprise more, fewer and / or different components from those depicted in Figure 12. The apparatus 1200 may be located in a single location or may be distributed in multiple locations. Such locations may be local or remote. The techniques described herein may be implemented in software or hardware, or may be implemented using a combination of software and hardware. They may include configuring an apparatus to carry out and / or support any or all of techniques described herein. A bitstream may be provided. The bitstream may comprise a base layer encoding of an image comprising XR content and being for an XR display device. The bitstream may comprise an enhancement layer encoding of a first image region of the image. The enhancement layer encoding of the first image region may have been generated in accordance with a first encoding priority. In some examples, the bitstream comprises an enhancement layer encoding of the second image region, where enhancement layer encoding of the second image region has been generated in accordance with a second, lower encoding priority. In other examples, the bitstream does not comprise an enhancement layer encoding of the second image region. In examples described above, the image is a stereoscopic image. However, other types of image having left and right image portions may be used. For example, one of the left and right image portions may comprise a view of a scene and the other of the left and right image portions may comprise a corresponding depth map. It may be acceptable to sacrifice some quality of the depth map in favour of the view of the scene, for instance. In another example, one of the left and right image portions may comprise a view of a sporting event from a first angle, and the other of the left and right image portions may comprise a view of the sporting event from a second, different angle. It may be acceptable to sacrifice some quality of one of the views of the sporting event (for example a zoomed-out view) in favour of the other view of the sporting event. Instead of the example techniques described above, a base layer encoding of an image may be obtained in which a first image region of the image is not distorted and in which a second image region of the image is distorted. Such distortion may involve removing some pixel values in the second image region. A decoder may then interpolate the remaining pixel values in the second image region to attempt to reconstruct the removed pixel values. Imperfections, resulting from the interpolation, may be tolerable in the second image region where the second image region is outside the FOV of a viewer. However, such an interpolation-based technique can introduce latency, which can negatively impact performance in the context of XR It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

Claims

1. A method of processing an image, the image comprising extended reality, XR, content and being for an XR display device, the method comprising:obtaining a base layer encoding of the image;performing first enhancement layer processing for a first image region of the image, wherein performing first enhancement layer processing comprises obtaining an enhancement layer encoding of the first image region, the enhancement layer encoding of the first image region being in accordance with a first encoding priority; andperforming second enhancement layer processing for a second image region of the image, wherein performing second enhancement layer processing comprises:inhibiting an enhancement layer encoding of the second image regionfrom being obtained; orobtaining an enhancement layer encoding of the second image region, the enhancement layer encoding of the second image region being in accordance with a second, lower encoding priority.

2. A method according to claim 1, wherein positions of the first and second image regions are independent of pixel values of pixels of the image.

3. A method according to claim 1 or 2, comprising:outputting the encodings for display on the XR display device.

4. A method according to any of claims 1 to 3, comprising:outputting, for display on the XR display device, base layer and enhancement layer encodings of additional images, the image and the additional images being part of a sequence of images comprising XR content for the XR display device,wherein the first and second image regions have fixed positions across the sequence of images.

5. A method according to any of claims 1 to 3, comprising:outputting, for display on the XR display device, base layer and enhancement layer encodings of additional images, the image and the additional images being part of a sequence of images comprising XR content for the XR display device,wherein the first and second image regions have dynamic positions across the sequence of images.

6. A method according to claim 5, comprising:receiving eye-tracking data from the XR display device; anddetermining the positions of the first and second image regions based on the eye-tracking data.

7. A method according to any of claims 1 to 6, wherein performing the second enhancement layer processing comprises:inhibiting the enhancement layer encoding of the second image region from being obtained.

8. A method according to any of claims 1 to 6, wherein performing the second enhancement layer processing comprises:obtaining the enhancement layer encoding of the second image region.

9. A method according to any of claims 1 to 8, wherein obtaining the enhancement layer encoding of the second image region comprises:obtaining a set of residuals corresponding to the second image region, the set of residuals corresponding to the second image region being quantised to zero.

10. A method according to any of claims 1 to 8, wherein obtaining the enhancement layer encoding of the first image region comprises:obtaining a set of residuals corresponding to the first image region, the set of residuals corresponding to the first image region being quantised with a first quantisation step width, andwherein generating the enhancement layer encoding of the second image region comprises:obtaining a set of residuals corresponding to the second image region, the set of residuals corresponding to the second image region being quantised with a second, higher quantisation step width.

11. A method according to any of claims 1 to 10, wherein the second image region surrounds the first image region.

12. A method according to any of claims 1 to 11, wherein the first image region corresponds to an expected field of view of a user of the XR display device.

13. A method according to any of claims 1 to 12, wherein the image is a stereoscopicimage having a left-side image portion and a right-side image portion.

14. A method according to claim 13, wherein the second image region includes a boundary between the left-side and right-side image portions.

15. A method according to claim 13 or 14, wherein the first image region comprises first and second sub-regions, and wherein the first and second sub-regions are nonoverlapping.

16. A method according to claim 15, wherein the first sub-region is central to one of the left-side and right-side image portions, and wherein the second sub-region is central to the other of the left-side and right-side image portions.

17. A method according to any of claims 1 to 16, wherein inhibiting the enhancement layer encoding of the second image region from being obtained comprises:causing an enhancement layer mask to be applied to the second image region.

18. A method according to any of claims 1 to 17, wherein the second image region includes at least part of an edge of the image.

19. A method according to any of claims 1 to 18, comprising:performing third enhancement layer processing for a third image region of the image, wherein performing third enhancement layer processing comprises obtaining an enhancement layer encoding of the third image region in accordance with a third, different encoding priority.

20. A method according to claim 19, wherein the third image region is between the first and second image regions, and the third encoding priority is between the first and second encoding priorities.

21. A method according to any of claims 1 to 20, wherein the enhancement layer comprises residuals resulting from upsampler-downsampler asymmetry and / or residuals resulting from encoder-decoder errors.

22. A method according to any of claims 1 to 21, wherein the base layer encoding of the image comprises:a first base layer encoding portion corresponding to the first image region; and a second base layer encoding portion corresponding to the second image region, wherein the first base layer encoding portion has a first encoding quality and wherein the second base layer encoding portion has a second, lower encoding quality.

23. A method of processing an image, the image comprising extended reality, XR, content and being for an XR display device, the method comprising:obtaining a base layer encoding of the image;obtaining an enhancement layer encoding of a first image region of the image, the enhancement layer encoding of the first image region being in accordance with a first encoding priority; andeither:obtaining an enhancement layer encoding of the second image region, the enhancement layer encoding of the second image region being in accordance with a second, lower encoding priority; ornot obtaining an enhancement layer encoding of the second image region,wherein the method further comprises:using the obtained encodings to reconstruct the image; andcausing the reconstructed image to be displayed on the XR display device.

24. Apparatus configured to perform a method according to any of claims 1 to 23.

25. A bitstream comprising:abase layer encoding of an image comprising extended reality, XR, content and being for an XR display device; andan enhancement layer encoding of a first image region of the image, the enhancement layer encoding of a first image region having been generated in accordance with a first encoding priority,wherein either:the bitstream comprises an enhancement layer encoding of the second image region, the enhancement layer encoding of the second image region having been generated in accordance with a second, lower encoding priority; or the bitstream does not comprise an enhancement layer encoding of the second image region.

Citation Information

Patent Citations

  • Virtual reality video streaming using viewport information

    WO2019117629A1