Sensor fusion

The hierarchical data signal processing system addresses memory and bandwidth limitations by using tiered quality levels and residual data to efficiently reconstruct signals, enhancing accuracy and reducing resource consumption for improved XR experiences.

GB2642372APending Publication Date: 2026-01-07V NOVA INT LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
GB2024013338
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2026-01-07

AI Technical Summary

Technical Problem

Existing sensor systems face challenges in implementing diverse perspectives of an environment due to computer memory availability and bandwidth bottlenecks, making it difficult to achieve a comprehensive view.

Method used

A hierarchical data signal processing system with tiered levels of quality, utilizing encoding and decoding techniques to efficiently transmit and reconstruct signals, including video and image data, using residual data and configuration information to enhance reconstruction accuracy and reduce resource usage.

Benefits of technology

The system enables accurate and efficient reconstruction of high-quality signals with reduced resource consumption, particularly beneficial for mobile devices and XR applications, while minimizing latency and optimizing bandwidth usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
  • Figure 00000000_0001_ABST
    Figure 00000000_0001_ABST
Patent Text Reader

Abstract

Controlling device, comprising: obtaining first signal (1602) representing environment, having been captured using first sensor (Fig. 12; 1208); receiving down-scaled version second signal (1650) repr
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field The present disclosure relates to sensor fusion. More particularly, but not exclusively, the present disclosure relates to sensor fusion in a swarm of robots and / or unmanned aerial vehicles (UAVs). Background A device comprising a sensor (for example, a camera) can obtain a perspective of an environment in which the device is located. Multiple devices in the environment with respective sensors can also obtain respective perspectives of the environment. In principle, different perspectives may be used, for example to get a more complete perspective of the environment. However, this may be difficult to implement in practice, for example in view of computer memory availability, usage and / or bandwidth bottlenecks. Summary Various aspects of the present disclosure are set out in the appended claims. Further features and advantages will become apparent from the following description of preferred embodiments, given by way of example only, which is made with reference to the accompanying drawings. Brief Description of the Drawings Figure 1 shows a schematic block diagram of an example of a signal processing system; Figures 2A and 2B show schematic block diagrams of another example of a signal processing system; Figure 3 shows a schematic diagram of an example of a hierarchical data signal processing arrangement; Figure 4 shows a schematic diagram of another example of a hierarchical data signal processing arrangement; Figure 5 shows a schematic diagram of an example encoder for a hierarchical encoding scheme; Figure 6 shows a schematic diagram of a number of levels of quality within a first example hierarchical coding scheme; Figure 7 shows a schematic diagram of a number of levels of quality within a second example hierarchical coding scheme; Figure 8 shows a schematic diagram of an example of a bytestream structure for a frame; Figure 9 shows a schematic diagram of an example of a coding structure; Figure 10 shows a schematic block diagram of an example of a system for performing a statistical coding methodology; Figure 11 shows a schematic diagram of an example of pyramidal reconstruction; Figure 12 shows a schematic diagram of an example of a signal processing system; Figure 13 shows a schematic diagram of another example of a signal processing system; Figure 14 shows a schematic diagram of another example of a signal processing system; Figure 15 shows a schematic diagram of another example of a signal processing system; Figure 16 shows a schematic diagram of another example of a signal processing system; Figure 17 shows a schematic diagram of another example of a signal processing system; Figure 18 shows a schematic diagram of another example of a signal processing system; Figure 19 shows a schematic diagram of another example of a signal processing system; and Figure 20 shows a schematic block diagram of an example of an apparatus. Detailed Description Referring to Figure 1, there is shown an example of a signal processing system 100. The signal processing system 100 is used to process signals. Examples of types of signal include, but are not limited to, video signals, image signals, audio signals, volumetric signals such as those used in medical, scientific or holographic imaging, or other multidimensional signals. The signal processing system 100 includes a first apparatus 102 and a second apparatus 104. The first apparatus 102 and second apparatus 104 may have a clientserver relationship, with the first apparatus 102 performing the functions of a server device and the second apparatus 104 performing the functions of a client device. The signal processing system 100 may include at least one additional apparatus (not shown). The first apparatus 102 and / or second apparatus 104 may comprise one or more components. The one or more components may be implemented in hardware and / or software. The one or more components may be co-located or may be located remotely from each other in the signal processing system 100. Examples of types of apparatus include, but are not limited to, computerised devices, handheld or laptop computers, tablets, mobile devices, games consoles, smart televisions, set-top boxes, Extended Reality (XR) headsets (including Augmented Reality (AR) and / or Virtual Reality (VR) headsets) etc. The first apparatus 102 is communicatively coupled to the second apparatus 104 via a data communications network 106. Examples of the data communications network 106 include, but are not limited to, the Internet, a Local Area Network (LAN) and a Wide Area Network (WAN). The first and / or second apparatus 102, 104 may have a wired and / or wireless connection to the data communications network 106. In this example, the first apparatus 102 comprises an encoder 108. The encoder 108 is configured to encode data comprised in and / or derived based on the signal, which is referred to hereinafter as “signal data”. For example, where the signal is a video signal, the encoder 108 is configured to encode video data. Video data comprises a sequence of multiple images or frames. The encoder 108 may perform one or more further functions in addition to encoding signal data. The encoder 108 may be embodied in various different ways. For example, the encoder 108 may be embodied in hardware and / or software. The encoder 108 may encode metadata associated with the signal. The first apparatus 102 may use one or more than one encoder 108. Although in this example the first apparatus 102 comprises the encoder 108, in other examples the first apparatus 102 is separate from the encoder 108. In such examples, the first apparatus 102 is communicatively coupled to the encoder 108. The first apparatus 102 may be embodied as one or more software functions and / or hardware modules. In this example, the second apparatus 104 comprises a decoder 110. The decoder 110 is configured to decode signal data. The decoder 110 may perform one or more further functions in addition to decoding signal data. The decoder 110 may be embodied in various different ways. For example, the decoder 110 may be embodied in hardware and / or software. The decoder 110 may decode metadata associated with the signal. The second apparatus 104 may use one or more than one decoder 110. Although in this example the second apparatus 104 comprises the decoder 110, in other examples the second apparatus 104 is separate from the decoder 110. In such examples, the second apparatus 104 is communicatively coupled to the decoder 110. The second apparatus 104 may be embodied as one or more software functions and / or hardware modules. The encoder 108 encodes signal data and transmits the encoded signal data to the decoder 110 via the data communications network 106. The decoder 110 decodes the received, encoded signal data and generates decoded signal data. The decoder 110 may output the decoded signal data, or data derived using the decoded signal data. For example, the decoder 110 may output such data for display on one or more display devices associated with the second apparatus 104. The one or more display devices may be components of the second apparatus 104 or may otherwise be associated with the second apparatus 104. The one or more display devices may be operable to display XR content and may, therefore, be referred to as XR display devices. In some examples described herein, the encoder 108 transmits to the decoder 110 a representation of a signal at a given level of quality and information the decoder 110 can use to reconstruct a representation of some or all of the signal at one or more higher levels of quality. Such information may be referred to as “reconstruction data”. In some examples, “reconstruction” of a representation involves obtaining a representation that is not an exact replica of an original representation. The extent to which the representation is the same as the original representation may depend on various factors including, but not limited to, quantisation levels. A representation of a signal at a given level of quality may be considered to be a rendition, version or depiction of data comprised in the signal at the given level of quality. In some examples, the reconstruction data is included in the signal data that is encoded by the encoder 108 and transmitted to the decoder 110. For example, the reconstruction data may be in the form of metadata. In some examples, the reconstruction data is encoded and transmitted separately from the signal data. The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may comprise residual data, as described in more detail below. Residual data is an example of reconstruction data. The information the decoder 110 uses to reconstruct the representation of the signal at the one or more higher levels of quality may also comprise configuration data relating to processing of the residual data. The configuration data may indicate how the residual data has been processed by the encoder 108 and / or how the residual data is to be processed by the decoder 110. The configuration data may be signalled to the decoder 110, for example in the form of metadata. The first and / or second apparatuses 102, 104 may be configured to perform some or all of the techniques described herein. A computer program may be configured to perform some or all of the techniques described herein. Referring to Figures 2A and 2B, there is shown schematically an example of a signal processing system 200. The signal processing system 200 includes a first apparatus 202 and a second apparatus 204. In this example, the first apparatus 202 comprises an encoder and the second apparatus 204 comprises a decoder. However, as explained above, in other examples, the encoder is not comprised in the first apparatus 202 and / or the decoder is not comprised in the second apparatus 204. In each of the first apparatus 202 and the second apparatus 204, items are shown on two logical levels. The two levels are separated by a dashed line. Items on the first, highest level relate to data at a first level of quality. Items on the second, lowest level relate to data at a second level of quality. The first level of quality is higher than the second level of quality. The first and second levels of quality relate to a tiered hierarchy having multiple levels of quality. In some examples, the tiered hierarchy comprises more than two levels of quality. In such examples, the first apparatus 202 and the second apparatus 204 may include more than two different levels. There may be one or more other levels above and / or below those depicted in Figures 2 A and 2B. As described herein, in certain cases, the levels of quality may correspond to different spatial resolutions. Referring first to Figure 2A, the first apparatus 202 obtains a first representation of an image at the first level of quality 206. A representation of a given image is a representation of data comprised in the image. The image may be a given frame of a video. The first representation of the image at the first level of quality 206 will be referred to as “input data” hereinafter as, in this example, it is data provided as an input to the encoder in the first apparatus 202. The first apparatus 202 may receive the input data 206. For example, the first apparatus 202 may receive the input data 206 from at least one other apparatus. The first apparatus 202 may be configured to receive successive portions of input data 206, e.g. successive frames of a video, and to perform the operations described herein to each successive frame. For example, a video may comprise frames Fi, F2, ... Ft and the first apparatus 202 may process each of these in turn. The first apparatus 202 derives data 212 based on the input data 206. In this example, the data 212 based on the input data 206 is a representation 212 of the image at the second, lower level of quality. In this example, the data 212 is derived by performing a downsampling operation on the input data 206 and will therefore be referred to as “downsampled data” hereinafter. In other examples, the data 212 is derived by performing an operation other than a downsampling operation on the input data 206, or the data 212 is the same as the input data 206 (i.e. the input data 206 is not processed, e.g. downsampled). In this example, the downsampled data 212 is processed to generate processed data 213 at the second level of quality. In other examples, the downsampled data 212 is not processed at the second level of quality. As such, the first apparatus 202 may generate data at the second level of quality, where the data at the second level of quality comprises the downsampled data 212 or the processed data 213. In some examples, generating the processed data 213 involves the downsampled data 212 being encoded. Such encoding may occur within the first apparatus 202, or the first apparatus 202 may output the processed data 213 to an external encoder. Encoding the downsampled data 212 produces an encoded image at the second level of quality. The first apparatus 202 may output the encoded image, for example for transmission to the second apparatus 204. A series of encoded images, e g. forming an encoded video, as output for transmission to the second apparatus 204 may be referred to as a “base” stream or “base” layer. As explained above, instead of being produced in the first apparatus 202, the encoded image may be produced by an encoder that is separate from the first apparatus 202. The encoded image may be part of an H.264 or H.265 encoded video, or otherwise. Generating the processed data 213 may, for example, comprise generating successive frames of video as output by a separate encoder such as an H.264 or H.265 video encoder. An intermediate set of data for the generation of the processed data 213 may comprise the output of such an encoder, as opposed to any intermediate data generated by the separate encoder. Generating the processed data 213 at the second level of quality may further involve decoding the encoded image at the second level of quality. The decoding operation may be performed to emulate a decoding operation at the second apparatus 204, as will become apparent below. Decoding the encoded image produces a decoded image at the second level of quality. In some examples, the first apparatus 202 decodes the encoded image at the second level of quality to produce the decoded image at the second level of quality. In other examples, the first apparatus 202 receives the decoded image at the second level of quality, for example from an encoder and / or decoder that is separate from the first apparatus 202. The encoded image may be decoded using an H.264 or H.265 decoder. The decoding by a separate decoder may comprise inputting encoded video, such as an encoded data stream configured for transmission to a remote decoder, into a separate black-box decoder implemented together with the first apparatus 202 to generate successive decoded frames of video. Processed data 213 may thus comprise a frame of video data that is generated via a complex non-linear encoding and decoding process, where the encoding and decoding process may involve modelling spatio-temporal correlations as per a particular encoding standard such as H.264 or H.265. However, because the output of any encoder is fed into a corresponding decoder, this complexity is effectively hidden from the first apparatus 202. In an example, generating the processed data 213 at the second level of quality further involves obtaining correction data based on a comparison between the downsampled data 212 and the decoded image obtained by the first apparatus 202, for example based on the difference between the downsampled data 212 and the decoded image. The correction data can be used to correct for encoder-decoder errors (which may also be referred to as “encode-decode errors”), namely errors introduced in encoding and decoding the downsampled data 212. In some examples, the first apparatus 202 outputs the correction data, for example for transmission to the second apparatus 204, as well as the encoded signal. This allows the recipient to correct for the encoder-decoder errors introduced in encoding and decoding the downsampled data 212. This correction data may also be referred to as a “first enhancement” stream. As the correction data may be based on the difference between the downsampled data 212 and the decoded image it may be seen as a form of residual data (e.g. that is different from the other set of residual data described later below). An item of residual data may be referred to as “a residual”. In some examples, generating the processed data 213 at the second level of quality further involves correcting the decoded image using the correction data. For example, the correction data as output for transmission may be placed into a form suitable for combination with the decoded image, and then added to the decoded image. This may be performed on a frame-by-frame basis. In other examples, rather than correcting the decoded image using the correction data, the first apparatus 202 uses the downsampled data 212. For example, in certain cases, just the encoded then decoded data may be used and, in other cases, encoding and decoding may be replaced by other processing. In some examples, generating the processed data 213 involves performing one or more operations other than the encoding, decoding, obtaining, and correcting acts described above. The first apparatus 202 obtains data 214 based on the data at the second level of quality. As indicated above, the data at the second level of quality may comprise the processed data 213, or the downsampled data 212 where the downsampled data 212 is not processed at the lower level. As described above, in certain cases, the processed data 213 may comprise a reconstructed video stream (e.g. from an encoding-decoding operation) that is corrected using correction data. In the example of Figures 2A and 2B, the data 214 is a second representation of the image at the first level of quality, the first representation of the image at the first level of quality being the input data 206. The second representation at the first level of quality may be considered to be a preliminary or predicted representation of the image at the first level of quality. In this example, the first apparatus 202 derives the data 214 by performing an upsampling operation on the data at the second level of quality. The data 214 will be referred to hereinafter as “upsampled data”. However, in other examples one or more other operations could be used to derive the data 214, for example where data 212 is not derived by downsampling the input data 206. The input data 206 and the upsampled data 214 are used to obtain residual data 216. The residual data 216 is associated with the image. The residual data 216 may be in the form of a set of residual elements, which may be referred to as a “residual frame” or a “residual image”. A residual element may be referred to as “a residual”. A residual element in the set of residual elements 216 may be associated with a respective image element in the input data 206. An example of an image element is a pixel. In this example, a given residual element is obtained by subtracting a value of an image element in the upsampled data 214 from a value of a corresponding image element in the input data 206. As such, the residual data 216 is useable in combination with the upsampled data 214 to reconstruct the input data 206. The residual data 216 may also be referred to as “reconstruction data” or “enhancement data”. In one case, the residual data 216 may form part of a “second enhancement” stream. The residual data 216 may therefore result from upsampler-downsampler asymmetry. Upsampler-downsampler asymmetry may also be referred to as “upsampling-downsampling asymmetry”, “upsample-downsample” asymmetry or the like. The first apparatus 202 obtains configuration data relating to processing of the residual data 216. The configuration data indicates how the residual data 216 has been processed and / or generated by the first apparatus 202 and / or how the residual data 216 is to be processed by the second apparatus 204. The configuration data may comprise a set of configuration parameters. The configuration data may be useable to control how the second apparatus 204 processes data and / or reconstructs the input data 206 using the residual data 216. The configuration data may relate to one or more characteristics of the residual data 216. The configuration data may relate to one or more characteristics of the input data 206. Different configuration data may result in different processing being performed on and / or using the residual data 216. The configuration data is therefore useable to reconstruct the input data 206 using the residual data 216. As described below, in certain cases, configuration data may also relate to the correction data described herein. In this example, the first apparatus 202 transmits to the second apparatus 204 data based on the downsampled data 212, data based on the residual data 216, and the configuration data (or data based on the configuration data), to enable the second apparatus 204 to reconstruct the input data 206. Turning now to Figure 2B, the second apparatus 204 receives data 220 based on (e.g. derived from) the downsampled data 212. The second apparatus 204 also receives data based on the residual data 216. For example, the second apparatus 204 may receive a “base” stream (data 220), a “first enhancement stream” (any correction data) and a “second enhancement stream” (residual data 216). The base stream may be referred to as a “base layer”. The first and / or second enhancement stream may be referred to as, and / or may be comprised in, an “enhancement layer”. The second apparatus 204 also receives the configuration data relating to processing of the residual data 216. The data 220 based on the downsampled data 212 may be the downsampled data 212 itself, the processed data 213, or data derived from the downsampled data 212 or the processed data 213. The data based on the residual data 216 may be the residual data 216 itself, or data derived from the residual data 216. In some examples, the received data 220 comprises the processed data 213, which may comprise the encoded image at the second level of quality and / or the correction data. In some examples, for example where the first apparatus 202 has processed the downsampled data 212 to generate the processed data 213, the second apparatus 204 processes the received data 220 to generate processed data 222. Such processing by the second apparatus 204 may comprise decoding an encoded image (e.g. that forms part of a “base” encoded video stream) to produce a decoded image at the second level of quality. In some examples, the processing by the second apparatus 204 comprises correcting the decoded image using obtained correction data. Hence, the processed data 222 may comprise a frame of corrected data at the second level of quality. In some examples, the encoded image at the second level of quality is decoded by a decoder that is separate from the second apparatus 204. The encoded image at the second level of quality may be decoded using an H.264 decoder. In other examples, the received data 220 comprises the downsampled data 212 and does not comprise the processed data 213. In some such examples, the second apparatus 204 does not process the received data 220 to generate processed data 222. The second apparatus 204 uses data at the second level of quality to derive the upsampled data 214. As indicated above, the data at the second level of quality may comprise the processed data 222, or the received data 220 where the second apparatus 204 does not process the received data 220 at the second level of quality. The upsampled data 214 is a preliminary representation of the image at the first level of quality. The upsampled data 214 may be derived by performing an upsampling operation on the data at the second level of quality. The second apparatus 204 obtains the residual data 216. The residual data 216 is useable with the upsampled data 214 to reconstruct the input data 206. The residual data 216 is indicative of a comparison between the input data 206 and the upsampled data 214. The second apparatus 204 also obtains the configuration data related to processing of the residual data 216. The configuration data is useable by the second apparatus 204 to reconstruct the input data 206. For example, the configuration data may indicate a characteristic or property relating to the residual data 216 that affects how the residual data 216 is to be used and / or processed, or whether the residual data 216 is to be used at all. In some examples, the configuration data comprises the residual data 216. There are several considerations relating to such processing. One such consideration is the amount of information that is generated, stored, transmitted and / or processed. The more information that is used, the greater the amount of resources that may be involved in handling such information. Examples of such resources include transmission resources, storage resources and processing resources. Some signal processing techniques allow a relatively small amount of information to be used. This may reduce the amount of data transmitted via the data communications network 106. The savings may be particularly relevant where the data relates to high quality video data, where the amount of information transmitted can be especially high. Another consideration is latency. Complex image processing may introduce latency, which may negatively impact performance. Other considerations include the ability of the decoder to perform image reconstruction accurately, reliably, and / or efficiently. Performing image reconstruction accurately and reliably may affect the ultimate visual quality of the displayed image and consequently may affect a viewer’s engagement with the image and / or with a video comprising the image. This can be especially relevant to XR. Efficient reconstruction is especially effective for mobile computing devices, which may readily be used in XR applications. Referring to Figure 3, there is shown schematically an example of a hierarchical data signal processing arrangement 300. The hierarchical data signal processing arrangement 300 represents multiple different levels of quality (LOQs). The levels of quality may relate to different levels of quality of data associated with the data signal. A factor that can be used to determine quality of image and / or video data is resolution. A higher resolution corresponds to a higher level of quality. The resolution may be spatial and / or temporal. Other factors that can be used to determine quality of image and / or video data include, but are not limited to, a level of quantization of the data, a level of frequency filtering of the data, peak signal-to-noise ratio of the data, a structural similarity (SSIM) index, etc. In this example, the hierarchical data signal processing arrangement 300 has three different layers (or ‘levels’), namely a first layer 302, a second layer 304 and a third layer 306. The hierarchical data signal processing arrangement could however have a different number of layers. The first layer 302 may be considered to be a base layer in that it represents a base level of quality and the second and third layers 304, 306 may be considered to be enhancement layers in that they represent enhancements in terms of quality over that associated with the base layer. The first layer 302 corresponds to a first level of quality LOQ1. The second layer 304 corresponds to a second level of quality LOQ2. The second level of quality LOQ2 is higher than the first level of quality LOQ1. The third layer 306 corresponds to a third level of quality LOQ2. The third level of quality LOQ3 is higher than the second level of quality LOQ2. The second level of quality LOQ2 is between (or ‘intermediate’) the first level of quality LOQ1 and the third level of quality LOQ3. In some examples, the hierarchical data signal processing arrangement 300 represents multiple different levels of quality of video data. For example, the first layer 302 may correspond to standard definition (SD) quality video, the second layer 304 may correspond to high definition (HD) quality video and the third layer 306 may correspond to ultra-high definition (UHD) video for example. Different combinations of different resolutions may be used, e.g. SD and UHD, HD and UHD, or SD, HD and UHD. In this example, each of the layers 302, 304, 306 is associated with respective enhancement data. Enhancement data may be used to generate data, such as image and / or video data, at a level of quality associated with the respective layer. Referring to Figure 4, there is shown schematically an optional variation of a hierarchical data signal processing arrangement 400. This may only be used for certain implementations and may not be used for typical implementations. The hierarchical data signal processing arrangement 400 shown in Figure 4 is similar to the hierarchical data signal processing arrangement 300 shown in Figure 3 in that it includes three layers 402, 404, 406. In this example, each of the layers 402, 404, 406 includes a set of sub-layers (or ‘sub-levels’). In this specific example, each of the layers includes four sub-layers. The first layer 402 is associated with a first level of quality LOQ1. Each of the sub-layers of the first layer 402 is associated with a respective level of quality. A first sub-layer of the first layer 402 is associated with level of quality LOQ11, a second sublayer of the first layer 402 is associated with level of quality LOQh, a third sub-layer of the first layer 402 is associated with level of quality LOQh and a fourth sub-layer of the first layer 402 is associated with level of quality LOQ h. Similarly, the second layer 404 is associated with a second level of quality LOQ2, the third layer 406 is associated with a third level of quality LOQ3 and the sub-layers of the second and third layers 404, 406 are associated with respective, increasing levels of quality. The level of quality associated with layers and sub-layers higher in the hierarchical data signal processing arrangement 400 is higher than the level of quality associated with layers and sub-layers lower in the hierarchical data signal processing arrangement 400. As such, the level of quality increases from the bottom to the top of the hierarchical data signal processing arrangement 400. Some or all of the layers 402, 404, 406 could have a number of sub-layers other than four. Some of all of the layers 402, 404, 406 could have a different number of sublayers than each other. Some of the layers 402, 404, 406 may not have any sub-layers. In this example, each of the sub-layers is associated with respective enhancement data. Enhancement data may be used to generate data, such as image and / or video data, at a level of quality associated with the respective sub-layer. The hierarchical data signal processing arrangement 400 therefore comprises a first layer having a first set of sub-layers and a second layer having a second set of sublayers. Each of the sub-layers is associated with a respective level of quality and is associated with respective enhancement data. Using a hierarchical data signal processing arrangement such as the hierarchical data signal processing arrangement 400 may allow some devices to reconstruct at a specific layer, for example LOQ3, but only using say the first two sub-layers, LOQ31 and LOQ32, of that layer. This may be for efficiency, battery saving or limited capacity purposes. Using only some sub-layers may be considered to be partial reconstruction. Other devices may use all four of the sublayers in LOQ3, namely LOQ31, LOQ32, LOQ3s and LOQ34, and reconstruct the signal completely. Using all of the sub-layers may be considered to be full reconstruction. The reader is referred to UK patent application no. GB1603727.7, which describes a hierarchical arrangement in more detail. The entire contents of GB1603727.7 are incorporated herein by reference. Figure 5 also shows 500 how a data plane of a data frame is received and is successively downsampled to generate a plurality of layers (0 to n are shown). The base layer may be a lowest layer and may represent an output of a last downsampling step (or a difference between that layer and a scalar offset). Higher layers are represented as residual data where an upsampled version of a lower layer is compared with an input (non-upsampled) version (e.g. following downsampling for layer Rn-i as shown), e.g. by subtracting a reconstructed upsampled version of a lower layer from the input version. Residual data allows efficient representation of the data frames, and may be particularly useful for sparse data where there may only be a few values within the residual data. On a decoding side the process may be reversed on receipt of encoded streams derived from Ro to Rn representing the layers. Figure 5 does not show encoding operations, which may comprise transforming, quantising and entropy encoding residual data. Some examples of hierarchical coding are set out in the SMPTE VC-6 standard, and the MPEG-5, Part-2, LCEVC standard (the specifications for both standards, including working drafts, being incorporated herein by reference). Additional descriptions of hierarchical coding may be found in one or more of U.S. Patent No. 8,977,065, filed on July 21, 2011, entitled “Inheritance in a tiered signal quality hierarchy,” the contents of which are hereby incorporated by reference in their entirety; U.S. Patent No. 8,948,248, filed on July 21, 2011, entitled “Tiered signal decoding and signal reconstruction,” the contents of which are hereby incorporated by reference in their entirety; U.S. Patent No. 8,711,943, filed on July 21, 2011, entitled “Signal processing and tiered signal encoding,” the contents of which are hereby incorporated by reference in their entirety; U.S. PatentNo. 9,129,411, filed on July 21, 2011, entitled “Upsampling in a tiered signal quality hierarchy,” the contents of which are hereby incorporated by reference in their entirety; and U.S. Patent No. 8,531,321, filed on July 21, 2011, entitled “Signal processing and inheritance in a tiered signal quality hierarchy,” the contents of which are hereby incorporated by reference in their entirety. In certain examples, in very constrained or public networks, streaming may be configured to use MPEG-5 Part 2 (LCEVC), as an alternative to the full-codec approach of SMPTE VC-6. For example, hierarchical encoding according to SMPTE VC-6 may be preferred for higher quality local transmissions where requirements are similar to video production (e.g. neighbouring rooms in a hospital or education establishment), but MPEG-5 Part 2 may be preferred for wider streaming over secure or unsecure networks (e.g. streaming over the Internet). Figure 6 shows 600 the approach of SMPTE VC-6, where there are a plurality of layers of quality going down to a lowest layer of quality, where all layers are encoded with a common codec. Figure 6 identifies layers 610-0, 610-1 and 610-n. Figure 7 shows 700 the approach of MPEG-5, Part 2, wherein there is a base codec and one or more layers that are encoded using a separate enhancement codec. When bandwidth is limited, MPEG-5, Part 2 may degrade to just use a base stream as generated by a base or legacy codec. This may be at a lower resolution to a full desired stream, wherein the enhancement levels or layers may allow a higher resolution. Instead of needing the base codec to encode 100% of the initial resolution, it may only need to encode 25% of the initial resolution, thus reducing the load on the base codec and allowing base encoded frames to be received even if resources are limited. If the enhancement codec is more efficient than the base codec, MPEG-5 Part 2 may allow a legacy codec to continue operating within its comfort zone while the quality at the higher resolution remains unaffected. Figure 7 shows a base codec 710, and two enhancement layers 720-1, 720-2. As shown in Figure 8, a bytestream 800 may include multiple fields, namely one or more headers and a payload. In general, a payload includes the actual data to be decoded, whilst the headers provide information needed when decoding the payload. The payload may include information about a plurality of planes. In other words, the payload is subdivided in portions, each portion corresponding to a plane. Each plane further comprises multiple sub-portions, each sub-portion associated with a level of quality. The logical structure of a Payload is an array of multi-tiered Tableaux, which precedes the Tile Tier with Tiles containing Residuals at their Top Layer. The data in the Payload that represents a Tessera is a a Stream. In the present example, Streams are ordered by LoQ, then by Plane, then by direction and then by Tier. However, the Streams can be ordered in any other way, for example first direction, then LoQ, the Plane, then Tier. The order between directions, LoQ and Planes can be done in any way, and the actual order can be inferred by using the information in the header, for example the stream offsets info. The payload contains a series of streams, each stream corresponding to an encoded tessera. For the purpose of this example, we assume that the size of a tessera is 16x16. First, the decoding module would derive a root tableau (for example, associated with a first direction of a first LoQ within a first plane). From the root tableau, the decoding module would derive up to 256 attributes associated with the corresponding up to 256 tesserae associated with it and which lie in the tier above the root tier (first tier). In particular, one of the attributes is the length of the stream associated with the tessera. By using said streamlengths, the decoding module can identify the individual streams and, if implemented, decode each stream independently. Then, the decoding module would derive, from each of said tessera, attributes associated with the 256 tesserae in the tier above (second tier). One of these attributes is the length of the stream associated with the tessera. By using said streamlengths, the decoding module can identify the individual streams and, if implemented, decode each stream independently. The process will continue until the top tier is reached. Once the top tier has been reached, the next stream in the bytestream would correspond to a second root tableau (for example, associated with a second direction of a first LoQ within a first plane), and the process would continue in the same way. The bytestream may include a fixed-sized header, i.e. a header whose byte / bit length is fixed. The header may include a plurality of fields. The fixed-sized header may include a first field indicating a version of the bytestream format (B.l - also described as format version : unit8). In an embodiment, this first field may include 8 bits (or equivalently 1 byte). This field may allow flexibility in the encoding / decoding process to use, adapt and / or modify the version of the bytestream format and inform a decoding module of said version. In this way, it is possible to use multiple different version of the encoding / decoding format and allow the decoding module to determine the correct version to be used. A decoding module would obtain said first field from the bytestream and determine, based on the value included in said first field, a version of the encoding format to be used in the decoding process of said bytestream. The decoding module may use and / or implement a decoding process to adapt to said version. The fixed-sized header may include a second field indicating a size of the picture frame encoded with a specific bytestream (B.2 - also described as picture size : unit32). The size of the picture frame may actually correspond to the size of the bytestream associated with that picture frame. In an embodiment, this first field may include 32 bits (or equivalently 4 bytes). The size of the picture frame may be indicated in units of bytes, but other units may be used. This allows the encoding / decoding process flexibility in encoding picture frames of different size (e.g., 1024x720 pixels, 2048x1540 pixels, etc.) and allow the decoding module to determine the correct picture frame size to be used for a specific bytestream. A decoding module would obtain said second field from the bytestream and determine, based on the value included in said second field, a size of a picture frame corresponding to said bytestream. The decoding module may use and / or implement a decoding process to adapt to said size, and in particular to reconstruct the picture frame from the encoded bytestream to fit into said size. The fixed-sized header may include a third field indicating a recommended number of bits / bytes to fetch / retrieve at the decoding module when obtaining the bytestream (B.3 - also described as recommendedfetchsize : unit32). In an embodiment, this first field may include 32 bits (or equivalently 4 bytes). This field may be particularly useful in certain applications and / or for certain decoding modules when retrieving the bytestream from a server, for example to enable the bytestream to be fetched / retrieved at the decoding module in “portions”. For example, this may enable partial decoding of the bytestream (as further described, for example, in European patent application No 17386047.9 filed on 6th December 2017 by the same applicant whose contents are included in their entirety by reference) and / or optimise the retrieval of the bytestream by the decoding module (as for example further described in European patent application No 12759221.0 filed on 20th July 2012 by the same applicant whose contents are included in their entirety by reference). A decoding module would obtain said third field from the bytestream and determine, based on the value included in said third field, a number of bits and / or bytes of the bytestream to be retrieved from a separate module (for example, a server and / or a content delivery network). The decoding module may use and / or implement a decoding process to request to the separate module said number of bits and / or bytes from the bytestream, and retrieve them from the separate module. The fixed-sized header may include another field indicating a generic value in the bytestream (B.3.1 - also described as element interpretation : uint8). In an embodiment, this first field may include 8 bits (or equivalently 1 byte). A decoding module would obtain said another field from the bytestream and determine, based on the value included in said another field, a value indicated by the field. The fixed-sized header may include a fourth field indicating various system information, including the type of transform operation to be used in the decoding process (B 4 - also described as pipeline : unit8). In an embodiment, this first field may include 8 bits (or equivalently 1 byte). A transform operation is typically an operation that transform a value from an initial domain to a transformed domain. One example of such a transform is an integer composition transform. Another example of such a transform is a composition transform. The composition transform (integer and / or standard) are further described in European patent application No. 13722424.2 filed on 13th May 2013 by the same applicant and incorporated herein by reference. A decoding module would obtain said fourth field from the bytestream and determine, based on at least one value included in said fourth field, a type of transform operation to be used in the decoding process. The decoding module may configure the decoding process to use the indicated transform operation and / or implement a decoding process which uses the indicated transform operation when converting one or more decoded transformed coefficient and / or value (e.g., a residual) into an original nontransform domain. The fixed-sized header may include a fifth field indicating a type of up-sampling filtering operation to be used in the decoding process (B.5 - also described as upsampler : unit8). In an embodiment, this first field may include 8 bits (or equivalently 1 byte). An up-sampling filtering operation comprises a filter which applies certain mathematical operations to a first number of samples / values to produce a second number of samples / values, wherein the second number is higher than the first number. The mathematical operations can either be pre-defined, adapted either based on an algorithm (e.g., using a neural network or some other adaptive filtering technique) or adapted based on additional information received at the decoding module. Examples of such up-sampling filtering operations comprise a Nearest Neighbour filtering operation, a Sharp filtering operation, a Bi-cubic filtering operation, and a Convolutional Neural Network (CNN) filtering operations. These filtering operations are described in further detail in the present application, as well as in UK patent application No. 1720365.4 filed on 6th December 2017 by the same applicant and incorporated herein by reference. A decoding module would obtain said fifth field from the bytestream and determine, based on at least one value included in said fifth field, a type of up-sampling operation to be used in the decoding process. The decoding module may configure the decoding process to use the indicated up-sampling operation and / or implement a decoding process which uses the indicated up-sampling operation. The indication of the upsampling operation to be used allows flexibility in the encoding / decoding process, for example to better suit the type of picture to be encoded / decoded based on its characteristics. The fixed-sized header may include a sixth field indicating one or more modifying operations used in the encoding process when building the fixed-sized header and / or other headers and / or to be used in the decoding process in order to decode the bytestream (see below) (B.6 - also described as shortcuts : shortcuts_t). These modifying operations are also called shortcuts. The general advantage provided by these shortcuts is to reduce the amount of data to be encoded / decoded and / or to optimise the execution time at the decoder, for example by optimising the processing of the bytestream. A decoding module would obtain said sixth field from the bytestream and determine, based on at least one value included in said sixth field, a type of shortcut used in the encoding process and / or to be used in the decoding process. The decoding module may configure the decoding process to adapt its operations based on the indicated shortcut and / or implement a decoding process which uses the indicated shortcut. The fixed-sized header may include a seventh field indicating a first number of bits to be used to represent an integer number and a second number of bits to be used to represent a fractional part of a number (B.7 - also described as element descriptor : tuple (uint5, utin3)). In an embodiment, this first field may include 8 bits (or equivalently 1 byte) subdivided in 5 bits for the first number of bits and 3 bits for the second number of bits. A decoding module would obtain said seventh field from the bytestream and determine, based on at least one value included in said seventh field, how many bits to dedicate to represent the integer part of a number that has both integer and fractional parts and how many bits to dedicate to a fractional number. The fixed-sized header may include an eighth field indicating a number of planes forming a frame and to be used when decoding the bytestream (B.8 - also described as num_plane : unit8). In an embodiment, this first field may include 8 bits (or equivalently 1 byte). A plane is defined in the present application and is, for example, one of the dimensions in a color space, for examples the luminance component Y in a YUV space, or the red component R in an RGB space. A decoding module would obtain said eighth field from the bytestream and determine, based on at least one value included in said fifth field, the number of planes included in a picture. The fixed-sized header may include a ninth field indicating a size of an auxiliary header portion included in a separate header - for example the First Variable-Size Header or the Second Variable-Size Header (B.9 - also described as aux header size : uuntl6). In an embodiment, this first field may include 16 bits (or equivalently 2 byte). This field allows the encoding / decoding process to be flexible and define potential additional header fields. A decoding module would obtain said ninth field from the bytestream and determine, based on at least one value included in said ninth field, a size of an auxiliary header portion included in a separate header. The decoding module may configure the decoding process to read the auxiliary header in the bytestream. The fixed-sized header may include a tenth field indicating a number of auxiliary attributes (B.10 - also described as num aux tile attribute : uint4 and numauxtableauattribute : uint4). In an embodiment, this first field may include 8 bits (or equivalently 1 byte) split into two 4-bits sections. This field allows the encoding / decoding process to be flexible and define potential additional attributes for both Tiles and Tableaux. These additional attributes may be defined in the encoding / decoding process. A decoding module would obtain said tenth field from the bytestream and determine, based on at least one value included in said tenth field, a number of auxiliary attributes associated with a tile and / or a number of auxiliary attributes associated with a tableau. The decoding module may configure the decoding process to read said auxiliary attributes in the bytestream. The bytestream may include a first variable-sized header, i.e. a header whose byte / bit length is changeable depending on the data being transmitted within it. The header may include a plurality of fields. The first variable-sized header may include a first field indicating a size of a field associated with an auxiliary attribute of a tile and / or a tableau (C. 1 - also described as aux attribute sizes : unti!6[num aux tile attribute + num aux tableau attribute]). In an embodiment, the second field may include a number of sub-fields, each indicating a size for a corresponding auxiliary attribute of a tile and / or a tableau. The number of these sub-fields, and correspondingly the number of auxiliary attributes for a tile and / or a tableau, may be indicated in a field of a different header, for example the fixed header described above, in particular in field B.10. In an embodiment, this first field may include 16 bits (or equivalently 2 bytes) for each of the auxiliary attributes. Since the auxiliary attributes may not be included in the bytestream, this field would allow the encoding / decoding process to define the size of the auxiliary attributes were they to be included in the bytestream. This contrasts, for example, with the attributes (see for example C.2 below) which typically are pre-defined in size and therefore their size does not need to be specified and / or communicated. A decoding module would obtain said first field from the bytestream and determine, based on a value included in said first field, a size of an auxiliary attribute associated with a tessera, (i.e., either a tile or a tableau). In particular, the decoding module may obtain from said first field in the bytestream, a size of an auxiliary attribute for each of the auxiliary attributes which the decoding module is expecting to decode, for example based on information received separately about the number of auxiliary attributes to be specified. The decoding module may configure the decoding process to read the auxiliary attributes in the bytestream. The first variable-sized header may include a second field indicating, for each attribute of a tile and / or a tableau, a number of different versions of the respective attribute (C.2 - also described as nums attribute : until6[4 + num aux tile attribute + num aux tableau attribute]). The second field may include a number of sub-fields, each indicating for a corresponding attribute a number of different version of said respective attribute. The number of these sub-fields, and correspondingly the number of standard attributes and auxiliary attributes for a tile and / or a tableau, may be indicated at least in part in a field of a different header, for example the fixed header described above, in particular in field B.10. The attributes may comprise both standard attributes associated with a tile and / or a tableau and the auxiliary attributes as described above. In an embodiment, there are three standard attributes associated with a tile (e.g., Residual Statistics, T-Node Statistics and Quantization Parameters) and two standard attributes associated with a tableau (e.g., Streamlengths Statistics and T-Node Statistics). In an embodiment, since the T-Node Statistics for the tiles and the tableaux may be the same, they may only require to be specified once. In such embodiment, only four different standard attributes will need to be included (and therefore only four sub-fields, C.2.1 to C.2.4, each associated with one of the four standard attributes Residual Statistics, T-Node Statistics, Quantization Parameters and Streamlengths Statistics, are included in the second field, each indicating a number of different versions of the respective attribute). Accordingly, there may be four different sub-fields in said second field, each indicating the number of standard attributes for a tile and / or a tableau which need to be specified for the decoding process. By way of example, if the sub-field associated with the T-Node Statistics indicate a number 20, it means that there will be 20 different available versions of T-Node Statistics to use for tiles and / or attributes. A decoding module would obtain said second field from the bytestream and determine, based on a value included in said second field, a number of different versions of a respective attribute, said attribute associated with a tile and / or a tableau. The decoding module may configure the decoding process to use the available versions of the corresponding attributes. The first variable-sized header may include a third field indicating a number of different groupings of tiles, wherein each grouping of tiles is associated with a common attribute (C.3 - also described as num tileset: uintl6). In an embodiment, this first field may include 16 bits (or equivalently 2 bytes). In an embodiment, the common attribute may be the T-Node Statistics for a tile. For example, if a grouping of tiles (also known as “tileset”) is associated with the same T-node Statistics, it means that all the tiles in that grouping shall be associated with the same T-Node Statistics. The use of grouping of tiles sharing one or more common attributes allows the coding and decoding process to be flexible in terms of specifying multiple versions of a same attribute and associate them with the correct tiles. For example, if a group of tiles belongs to “Group A”, and “Group A” is associated with “Attribute A” (for example, a specific T-Node Statistics), then all the tiles in Group A shall use that Attribute A. Similarly, if a group of tiles belongs to “Group B”, and “Group B” is associated with “Attribute B” (for example, a specific T-Node Statistics different from that of Group A), then all the tiles in Group B shall use that Attribute B. This is particularly useful in allowing the tiles to be associated with a statistical distribution as close as possible to that of the tile but without having to specify different statistics for every tile. In this way, a balance is reached between optimising the entropy encoding and decoding (optimal encoding and decoding would occur if the distribution associated with the tile is the exact distribution of that tile) whilst minimising the amount of data to be transmitted. Tiles are grouped, and a “common” statistics is used for that group of tiles which is as close as possible to the statistics of the tiles included in that grouping. For example, if we have 256 tiles, in an ideal situation we would need to send 256 different statistics, one for each of the tiles, in order to optimise the entropy encoding and decoding process (an entropy encoder / decoder is more efficient the more the statistical distribution of the encoded / decoded symbols is close to the actual distribution of said symbols). However, sending statistics is impractical and expensive in terms of compression efficiency. So, typical systems would send only one single statistics for all the 256 tiles. However, if the tiles are grouped into a limited number of groupings, for example 10, with each tile in each grouping having similar statistics, then only 10 statistics would need to be sent. In this way, a better encoding / decoding would be achieved than if only one common statistics was to be sent for all the 256 tiles, whilst at the same time sending only 10 statistics and therefore not compromising too much the compression efficiency. A decoding module would obtain said third field from the bytestream and determine, based on a value included in said third field, a number of different groupings of tiles. The decoding module may configure the decoding process to use, when decoding a tile corresponding to a specific grouping, one or more attributes associated with said grouping. The first variable-sized header may include a fourth field indicating a number of different groupings of tableaux, wherein each grouping of tableaus is associated with a common attribute (C.4 - also described as num tableauset : uintl6). In an embodiment, this fourth field may include 16 bits (or equivalently 2 bytes). This field works and is based on the same principles as the third field, except that in this case it refers to tableaux rather than tiles. A decoding module would obtain said fourth field from the bytestream and determine, based on a value included in said fourth field, a number of different groupings of tableaux. The decoding module may configure the decoding process to use, when decoding a tableau corresponding to a specific grouping, one or more attributes associated with said grouping. The first variable-sized header may include a fifth field indicating a width for each of a plurality of planes (C.5 - also described as widths : uintl6[num_plane]). In an embodiment, this fifth field may include 16 bits (or equivalently 2 bytes) for each of the plurality of planes. A plane is further defined in the present specification, but in general is a grid (usually a two-dimensional one) of elements associated with a specific characteristic, for example in the case of video the characteristics could be luminance, or a specific color (e.g. red, blue or green). The width may correspond to one of the dimensions of a plane. Typically, there are a plurality of planes. A decoding module would obtain said fifth field from the bytestream and determine, based on a value included in said fifth field, a first dimension associated with a plane of elements (e.g., picture elements, residuals, etc.). This first dimension may be the width of said plane. The decoding module may configure the decoding process to use, when decoding the bytestream, said first dimension in relation to its respective plane. The first variable-sized header may include a sixth field indicating a width for each of a plurality of planes (C.6 - also described as heights : uintl6[num_plane]). In an embodiment, this sixth field may include 16 bits (or equivalently 2 bytes) for each of the plurality of planes. The height may correspond to one of the dimensions of a plane. A decoding module would obtain said sixth field from the bytestream and determine, based on a value included in said sixth field, a second dimension associated with a plane of elements (e.g., picture elements, residuals, etc.). This second dimension may be the height of said plane. The decoding module may configure the decoding process to use, when decoding the bytestream, said second dimension in relation to its respective plane. The first variable-sized header may include a seventh field indicating a number of encoding / decoding levels for each of a plurality of planes (C.7 - also described as num loqs : uint8[num_plane]). In an embodiment, this seventh field may include 16 bits (or equivalently 2 bytes) for each of the plurality of planes. The encoding / decoding levels correspond to different levels (e.g., different resolutions) within a hierarchical encoding process. The encoding / decoding levels are also referred in the application as Level of Quality. A decoding module would obtain said seventh field from the bytestream and determine, based on a value included in said seventh field, a number of encoding levels for each of a plurality of planes (e.g., picture elements, residuals, etc.). The decoding module may configure the decoding process to use, when decoding the bytestream, said number of encoding levels in relation to its respective plane. The first variable-sized header may include an eighth field containing information about the auxiliary attributes (C.8 - also described as aux header : uint8[auxjreader_size]). In an embodiment, this eight field may include a plurality of 8 bits (or equivalently 1 byte) depending on a size specified, for example, in a field of the fixed header (e.g., B.9) A decoding module would obtain said eighth field from the bytestream and determine information about the auxiliary attributes. The decoding module may configure the decoding process to use, when decoding the bytestream, said information to decode the auxiliary attributes. The bytestream may include a second variable-sized header, i.e. a header whose byte / bit length is changeable depending on the data being transmitted within it. The header may include a plurality of fields. The second variable-sized header may include a first field containing, for each attribute, information about one or more statistics associated with the respective attribute (see D 1). The number of statistics associated with a respective attribute may be derived separately, for example via field C.2 as described above. The statistics may be provided in any form. In an embodiment of the present application, the statistics is provided using a particular set of data information which includes information about a cumulative distribution function (type residualstatt). In particular, a first group of sub-fields in said first field may contain information about one or more statistics associated with residuals values (also D.l.l -also described as residual stats : residual_stat_t[nums_attribute[O]]). In other words, the statistics may identify how a set of residual data are distributed. The number of statistics included in this first group of sub-fields may be indicated in a separate field, for example in the first sub-field C.2.1 of field C.2 as described above (also indicated as nums_attribute[O]). For example, if nums_attribute[O] is equal to 10, then there would be 10 different residuals statistics contained in said first field. For example, the first 10 sub-fields in the first field correspond to said different 10 residuals statistics. A second group of sub-fields in said first field may contain information about one or more statistics associated with nodes within a Tessera (also D.1.2 - also described as tnode stats : tnode_stat_t[nums_attribute[l]]). In other words, the statistics may identify how a set of nodes are distributed. The number of statistics included in this second group of sub-fields may be indicated in a separate field, for example in the second sub-field C.2.2 of field C.2 as described above (also indicated as nums_attribute[l]). For example, if nums_attribute[l] is equal to 5, then there would be 5 different t-node statistics contained in said first field. For example, considering the example above, after the first 10 sub-fields in the first field, the next 5 sub-fields correspond to said 5 different t-node statistics. A third group of sub-fields in said first field may contain information about one or more quantization parameters (also D.1.3 - also described as quantization_parameters : quantization_parameters_t[nums_attribute[2]]). The number of quantization parameters included in this third group of sub-fields may be indicated in a separate field, for example in the third sub-field C.2.3 of field C.2 as described above (also indicated as nums_attribute[2]). For example, if nums_attribute[2] is equal to 10, then there would be 10 different quantization parameters contained in said first field. For example, considering the example above, after the first 15 sub-fields in the first field, the next 10 sub-fields correspond to said 10 different quantization parameters. A fourth group of sub-fields in said first field may contain information about one or more statistics associated with streamlengths (also D.1.4 - also described as stream_length_stats : streamlengthstat^ In other words, the statistics may identify how a set of streamlengths are distributed. The number of statistics included in this fourth group of sub-fields may be indicated in a separate field, for example in the fourth sub-field C.2.4 of field C.2 as described above (also indicated as nums_attribute[3]). For example, if nums_attribute[4] is equal to 12, then there would be 12 different streamlengths statistics contained in said first field. For example, considering the example above, after the first 25 sub-fields in the first field, the next 12 sub-fields correspond to said 12 different streamlengths statistics. Further groups of sub-fields in said first field may contain information about auxiliary attributes (also described as auxatttributes : uintl[aux_attributes_size[i]] [num^aux^tile^attribute + num aux tableau attribute]). The number of auxiliary attributes may be indicated in another field, for example in field C.2 as described above. Specifying one or more versions of the attributes (e.g., statistics) enables flexibility and accuracy in the encoding and decoding process, because for instance more accurate statistics can be specified for a specific grouping of tesserae (tiles and / or tableaux), thus making it possible to encode and / or decode said groupings in a more efficient manner. A decoding module would obtain said first field from the bytestream and determine, based on the information contained in said first field, one or more attributes to be used during the decoding process. The decoding module may store the decoded one or more attributes for use during the decoding process. The decoding module may, when decoding a set of data (for example, a tile and / or a tableau) and based on an indication of attributes to use in relation to that set of data, retrieve the indicated attributes from the stored decoded one or more attributes and use it in decoding said set of data. The second variable-sized header may include a second field containing, for each of a plurality of grouping of tiles, an indication of a corresponding set of attributes to use when decoding said grouping (D.2 - also described as tilesets : uintl6[3 + num aux tile attributes] [num tiles]). The number of groupings of tiles may be indicated in a separate field, for example in field C.3 described above. This second field enables the encoding / decoding process to specify which of the sets of attributes indicated in field D.l described above is to be used when decoding a tile. A decoding module would obtain said second field from the bytestream and determine, based on the information contained in said second field, which of a set of attributes is to be used when decoding a respective grouping of tiles. The decoding module would retrieve from a repository storing all the attributes the ones indicated in said second field, and use them when decoding the respective grouping of tiles. The decoding process would repeat said operations when decoding each of the plurality of grouping of tiles. By way of example, and using the example described above in relation to field D.l, let’s assume that for a first grouping of tiles the set of attributes indicated in said second field corresponds to residuals statistics No. 2, t riode statistics No. 1 and to quantization parameter No. 4 (we assume for simplicity that there are no auxiliary attributes). When the receiving module receives said indication, it would retrieve from the stored attributes (as described above) the second residuals statistics from the 10 stored residuals statistics, the first tnode statistics from the 5 stored tnode statistics and the fourth quantization parameter from the 10 stored quantization parameters. The second variable-sized header may include a fourth field containing, for each of a plurality of grouping of tableaux, an indication of a corresponding set of attributes to use when decoding said grouping (D.4 - also described as tableausets : uintl6[2 + num aux tableaux attributes] [num tableaux]). The number of groupings of tableaux may be indicated in a separate field, for example in field C.4 described above. This fourth field enables the encoding / decoding process to specify which of the sets of attributes indicated in field D. 1 described above is to be used when decoding a tableau. The principles and operations behind this fourth field corresponds to that described for the second field, with the difference that in this case it applies to tableaux rather than tiles. In particular, a decoding module would obtain said fourth field from the bytestream and determine, based on the information contained in said fourth field, which of a set of attributes is to be used when decoding a respective grouping of tableaux. The decoding module would retrieve from a repository storing all the attributes the ones indicated in said fourth field, and use them when decoding the respective grouping of tableaux. The decoding process would repeat said operations when decoding each of the plurality of grouping of tableaux. The second variable-sized header may include a fifth field containing, for each plane, each encoding / decoding level and each direction, an indication of a corresponding set of attributes to use when decoding a root tableau (D.5 - also described as root tableauset indices : uintl6[loq_idx][num_planes][4]). This fifth field enables the encoding / decoding process to specify which of the sets of attributes indicated in field D.l described above is to be used when decoding a root tableau. A decoding module would obtain said fifth field from the bytestream and determine, based on the information contained in said fifth field, which of a set of attributes is to be used when decoding a respective root tableau. The decoding module would retrieve from a repository storing all the attributes the ones indicated in said fifth field, and use them when decoding the respective grouping of tiles. In this way, the decoding module would effectively store all the possible attributes to be used when decoding tiles and / or tableaux associated with that bytestream, and then retrieve for each of a grouping of tiles and / or tableaux only the sub-set of attributes indicated in said second field to decode the respective grouping of tiles and / or tableaux. The second variable-sized header may include a third field containing information about the statistics of the groupings of tiles (D.3 - also described as cdf tilesets : line_segments_cdfl5_t<tilese_index_t>). The statistics may provide information about how many times a certain grouping of tiles occurs. The statistics may be provided in the form of a cumulative distribution function. In the present application, the way the cumulative distribution function is provided is identified as a function type, specifically type line_segments_cdfl5_t<x_axis_type>. By using said statistics, the encoding / decoding process is enabled to compress the information about the grouping of tiles (e.g., the indices of tiles) and therefore optimise the process. For example, if there are N different groupings of tiles, and correspondingly N different indexes, rather than transmitting these indexes in an uncompressed manner, which would require 2^^2^1 bits (where [.] is a ceiling function), the grouping can be compressed using an entropy encoder thus reducing significantly the number of bits required to communicate the groupings of tiles. This may represent a significant savings. For example, assume that there are 10,000 tiles encoded in the bytestream, and that these tiles are divided in 100 groupings. Without compressing the indexes, an index needs to be sent together with each tile, meaning that at least 2''°52'°0'====7 bits per tile, which means a total of 70,000 bits. If instead the indexes are compressed using an entropy encoder to an average of 1.5 bits per index, the total number of bits to be used would be 15,000, reducing the number of bits to be used by almost 80%. A decoding module would obtain said third field from the bytestream and determine, based on the information contained in said third field, statistical information about the groupings of tiles. The decoding module would use said statistical information when deriving which grouping a tile belongs to. For example, the information about the tile grouping (e.g., tileset index) can be compressed using said statistics and then reconstructed at the decoder using the same statistics, for example using an entropy decoder. The second variable-sized header may include a sixth field containing information about the statistics of the groupings of tableaux (D.6 - also described as cdf tableausets : line_segments_cdfl5_t<tableauset_index_t>). The statistics may provide information about how many times a certain grouping of tableaux occurs. The statistics may be provided in the form of a cumulative distribution function. This field works in exactly the same manner as the third field but for grouping of tableaux rather than grouping of tiles. In particular, a decoding module would obtain said sixth field from the bytestream and determine, based on the information contained in said sixth field, statistical information about the groupings of tableaux. The decoding module would use said statistical information when deriving which grouping a tableau belongs to. For example, the information about the tableau grouping (e.g., tableauset index) can be compressed using said statistics and then reconstructed at the decoder using the same statistics, for example using an entropy decoder. The second variable-sized header may include a seventh field containing, for each plane, each encoding / decoding level and each direction, an indication of a location, within a payload of the bytestream, of one or more sub-streams (e.g., a Surface) of bytes associated for that respective plane, encoding / decoding level and direction (D.7 - also described as rootstreamoffsets root_stream_offset_t[loq_idx][num_planes][4]). The location may be indicated as an offset with respect to the start of the payload. By way of example, assuming 3 planes, 3 encoding / decoding levels and 4 directions, there will be 3*3*4=36 different substreams, and correspondingly there will be 36 different indication of locations (e.g., offsets). A decoding module would obtain said seventh field from the bytestream and determine, based on the information contained in said seventh field, where to find in the payload a specific sub-stream. The sub-stream may be associated with a specific direction contained in a specific plane which is within a specific encoding / decoding level. The decoding module would use said information to locate the sub-stream and decode said sub-stream accordingly. The decoding module may implement, based on this information, decoding of the various sub-stream simultaneously and / or in parallel. This can be advantageous for at least two reasons. First, it would allow flexibility in ordering of the sub-streams. The decoder could reconstruct, based on the location of the sub-streams, to which direction, plane and encoding / decoding level the sub-stream belongs to, without the need for that order to be fixed. Second, it would enable the decoder to decode the sub-streams independently from one another as effectively each sub-stream is separate from the others. The second variable-sized header may include an eighth field containing, for each plane, each encoding / decoding level and each direction, a size of the Stream of bytes associated with the root tableau (D.8 - also described as root stream lengths : root_stream_length_t[loq_idx][num_planes][4]). A decoding module would obtain said eighth field from the bytestream and determine, based on the information contained in said eighth field, the length of a stream associated with a root tableau. Figure 9 shows a more general view of an example coding structure 900. In particular, and as already described in the present application and / or in other patent applications by the same applicant (such as European Patent No. 17386046.1 filed on 6th December 2017 and incorporated herein by reference), there are various Tiers in the coding structure, with Tier 0 being formed by Tiles, Tier -m formed of Tableaux and Root Tier formed of the single Root Tier. For example, one can see that a Summit, corresponding to maximum Grid size that could be encoded, is present both at a specific Tier (i.e., the union of all the values in Layer 0 of each Tier) or could also be seen for each tessera (i.e., the 256 values in Layer 0 of each tessera). Another important concept is that of Active Volume which effectively corresponds to that portion of the coding structure associated with the area occupied by the Surface. In other words, all those tesserae which are associated with at least one element of the Surface belongs to the Active Volume. The Surface, on the other hand, corresponds to the area of the Summit where there are actual data (e.g., residuals or metadata) to be used. Figure 10 is a block diagram of a system 1000. In Figure 10, there is shown the system 1000, the system 1000 comprising a streaming server 1002 connected via a network 1014 to a plurality of client devices 1030, 1032. The streaming server 1002 comprising an encoder 1004, the encoder configured to receive and encode a first video stream utilising the methodology described herein. The streaming server 1004 is configured to deliver an encoded video stream 1006 to a plurality of client devices such as set-top boxes smart TVs, smartphones, tablet computers, laptop computers etc., 1030 and 1032. Each client device 1030 and 1032 is configured to decode and render the encoded video stream 1006. The client devices and streaming server 1004 are connected via a network 1014. For ease of understanding the system 1000 of Figure 10 is shown with reference to a single streaming server 1002 and two recipient client devices 1030, 1032 though in further embodiments the system 1000 may comprise multiple servers (not shown) and several tens of thousands of client devices. The streaming server 1002 can be any suitable data storage and delivery server which is able to deliver encoded data to the client devices over the network. Streaming servers are known in the art, and use unicast and / or multicast protocols. The streaming server is arranged to encode and store the encoded data stream, and provide the encoded video data in one or more encoded data streams 1006 to the client devices 1030 and 1032. The encoded video stream 1006 is generated by the encoder 1004. The encoder 1004 in Figure 10 is located on the streaming server 1002, though in further embodiments the encoder 1004 is located elsewhere in the system 1000. The encoder 1004 generates the encoded video stream utilising the techniques described herein. The encoder further comprises a statistical module 1008 configured to determine and calculate statistical properties of the video data. The client devices 1030 and 1032 are devices known in the art and comprise the known elements required to receive and decode a video stream such as a processor, communications port and decoder. A Level of Quality (LoQ) represents Pels (elements of a picture) encoded at a certain resolution. With reference to Figure 11, Four Cycles of Pyramidal Reconstruction for Initial LoQ and three higher quality successor LoQs. On all higher Cycles, Residuals are dequantized and Composed prior to adding a predicted picture to obtain a Reconstructedlmage. The predicted picture comes from upsampling the Reconstructedlmage belonging to the previous LoQ. (The Initial Cycle differs since it does not add any predicted picture.) Figure 11 shows schematically 1100 that higher LoQs are predicted from lower LoQs and are then corrected using Residuals. The process leading from one LoQ to its successor, or leading to an Initial LoQ, is called a Cycle. In Figure 11 the use of Tiers of Tesserae, discussed elsewhere, to decode Residuals is not emphasized, but is represented by the rectangles entitled “De-sparsification” and “Entropy Decoding”. Every successor LoQ, LoQ -n+1, has a Reconstructedlmage that is generated from an upsampled version of the Reconstructedlmage from LoQ -n (the prediction data) and from decoding and applying a Composition transform to additional Residual data. Appropriate Residual data is contained in the Residual Surfaces connected with the successor LoQ. This recursive approach is called Pyramidal Reconstruction. To initiate the process, the Initial LoQ is decoded first. (Note: The highest Level of Quality is LoQ 0, the next highest LoQ -1, etc.). Within an LoQ, the Composition Transform allows a single ComposedResidualArray of Pels to be recovered from 4 Residual Surfaces of transformed Pels. Without loss of generality, examples will now be described in which a number of autonomous agents (such as drones and / or agents) cooperate to make sense of an environment. Each autonomous agent has its own sensor(s). However, each autonomous agent can also communicate with other autonomous agents, for example nearby autonomous agents, to leverage their perspectives of the same environment. This ultimately allows a better reconstruction of the environment. In other words, a collective of sensors has a better perspective (for example, view) of the environment than a single sensor. Each autonomous agent may query other autonomous agents to determine what they can sense. Such communication facilitates sensor fusion. For example, a first autonomous agent may have a left-side and front view of an object, and a second autonomous agent may have a front and right-side view of the object. The first autonomous agent may request the right-side view of the object from the second autonomous agent to get a better overall view of the object. Similarly, the first autonomous agent may provide the second autonomous agent with the left-side view. The first and second autonomous agents may therefore sensor fuse to obtain a combined view that incorporates both the views of the first autonomous agent and the views of the second autonomous agent. In accordance with examples, a hierarchical data representation with partial recall and partial decoding may allow each autonomous agent to request the type of signal and / or the area(s) of that signal that are most important to that autonomous agent. In this manner, autonomous agents may not transmit all sensor data to all other autonomous agents. A partial file recall may be performed, when appropriate, at a maximum level of detail, and of just the regions of interest (ROIs) of one or more particular signals, to create a better comprehension of what a particular autonomous agent is trying to understand. In such examples, and as will be explained in more detail below, each autonomous agent may keep all of the signals it captures at the maximum level of quality. However, it may transmit only a fraction of those signals, at lower levels of quality, according to the needs of the other autonomous agents in a group. This may involve the raw output of the sensor(s) and / or the processed outcome of what the autonomous agent understands from the sensor(s). For example, each autonomous agent may recreate a volumetric signal based on its available information. Each autonomous agent may then perfect that reconstruction based on the reconstructions of other autonomous agents that potentially have a better perspective on certain objects or items in the environment. Use of hierarchical data formats that can be heavily parallelised and that support partial decoding is particularly effective. For example, the captured signals may be represented in a hierarchical-compatible and ROI-compatible manner, such that devices may communicate efficiently and effectively between themselves. Signals may be communicated at lower resolutions, improving transmission latency and bandwidth which may be heavily constrained in some example deployments. Signals may generally be kept in compressed form, reconstructing only ROIs at higher qualities. Such examples may be applicable to videogrammetry. For example, volumetric capture may currently be performed using a large number of sensors, for example fifty cameras. Crossing may be performed for each pixel of each image from each camera. Examples described herein that leverage hierarchical coding and hierarchical ROIs may involve fewer cameras. Additionally, artificial intelligence (AI) may be used in place of some of the cameras. For example, AI may be used to resolve conflicts only in regions in which there is particularly important detail. For example, it may be possible not to fuse all signals from all cameras where processing is performed hierarchically in this manner. Crossing a hierarchical ROI in videogrammetry may be more efficient than current videogrammetry. For example, using fifty frames of eight-megapixel images may involve significant processing. Using fewer frames of two-megapixel images, where ROIs of those images can be upscaled (for example to eight megapixels) where appropriate, can be more efficient. Such increased efficiency may be in terms of random access memory (RAM) usage and availability, memory bandwidth bottlenecks, etc. In contrast, by keeping data in compressed form, only expanding ROIs and performing videogrammetry only on the ROIs, memory bandwidth may be saved, and fewer views (for example three or four) may be sufficient. This may simplify the process of generating a sensor-fused representation by an order of magnitude. Referring to Figure 12, there is shown a schematic diagram of an example of a system 1200. The example system 1200 comprises a plurality of devices 1202, 1204, 1206. In this specific example, the example system 1200 comprises three devices 1202, 1204, 1206. However, such a system may comprise more or fewer devices in other examples. In this example, the devices 1202, 1204, 1206 comprise autonomous devices. Autonomous devices may operate at one or more levels of autonomy. In this specific example, the autonomous devices 1202, 1204, 1206 comprise autonomous agents in the form of UAVs, which are also known as “drones”. However, the autonomous devices 1202,1204,1206 may take other forms in other examples. For instance, the autonomous devices 1202, 1204,1206 may comprise autonomous agents in the form of robots. Thus, the system 1200 may comprise a swarm of robots and / or drones. In this specific example, the devices 1202, 1204, 1206 are all the same type of device as each other, i.e. they are all UAVs. In other examples, some or all of the devices 1202, 1204, 1206 may be different types of devices. For example, the devices 1202, 1204, 1206 may comprise a UAV, a robot and a ground-mounted camera. More generally, the devices 1202, 1204, 1206 may take various different forms including, but not limited to, UAVs, robots, fixed-location cameras, mobile-location cameras, servers, smartphones, tablet computing devices, headsets and so on. In this example, each device 1202, 1204, 1206 comprises one or more sensors 1208, 1210, 1212. In this specific example, the sensors 1208, 1210, 1212 comprise cameras. Thus, in this specific example, each UAV 1202, 1204, 1206 comprises a camera 1208, 1210, 1212. The sensors 1208, 1210, 1212 may take other forms in other examples. For example, the sensors 1208, 1210, 1212 may comprise one or more microphones and / or one or more thermal sensors. Each sensor 1208, 1210, 1212 has a respective perspective 1214, 1216, 1218 of an environment in which the system 1200 is located. In this specific example in which the sensors 1208, 1210, 1212 comprise cameras, each camera 1208, 1210, 1212 has a respective view 1214, 1216, 1218 of the environment. A view of an environment may be considered to be a visual perspective of an environment. Another type of perspective is a sonic perspective. For example, multiple microphones may capture different sonic perspectives of an environment. In this example, the environment comprises an object 1220. In this specific example, the object 1220 comprises a person. However, the object 1220 may take a different form in other examples. An environment may comprise more than one object 1220. All of the devices 1202, 1204, 1206 may have perspectives on the same object(s), or some of the devices 1202, 1204, 1206 may have perspectives on different objects. Referring to Figure 13, there is shown a schematic diagram of another example of a system 1300. The example system 1300 shown in Figure 13 corresponds generally to the system 1200 shown in Figure 12. Reference signs used in Figure 13 are the same as those used in Figure 12 for the same or similar features, but incremented by 100. For ease of understanding, the perspectives 1214, 1216, 1218 shown in Figure 12 are not reproduced in Figure 13. In this specific example, each device 1302, 1304, 1306 can communicate with each other device 1302, 1304, 1306. For instance, each device 1302, 1304, 1306 may be able to communicate with each other device 1302, 1304, 1306 directly. Such direct communication may be via Wi-Fi™, Bluetooth™ or in another manner. Each device 1302, 1304, 1306 may be able to communicate with each other device 1302, 1304, 1306 indirectly. Such indirect communication may be via a ground-based relay, a cellular network or in another manner. More specifically, in this example, a first device 1302 can communicate with a second device 1304 via a first communications channel 1322, the second device 1304 can communicate with a third device 1306 via a second communications channel 1324, and the third device 1306 can communicate with the first device 1302 via a third communications channel 1326. Referring to Figure 14, there is shown a schematic diagram of another example of a system 1400. The example system 1400 shown in Figure 14 corresponds generally to the system 1300 shown in Figure 13. Reference signs used in Figure 14 are the same as those used in Figure 13 for the same or similar features, but incremented by 100. In this specific example, the first device 1402 can communicate with the second device 1404 via the first communications channel 1422, and the second device 1404 can communicate with the third device 1406 via the second communications channel 1424. However, in this example, a communication channel (corresponding to the third communication link 1326 in the example system 1300) is not present between the first and third devices 1402, 1406. In this specific example, the devices 1402, 1404, 1406 communicate with each other via the second device 1404. The second device 1404 may therefore be considered to be a master device. Referring to Figure 15, there is shown a schematic diagram of another example of a system 1500. The example system 1500 shown in Figure 15 corresponds generally to the system 1400 shown in Figure 14. Reference signs used in Figure 15 are the same as those used in Figure 14 for the same or similar features, but incremented by 100. In addition to the devices 1502,1504, 1506, the example system 1500 comprises a further device 1528. In this example, the devices 1502, 1504, 1506 are the same type of device as each other (in this specific example, UAVs), and the further device 1528 is a different type of device. More specifically, in this specific example, the further device 1528 comprises a server. Although, in this specific example, the devices 1502, 1504, 1506 are the same type of device as each other, they may comprise different types of device in other examples. In this example, the devices 1502, 1504, 1506 are not able to communicate with each other. Instead, in this example, the first device 1502 can communicate with the server 1528 via a first communications channel 1530, the second device 1504 can communicate with the server 1528 via a second communications channel 1532, and the third device 1506 can communicate with the server 1528 via a third communications channel 1534. Referring to Figure 16, there is shown a schematic diagram of an example of a signal processing system 1600. The example signal processing system 1600 comprises three devices 1602, 1604, 1606 corresponding to the three devices described above with reference to Figures 12 to 15. Different numbers and / or types of device may be used in other examples. The first device 1602 obtains a signal 1636. The first device 1602 may obtain the signal 1636 by capturing the signal 1636 with a sensor (not shown) of the first device 1602. However, the first device 1602 may obtain the signal 1636 in other ways. For example, the first device 1602 may receive the signal 1636, may retrieve the signal 1636 from memory, and so on. In this example, the first device 1602 stores the signal 1636 in memory. In this example, the signal 1636 is at a highest level of quality in a tiered hierarchy with multiple, different levels of quality. In this specific example, the tiered hierarchy comprises three different levels of quality. The highest, middle, and lowest levels of quality may be referred to as the “maximum”, “medium”, and “minimum” levels of quality respectively, or the like. In this example, the second device 1604 obtains a signal 1638. In this example, the second device 1604 stores the signal 1638 in memory. In this example, the second device 1604 identifies a sub-region 1640 of the signal 1638. The second device 1604 may identify the sub-region 1640 of the signal 1638 in various different ways. For example, the first device 1602 may identify the subregion 1640 to the second device 1604, the first device 1602 may identify an object to the second device 1604, the second device 1604 may identify that an object is in a subregion 1640 and / or otherwise. In this example, the second device 1604 downscales only the sub-region 1640 of the signal 1638 to obtain a sub-region 1642 of a downscaled rendition 1644 of the signal 1638. In other examples, the second device 1604 may downscale the signal 1638 to obtain the downscaled rendition 1644 of the signal 1638. In such other examples, the second device 1604 may select the sub-region 1642 of the downscaled rendition 1644 or may use the downscaled rendition 1644. In this example, the second device 1604 downscales only the sub-region 1642 of the downscaled rendition 1644 of the signal 1638 to obtain a sub-region 1646 of a further-downscaled rendition 1648 of the signal 1638. In other examples, the second device 1604 may downscale the downscaled rendition 1644 of the signal 1638 to obtain the further-downscaled rendition 1648 of the signal 1638. In such other examples, the second device 1604 may select the sub-region 1646 of the further-downscaled rendition 1648 or may use the further-downscaled rendition 1648. In this example, the first device 1602 obtains only the sub-region 1646 of the further-downscaled rendition 1648 of the signal 1638 from the second device 1604 via a communications channel 1650. In other examples, the first device 1602 obtains the further-downscaled rendition 1648 of the signal 1638 from the second device 1604 via the communications channel 1650. Processing at a third device 1606 corresponds to that at the second device 1604. For convenience and brevity, a detailed explanation of such processing is not repeated here. The position(s) of the sub-region(s) selected by the second and third devices 1604, 1606 may be different from each other, as depicted schematically in Figure 16. Referring to Figure 17, there is shown a schematic diagram of an example of a signal processing system 1700. The example signal processing system 1700 comprises one device 1702 corresponding to the first device described above with reference to Figures 12 to 16. In this example, the device 1702 has a signal 1736 corresponding to the signal 1636 described above with reference to Figure 16. Additionally, the device 1702 has a further-downscaled signal 1748 corresponding to the further-downscaled signal 1648 described above with reference to Figure 16, or has a sub-region 1746 thereof corresponding to the sub-region 1646 described above with reference to Figure 16. The device 1702 also has another further-downscaled signal or sub-region thereof (for example, received from the third device 1606 described above with reference to Figure 16). For convenience and brevity, a detailed explanation of the other further downscaled signal or sub-region thereof is not repeated here. In this example, the device 1702 upscales the further-downscaled signal 1748 or the sub-region 1746 thereof. If the device 1702 obtains the further-downscaled signal 1748, the device 1702 may upscale the further-downscaled signal 1748 or just the subregion 1746 thereof. The upscaling generates an upscaled signal 1744 corresponding to the downscaled signal 1644 described above with reference to Figure 16 or a sub-region 1742 thereof corresponding to the sub-region 1642 described above with reference to Figure 16. Where the upscaling generates the upscaled signal 1744, the device 1702 may select just the sub-region 1742 thereof for further use. In this example, the device 1702 upscales the upscaled signal 1744 or the subregion 1742 thereof to generate an upscaled signal 1738 corresponding to the signal 1638 described above with reference to Figure 16 or a sub-region 1740 thereof corresponding to the sub-region 1640 described above with reference to Figure 16. Where the upscaling generates the upscaled signal 1738, the device 1702 may select just the sub-region 1740 thereof for further use. In this example, the device 1702 combines the signal 1736, and all of part of the other upscaled signals, including the upscaled signal 1738 or the sub-region 1740 thereof, using a combiner 1752. Thus, with reference to Figures 16 and 17, the device 1602, 1702 may implement a method of controlling the device 1602, 1702. A first signal 1636, 1736 representing a first perspective of an environment may be obtained. The first perspective may be that of the device 1602, 1702. The first signal 1636, 1736 has been captured using afirst sensor. The device 1602, 1702 may comprise the first sensor (for example a camera), and the first signal 1636, 1736 may be obtained by capturing the first signal 1636, 1736 using the first sensor. A downscaled version 1646, 1648, 1746, 1748 of all or part of a second signal 1638, 1738 representing a second, different perspective of the environment may be received. The second perspective may be that of another device 1604. The second signal 1638, 1738 has been captured using a second sensor. The other device 1604 may comprise the second sensor, and receiving the downscaled version 1646, 1648, 1746, 1748 may comprise receiving the downscaled version 1646, 1648, 1746, 1748 from the other device 1604. The device 1602, 1702 and the other device 1604 may be the same type of device as each other. Alternatively, the device 1602, 1702 and the other device 1604 may be different types of device. The device 1602, 1702 and / or the other device 1604 may be, or may comprise, an autonomous device. The device 1602, 1702 and / or the other device 1604 may be, or may comprise, a UAV. The device 1602, 1702 and / or the other device 1604 may be, or may comprise, a robot. As will be explained in more detail below, the device 1602, 1702 may transmit a signal request to the other device 1604. The downscaled version 1646, 1648, 1746, 1748 may be received in response to the signal request. The signal request may identify a type of the downscaled version 1646,1648, 1746, 1748 and / or the second signal 1638, 1738. The signal request may identify an ROI in the environment. The downscaled version 1646, 1746 may correspond to the ROI. The downscaled version 1646, 1746 may be a downscaled version of only part of the second signal 1638, 1738. Alternatively, the downscaled version 1646, 1648, 1746, 1748 may be a downscaled version the second signal 1638, 1738. All or part of the downscaled version 1646, 1648, 1746, 1748 may be upscaled to generate an upscaled version 1742, 1744, 1738, 1740 of the downscaled version 1646, 1648, 1746, 1748. Such upscaling may comprise upscaling multiple times to reach a target level of quality in a tiered hierarchy having multiple different levels of quality. The target level of quality may be a maximum level of quality in the tiered hierarchy. The upscaled version 1742, 1744, 1738, 1740 may be used. The first signal 1636, 1736 may also be used. Using the first signal 1636, 1736 and using the upscaled version 1742, 1744, 1738, 1740 may comprise using the first signal 1636, 1736 and the upscaled version 1742, 1744, 1738, 1740 to generate (for example, using the combiner 1752) a combined signal representing a combined perspective of the environment. A set of residuals associated with the second signal 1638, 1738 may be received. The set of residuals may be used to generate the combined signal. A downscaled version of all or part of a third signal representing a third, different perspective of the environment may be obtained. The third perspective may be that of the third device 1606. The third signal has been captured using a third sensor (for example a camera of the third device 1606). All or part of the downscaled version of all or part of the third signal may be upscaled to generate an upscaled version of the downscaled version of all or part of the third signal. The upscaled version of the downscaled version of all or part of the third signal may be used to generate (for example, using the combiner 1752) the combined signal representing the combined perspective of the environment. Using the first signal 1636, 1736 may comprise analysing the first signal 1636, 1736. The downscaled version 1646, 1648, 1746, 1748 of all or part of the second signal 1638, 1738 may be received based on the analysing. Such analysing may comprise object detection and / or object recognition. The downscaled version 1646, 1648, 1746, 1748 may be analyed. The downscaled version 1646, 1648, 1746, 1748 may be selected for upscaling, in preference to a downscaled version of all or part of another signal representing another, different perspective of the environment, based on the analysing. The upscaling may be in response to the selecting. The device 1602, 1702 may comprise memory and the first signal 1636, 1736 may be stored in the memory of the device 1602, 1702. The other device 1604 may comprise memory and the second signal 1638 may be stored in the memory of the other device 1604. The first and / or second signal 1636, 1736, 1638, 1738 may comprise raw output of the first and / or second sensor respectively. The first and / or second signal 1636, 1736, 1638, 1738 may comprise comprises a processed version of raw output of the first and / or second sensor respectively. The first and / or second signal 1636, 1736, 1638, 1738 may comprise an image signal, a video signal, and / or a volumetric signal. The first and second signal 1636, 1736, 1638, 1738 may be the same type of signal as each other. Alternatively, the first and second signal 1636, 1736, 1638, 1738 may be different types of signal. Another method of controlling a device 1604 comprising a sensor is provided. A signal 1638 representing a perspective of an environment is captured using the sensor. A downscaled version 1642, 1644, 1646, 1648 of all or part of the captured signal 1638 is generated. An upscaled version of the downscaled version 1642, 1644, 1646, 1648 is generated. A set of residuals is generated using the captured signal 1638 and the upscaled version. The downscaled version and the set of residuals are transmitted to another device 1602, 1602. Reference is made in this connection to the residual generation and use described above with reference to Figures 2A and 2B. Another method of controlling a first device 1602, 1702 comprising a first sensor is provided. A first signal 1636, 1736 representing a first perspective of an environment is obtained. The first signal has been captured using the first sensor. A downscaled version 1646, 1648, 1746, 1748 of all or part of a second signal 1638, 1738 representing a second, different perspective of the environment is received from a second device 1604 comprising a second sensor. The second signal has been captured using the second sensor. Another method of controlling a first device 1602, 1702 comprising a first sensor is provided. A first signal 1636, 1736 representing a first perspective of an environment is obtained. The first signal has been captured using the first sensor. The first signal 1636, 1736 is analysed. All or part of a second signal 1638, 1640 representing a second, different perspective of the environment is requested based on the analysing and from a second device 1604. The second signal 1638 has been captured using a second sensor. The second device 1604 comprises the second sensor. Apparatus configured to perform such methods is also provided, for example in the form of the devices described herein. A computer program configured to perform such methods is also provided. Referring to Figure 18, there is shown a schematic diagram of an example of a signal processing system 1800. The example signal processing system 1800 comprises two devices 1802, 1804, corresponding to the two of devices described above with reference to Figures 12 to 17. Different numbers and / or types of device may be used in other examples. In this example, the first device 1802 transmits a signal request 1858 to the second device 1804. In some examples, the signal request 1858 identifies a type of the downscaled version and / or second signal. In some examples, the signal request 1858 identifies a region of interest in the environment, and the downscaled version corresponds to the region of interest. The signal request may request the second device 1804 to identify any perspectives of the environment and / or one or more objects in the environment that the second device 1804 can represent in one or more signals. In this example, the second device 1804 transmits a response 1860 to the first device 1802. The response 1860 is in response to the signal request 1858. In other examples, rather than the response 1860 being in response to the signal request 1858, the second device 1804 may transmit content from the response 1860 intermittently, periodically, or otherwise. For example, the second device 1804 may regularly broadcast such information. Referring to Figure 19, there is shown a schematic diagram of an example of a signal processing system 1900. The example signal processing system 1900 comprises two devices 1902, 1904, corresponding to the two of devices described above with reference to Figures 12 to 18. Different numbers and / or types of device may be used in other examples. In this example, the first device 1902 transmits a signal information request 1962 to the second device 1904. In this example, the second device 1904 transmits a response 1964 to the first device 1902. The response 1960 is in response to the signal information request 1962. In other examples, rather than the response 1964 being in response to the signal information request 1962, the second device 1904 may transmit content from the response 1964 intermittently, periodically, or otherwise. For example, the second device 1904 may regularly broadcast such information. The signal information request 1962 may differ from the signal request 1858 described above with reference to Figure 18 in that the signal request 1858 may request a signal and / or a region of a signal, whereas the signal information request 1962 may request information about the signals and / or regions of signal available from the second device 1804. For example, a signal information request may be transmitted to solicit information about available signals and / or regions of signal and a signal request may be transmitted based on a response to the signal information request. The signal information request 1962 may request information concerning quality and / or objects detected in a signal. In a further example, a device obtains a signal representing a perspective of an environment. The signal may comprise video, for example. The device may obtain supplemental information from another device. The device may process the supplemental information. The supplemental information may supplement the signal obtained by the device. The supplemental information may comprise a signal or any other type of supplemental information. The device may receive the supplemental information from one or more further devices. The signal(s) and / or supplemental information may be hierarchically structured. Using hierarchically structured data in this manner can be more efficient than using other data. For instance, a device could obtain low-quality versions of a target ROI from three other devices. The device may then determine which of those versions is most preferred and request higher levels of quality (LOQs) from the appropriate other device. This example may not involve using the signal obtained by the device to generate a combined perspective of the environment. Instead, the device may use the signal obtained by the device to identify one or more ROIs and may obtain high-quality versions of the ROI(s) from other devices that may have different perspectives of the ROI(s). In some examples, the device obtains a signal representing a perspective of an environment. The device obtains one or more further signals representing one or more further perspectives of the environment. The device then requests at least one further LOQ for at least one of the one or more further perspectives. This differs from an ROI being selected and a further LOQ being requested for that ROI. In some examples, a signal is captured. For example, the signal may be captured at a device. Based on an analysis of the signal, one or more signals may be requested from one or more further devices. Referring to Figure 20, there is shown a schematic block diagram of an example of an apparatus 2000. In an example, the apparatus 2000 comprises an encoder. In another example, the apparatus 2000 comprises a decoder. In other examples, the apparatus 2000 comprises neither an encoder nor a decoder but is configured to communicate with an encoder and / or a decoder. Examples of apparatus 2000 include, but are not limited to, a mobile computer, a personal computer system, a wireless device, base station, phone device, desktop computer, laptop, notebook, netbook computer, mainframe computer system, handheld computer, workstation, network computer, application server, storage device, a consumer electronics device such as a camera, camcorder, mobile device, video game console, handheld video game device, an XR headset, or in general any type of computing or electronic device. In this example, the apparatus 2000 comprises one or more processors 2001 configured to process information and / or instructions. The one or more processors 2001 may comprise a CPU. The one or more processors 2001 are coupled with a bus 2002. Operations performed by the one or more processors 2001 may be carried out by hardware and / or software. The one or more processors 2001 may comprise multiple colocated processors or multiple disparately located processors. In this example, the apparatus 2000 comprises computer-useable volatile memory 2003 configured to store information and / or instructions for the one or more processors 2001. The computer-useable volatile memory 2003 is coupled with the bus 2002. The computer-useable volatile memory 2003 may comprise random access memory (RAM). In this example, the apparatus 2000 comprises computer-useable non-volatile memory 2004 configured to store information and / or instructions for the one or more processors 2001. The computer-useable non-volatile memory 2004 is coupled with the bus 2002. The computer-useable non-volatile memory 2004 may comprise read-only memory (ROM). In this example, the apparatus 2000 comprises one or more data-storage units 2005 configured to store information and / or instructions. The one or more data-storage units 2005 are coupled with the bus 2002. The one or more data-storage units 2005 may for example comprise a magnetic or optical disk and disk drive or a solid-state drive (SSD). In this example, the apparatus 2000 comprises one or more input / output (I / O) devices 2006 configured to communicate information to and / or from the one or more processors 2001. The one or more I / O devices 2006 are coupled with the bus 2002. The one or more I / O devices 2006 may comprise at least one network interface. The at least one network interface may enable the apparatus 2000 to communicate via one or more data communications networks. Examples of data communications networks include, but are not limited to, the Internet and a Local Area Network (LAN). The one or more EO devices 2006 may enable a user to provide input to the apparatus 2000 via one or more input devices (not shown). The one or more input devices may include for example a remote control, one or more physical buttons etc. The one or more I / O devices 2006 may enable information to be provided to a user via one or more output devices (not shown). The one or more output devices may for example include a display screen. Various other entities are depicted for the apparatus 2000. For example, when present, an operating system 2007, image processing module 2008, one or more further modules 2009, and data 2010 are shown as residing in one, or a combination, of the computer-usable volatile memory 2003, computer-usable non-volatile memory 2004 and the one or more data-storage units 2005. The data signal processing module 2008 may be implemented by way of computer program code stored in memory locations within the computer-usable non-volatile memory 2004, computer-readable storage media within the one or more data-storage units 2005 and / or other tangible computer-readable storage media. Examples of tangible computer-readable storage media include, but are not limited to, an optical medium (e g., CD-ROM, DVD-ROM or Blu-ray), flash memory card, floppy or hard disk or any other medium capable of storing computer-readable instructions such as firmware or microcode in at least one ROM or RAM or Programmable ROM (PROM) chips or as an Application Specific Integrated Circuit (ASIC). The apparatus 2000 may therefore comprise a data signal processing module 2008 which can be executed by the one or more processors 2001. The data signal processing module 2008 can be configured to include instructions to implement at least some of the operations described herein. During operation, the one or more processors 2001 launch, run, execute, interpret or otherwise perform the instructions in the signal processing module 2008. Although at least some aspects of the examples described herein with reference to the drawings comprise computer processes performed in processing systems or processors, examples described herein also extend to computer programs, for example computer programs on or in a carrier, adapted for putting the examples into practice. The carrier may be any entity or device capable of carrying the program. It will be appreciated that the apparatus 2000 may comprise more, fewer and / or different components from those depicted in Figure 20. The apparatus 2000 may be located in a single location or may be distributed in multiple locations. Such locations may be local or remote. The techniques described herein may be implemented in software or hardware, or may be implemented using a combination of software and hardware. They may include configuring an apparatus to carry out and / or support any or all of techniques described herein. It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims. Key for Figure 9 • Tier 0 Top tier in the coding structure - used for residuals 5 • Tier -m Middle tiers in the coding structure - used for metadata for next level tesserae • Root Tier 10 Bottom tier in the coding structure - used for metadata for next level tesserae • Summit Maximum grid that could be encoded by an S-Tree 15 • Active Volume Volume of the sparse coding structure corresponding to the area occupied by the image LD • Surface CM 20 Fraction of the active volume in layer 0 - e.g., corresponding to the image above

Claims

1. A method of controlling a device, the method comprising:obtaining a first signal representing a first perspective of an environment, the 5 first signal having been captured using a first sensor;receiving a downscaled version of all or part of a second signal representing a second, different perspective of the environment, the second signal having been captured using a second sensor;upscaling all or part of the downscaled version to generate an upscaled version10 of the downscaled version; andusing the upscaled version.

2. A method according to claim 1, comprising using the first signal.15 3. A method according to claim 2, wherein using the first signal and using theupscaled version comprises:using the first signal and the upscaled version to generate a combined signal representing a combined perspective of the environment.20 4. A method according to claim 3, comprising:receiving a set of residuals associated with the second signal,wherein the set of residuals is used to generate the combined signal.

5. A method according to claim 3 or 4, comprising:25 obtaining a downscaled version of all or part of a third signal representing athird, different perspective of the environment, the third signal having been captured using a third sensor; andupscaling all or part of the downscaled version of all or part of the third signal to generate an upscaled version of the downscaled version of all or part of the third 30 signal,wherein the upscaled version of the downscaled version of all or part of the third signal is used to generate the combined signal representing the combined perspective of the environment.5 6. A method according to any of claims 2 to 5, wherein using the first signalcomprises analysing the first signal, and wherein the downscaled version is received based on the analysing.

7. A method according to any of claims 1 to 6, wherein the device comprises the 10 first sensor, and wherein obtaining the first signal comprises capturing the first signal using the first sensor.

8. A method according to any of claims 1 to 7, wherein the downscaled version is a downscaled version of only part of the second signal.

9. A method according to any of claims 1 to 8, wherein another device comprises the second sensor, and wherein receiving the downscaled version comprises receiving the downscaled version from the other device.20 10. A method according to claim 9, wherein the device and the other device are thesame type of device as each other.

11. A method according to claim 9 or 10, comprising transmitting a signal request to the other device, wherein the downscaled version is received in response to the signal25 request.

12. A method according to claim 11, wherein the signal request identifies a type of the downscaled version and / or the second signal.

13. A method according to claim 11 or 12, wherein the signal request identifies a region of interest in the environment, and wherein the downscaled version corresponds to the region of interest.5 14. A method according to any of claims 1 to 13, comprising:analysing the downscaled version of all or part of the second signal; andselecting the downscaled version of all or part of the second signal for upscaling, in preference to a downscaled version of all or part of another signal representing another, different perspective of the environment, based on the analysing,10 wherein the upscaling is in response to the selecting.

15. A method according to any of claims 1 to 14, wherein the device is an autonomous device.15 16. A method according to any of claims 1 to 15, wherein the device is an unmannedaerial vehicle.17.A method according to any of claims 1 to 15, wherein the device is a robot.20 18. A method according to any of claims 1 to 17, wherein the first and secondsignals comprise first and second images respectively, wherein the first and second perspectives comprise first and second views respectively, and wherein the first and second sensors comprise first and second cameras respectively.25 19. A method according to any of claims 1 to 18, wherein upscaling all or part ofthe downscaled version comprises upscaling multiple times to reach a target level of quality in a tiered hierarchy having multiple different levels of quality.

20. A method according to claim 19, wherein the target level of quality is a 30 maximum level of quality in the tiered hierarchy.04 11 2521. A method according to claim 19 or 20, wherein the first signal and the upscaled version have the same level of quality as each other in the tiered hierarchy.

22. A method according to any of claims 1 to 21, wherein the device comprises 5 memory and wherein the method comprises storing the first signal in the memory.

23. A method according to any of claims 1 to 22, wherein the first and / or second signal comprises raw output of the first and / or second sensor respectively.10 24. A method according to any of claims 1 to 23, wherein the first and / or secondsignal comprises a processed version of raw output of the first and / or second sensor respectively.

25. A method according to any of claims 1 to 24, wherein the first and / or second 15 signal comprises:an image signal;a video signal; and / or a volumetric signal.20 26. A method according to claim 25, wherein the video signal comprises H.264encoded video or H.265 encoded video.

27. A method according to any of claims 1 to 26, wherein the first and second signal are the same type of signal as each other.2528. Apparatus configured to perform a method according to any of claims 1 to 27.

29. A computer program configured to perform a method according to any of claims1 to 27.

Citation Information

Patent Citations

  • Method for producing high dynamic range images

    US20120105681A1

  • Stereoscopic Capture Using Cameras with Different Fields of View

    US20240406361A1