Bitstream

The bitstream format addresses the high processing and storage demands of three-dimensional representations by defining points with attributes and capture device information, enhancing transmission efficiency and reducing file sizes.

GB2703234APending Publication Date: 2026-07-22V NOVA INT LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
V NOVA INT LTD
Filing Date
2024-12-23
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Three-dimensional representations require substantial processing power and large file sizes, necessitating significant storage and bandwidth for transmission.

Method used

A bitstream format is developed to define points of a three-dimensional representation, including attributes, capture device information, and distance from the device, enabling efficient encoding and decoding of three-dimensional data.

Benefits of technology

The bitstream format reduces processing requirements and file sizes, facilitating efficient transmission and storage of three-dimensional data while maintaining image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A bitstream that defines a viewing zone 1 of one or more three-dimensional representations of a scene, the bitstream defining a location and / or a size of the viewing zone. preferably including inform
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Disclosure The present disclosure relates to methods, systems, and apparatuses for determining, transmitting, encoding, and decoding a bitstream. Background to the Disclosure Three-dimensional representations of environments are used in many contexts, including forthe generation of virtual reality videos, in which depth information for a plurality of points of the representation is used to generate different images for a left eye and a right eye of a user. Typically, substantial processing power is required to determine such a three-dimensional representation, and the file size of files associated with these representations is typically large so that substantial amounts of storage are needed to keep the files and substantial amounts of bandwidth are required to transfer the files. Summary of the Disclosure According to an aspect of the present disclosure, there is described a bitstream for defining a point of a three-dimensional representation. According to an aspect of the present disclosure, there is described a bitstream for defining a texture point of a three-dimensional representation. According to an aspect of the present disclosure, there is described a bitstream for defining a movement vector of a three-dimensional representation. According to an aspect of the present disclosure, there is described a bitstream for defining one or more of, or all of, a point, a texture point, and / or a movement vector of a three-dimensional representation. According to an aspect of the present disclosure, there is described a bitstream for defining a point of a three-dimensional representation, the definition of the point comprising: an attribute field indicating an attribute of the point; a capture device field indicating a capture device associated with the point; and a distance field indicating a distance of the point from the capture device. Preferably, the bitstream comprises (e.g. defines) a plurality of points. Preferably, the bitstream comprises, for each of the points: an attribute field indicating an attribute of said point; a capture device field indicating a capture device associated with said point; and a distance field indicating a distance of said point from the capture device. Preferably, the points are associated with different capture devices. Preferably, the definition of the point comprises an angle field indicating an angle of the point from the capture device indicated in the capture device field. Preferably, the angle field indicates the angle as an index, preferably an index of an angular bracket associated with the capture device indicated in the capture device field. Preferably, angle field indicates the angle with reference to an angle step. Preferably, the angle field indicating an absolute angle of the point from the capture device indicated in the capture device field. Preferably, the bitstream comprises an indication of an angular step associated with one or more capture devices. Preferably, the bitstream comprises indications of plurality of different angular steps. Preferably, different capture devices are associated with different angular steps. Preferably, different capture devices are associated with different angular steps, wherein the capture device is associated with a plurality of different angular steps, wherein a first angular step between a first pair of angular brackets associated with the capture device differs from a second angular step between a second pair of angular brackets associated with the capture device. Preferably, the attribute field comprises a plurality of sub-field, preferably a plurality of sub-fields that indicate, respectively, a chrominance, a luminance, and a luma of the point. Preferably, the capture device field comprises an index of the capture device. Preferably, the bitstream comprises a table of capture devices. Preferably, the table comprises one or more of: an index for each of the capture devices; a location for each of the capture devices; a relationship between a pair of capture devices; an order of capture devices; and a priority value for one or more capture devices. Preferably, the capture device identifier indicates a viewing zone associated with the point. Preferably, the definition of the point comprises a further attribute field indicating a further attribute of the point. Preferably, the attribute is associated with a first eye of the user and the further attribute is associated with a second eye of a user. Preferably, the further attribute field signals a value of the further attribute as a difference from the attribute signalled in the attribute field. Preferably, the further attribute field signals a that is an exclusive or (XOR) function of the further attribute and the attribute signalled in the attribute field. Preferably, each of the attribute field and the further attribute field comprise a plurality of sub-fields. Preferably, each sub-field of the further attribute field signals a value of a component of the further attribute as a difference from a component of the attribute signalled in a corresponding attribute sub-field. Preferably, the attribute field and the further attribute field are arranged to be combinable in order to determine an attribute value. Preferably, the further attribute field is arranged to indicate whether the attribute and the further attribute are different. Preferably, a value of 0 in the further attribute field signals that the attribute and the further attribute have the same value. Preferably, the distance is signalled using an index. Preferably, the distance is signalled with reference to a quantisation level and / or a quantisation curve, preferably a 1 / z quantisation curve. Preferably, the distance is defined based on a set of quantisation levels, wherein a step between the levels in the set of quantisation levels increases with distance from a viewing zone associated with the point. Preferably, the quantisation curve is identified and / or defined in the bitstream. Preferably, the distance is signalled with reference to a container that encompasses the point. Preferably, the distance is signalled with reference to a quantisation level within the container. Preferably, the definition of the point comprises a size field indicating a size of the point. Preferably, the size field indicates a height and / or a width of the point. Preferably, the size field indicates a number of angular brackets covered by the point, preferably wherein the size field indicates a number of elevational brackets covered by the point and a number of azimuthal angular brackets covered by the point. Preferably, the size field indicates an underscan value of the point. Preferably, the size of the point is signalled as an index. Preferably, the size of the point is signalled as a value from a table of possible values. Preferably, an arrangement and / or a coverage of the point is dependent on a size and / or index signalled in the size field. Preferably, an arrangement of angular brackets covered by a point is defined by a value of the size field of the point. Preferably, a size of 3x1 in the size field indicates that the point covers a central angular bracket of a 3x1 arrangement; a size of 1x3 in the size field indicates that the point covers a central angular bracket of a 1x3 arrangement; and / or a size of 3x3 in the size field indicates that the point covers a central angular bracket of a 3x3 arrangement. Preferably, the size indicates each of a height, a width, and an underscan value. Preferably, an underscan value that is above a threshold indicates that the height and / or width of the point are limited. Preferably, an underscan value of 3 indicates that the height and width are each in the range [0,1]. Preferably, the bitstream comprises: a height value of the size field is limited to the range [0, 2]; and / or a width value of the size field is limited to the range [0, 2]; and / or an underscan value of the size field is limited to the range [0, 3], Preferably, the size field comprises a value that is equal to: size field=9*underscan value+3*height value+width. Preferably, the bitstream comprises a table associating sizes of points to coverages of points. Preferably, the size indicates a type of the point. Preferably, the bitstream comprises a table defining possible sizes for points of the bitstream. Preferably, each possible size is associated with a size index. Preferably, the definition of the point comprises a type field that indicates a type of the point. Preferably, the type field indicates whether the point is a translucent point. Preferably, the type field indicates whether the point is a reflective point. Preferably, the type field comprises a glass field that indicates whether the point is a point on a glass surface. Preferably, the definition of the point comprises a relative movement field that indicates whether the point moves along with a viewing zone of the three-dimensional representation. Preferably, the relative movement field indicates whether a movement vector of the point is defined relative to a movement vector of the viewing zone. Preferably, the relative movement field comprises a flag that indicates that a movement vector of the point is defined relative to a movement vector of the viewing zone. Preferably, the definition of the point comprises a viewing zone field indicating a viewing zone associated with the point. Preferably, the viewing zone field comprises an index. Preferably, the definition ofthe point comprises a normal field indicating a normal of the point. Preferably, the normal is signalled using octahedron normal encoding. Preferably, the normal is defined with reference to an angle field ofthe point. Preferably, the normal field defines a difference between the normal of a point and an angle signalled in an angle field ofthe point. Preferably, the normal is signalled as an index. Preferably, the definition of the point comprises a transparency field indicating a transparency of the point. Preferably, the transparency field is arranged to signal that a point is a texture point based on a value of the transparency field. Preferably, the definition of the point comprises a movement vector flag that indicates whether the point is associated with a movement vector. Preferably, the definition of the point comprises a movement vector field that indicates a movement vector associated with the point. Preferably, the movement vector indicates one or more of: a direction of the movement vector; and a magnitude of the movement vector. Preferably, the movement vector field identifies whether a single movement vector is used for each corner of the point, preferably wherein a value of 0x1 FFFFF in the movement vector field indicates that the same movement vector is used for each corner of the point. Preferably, the point is associated with a plurality of movement vectors, preferably wherein each corner of the point is associated with a different motion vector. Texture points Preferably, the attribute field defines a texture index that indicates an attribute value of the point based on a source external to the definition of the point. Preferably, the bitstream comprises: a first point with an attribute field; and that directly defines an attribute value of the first point; and a second point with an attribute field that defines a texture index of the second point. Preferably, the definition of the point comprises one or more indications that a point is a texture point that indicates an attribute value of the point based on a source external to the definition of the point. Preferably, the texture point is signalled based on a transparency value and / or a size of the texture point. Preferably, the texture point is signalled based on a transparency value that indicates the texture point is opaque. Preferably, the definition of the point comprises a texture index that indicates an attribute value of the point with reference to a texture atlas signalled in the bitstream. Preferably, the texture index is specified as a row-major tile index into the texture atlas. Preferably, the definition of the point comprises a texture index that indicates an attribute value of the point based on a source external to the definition of the point. Preferably, the definition of the point comprises a definition of the external source. Preferably, the texture point has a predetermined size. Preferably, the predetermined size is defined in the bitstream. Preferably, the texture point is arranged to signal a texture index for a first eye and a texture index for a second eye. Preferably, the texture point comprises a plurality of texture index fields, wherein each texture index field indicates a texture patch for a different eye. Preferably, the texture index for the second eye is arranged to be inferred based on the texture index for the first eye. Preferably, the texture index for the second eye is adjacent the texture index for the first eye. Preferably, the texture point comprises a stereo flag, the stereo flag indicating a relationship between a first texture index for a first eye and a second texture index for a second eye. Preferably, the stereo flag indicates whether the second texture index is different to the same index. Preferably, the stereo flag indicates that the relationship is one of: the texture patches for the left and right eyes being the same; the index of the texture patch for a second eye being one greater than the index of a texture patch for a first eye; the texture patch for a second eye being contained in a different texture atlas to the texture patch for a first eye; and the texture patch for a second eye being defined as a difference to the texture patch for a first eye. Preferably, the definition of the point comprises a bending orientation field that indicates a bending line of the point. Preferably, the bending orientation field comprises a flag, preferably a flag that indicates whether the bending line is a bottom-left to top-right diagonal or a bottom-right to top-left diagonal. Preferably, the bending orientation field comprises a diagonal orientation field that indicates a diagonal between corners of the point along which the point bends. Preferably, the definition of the point comprises a bend field indicating a bending of the point, preferably indicating a bending of the point along a bending line of the point. Movement vectors Preferably, the bitstream defines a movement vector of a three-dimensional representation, the definition of the movement vector comprising: a direction field indicating of the movement vector; and a magnitude field indicating a magnitude of the movement vector. According to another aspect of the present disclosure, there is described a bitstream for defining a movement vector of a three-dimensional representation, the definition of the movement vector comprising: a direction field indicating of the movement vector; and a magnitude field indicating a magnitude of the movement vector. Preferably, the movement vector is associated with a point signalled in the bitstream. Preferably, the movement vector comprises an identifier field identifying a point associated with the movement vector. Preferably, the movement vector is located adjacent a point in the bitstream, the movement vector being associated with the point. Preferably, a component of the direction and / or the magnitude of the movement vector is signalled in a definition of a point associated with the movement vector. Parts Preferably, the bitstream comprises a plurality of parts. Preferably, a first, non-opaque, part of the bitstream comprises non-opaque points and a second, opaque, part of the bitstream comprises opaque points. Preferably, the bitstream comprises a first point in the non-opaque part of the bitstream, the first point having a transparency field that indicates the point is opaque, wherein the transparency field indicates that the first point is a texture point. Preferably, a first, first capture device, part of the bitstream comprises points associated with a first capture device and a second, second capture device, part of the bitstream comprises points associated with a second capture device. Preferably, a first, point, part of the bitstream comprises definitions of points and a second, movement vector, part of the bitstream comprises definitions of movement vectors. Preferably, the bitstream comprises a bit indicating a change in a capture device, the bit preferably being located between a definition of a first point and a definition of a second point, the first point being associated with a first capture device and the second point being associated with a second capture device. Preferably, the bitstream comprises a capture device step that indicates a step between capture devices. Preferably, the bitstream comprises capture device step information that indicates a difference in the locations of a plurality of capture devices. Preferably, the bitstream comprises capture device point information that indicates a number of points associated with one or more capture devices. Preferably, the bitstream comprises information indicating an angular step associated with one or more capture devices. Preferably, a plurality of capture devices are associated with different angular steps, preferably the bitstream comprises correspondences that associate each capture device with a respective angular step. Preferably, the bitstream comprises information indicating an origin point of one or more three-dimensional representations. Preferably, the bitstream comprises an indication of a location of one or more viewing zones and / or dimensions of one or more viewing zones. According to another aspect of the present disclosure, there is described a bitstream for providing capture device information relating to one or more capture devices referenced by points of one or more three-dimensional representations. According to another aspect of the present disclosure, there is described a bitstream for defining a viewing zone of one or more three-dimensional representations. Preferably, the bitstream comprises: a definition of one or more viewing zones of the three-dimensional representations; and a definition of one or more points of the three-dimensional representations; wherein one or more of the points is defined with reference to a viewing zone. Preferably, the bitstream comprises an indication of viewing zone information for a plurality of three-dimensional representations. Preferably, the viewing zone information differs for the plurality of representations. Preferably, the bitstream comprises a plurality of segments, where each segment is associated with a different viewing zone and wherein each segment defines one or more points associated with that viewing zone. Preferably, the bitstream defines a plurality of viewing zones. Preferably, one or more viewing zones are associated with a plurality of three-dimensional representations, preferably a plurality of three-dimensional representations defined in the bitstream. Preferably, the bitstream comprises a movement vector associated with a viewing zone. Preferably, the bitstream comprises an indication of possible sizes of points defined in the bitstream. Preferably, the bitstream comprises an indication of arrangements of points defined in the bitstream, preferably comprising an association between sizes of points and an arrangement of those sizes, more preferably wherein the arrangement indicates an arrangement of angular brackets covered by a point of a given size. Preferably, the bitstream comprises information about one or more of a resolution; quality; and frequency of one or more three-dimensional representations defined in the bitstream. Preferably, the bitstream comprises information about a device arranged to view a three-dimensional representation defined in the bitstream and / or arranged to be compatible with the bitstream. Preferably, the bitstream comprises a texture atlas that contains attribute values for one or more points, wherein the texture atlas is referenced by the points. According to another aspect of the present disclosure, there is described a bitstream comprising a texture atlas that contains attribute values for one or more points, wherein the texture atlas is referenced by the points. Preferably, the texture atlas comprises one or more (e.g. a plurality of) texture patches, with each texture patch comprising a plurality of values, wherein the texture patch is arranged to indicate a plurality of attribute values associated with a point that references the texture patch. Preferably, the texture atlas comprises a two-dimensional image comprising a plurality of tiles. Preferably, the bitstream comprises a plurality of tiles of fixed size and / or of the same size. Preferably, a second texture patch in the texture atlas is defined with reference to a first texture patch in the texture atlas. Preferably, a second texture patch in the texture atlas is defined as a difference to a first texture patch in the texture atlas. Preferably, the or each texture atlas is associated with a single three-dimensional representation. Preferably, the texture atlas is associated with a plurality of three-dimensional representations. Preferably, the bitstream comprises a plurality of texture atlases, preferably wherein the texture atlases are located in the same part of the bitstream. Preferably, the bitstream comprises a plurality of texture atlases encoded as a video, wherein each frame of the video comprises a texture atlas. Preferably, the bitstream comprises one or more of: a point section comprising definitions of one or more points; a movement vector section comprising definitions of one or more movement vectors; and a texture atlas section comprising definitions of one or more texture atlases. Preferably, the bitstream comprises a plurality of three-dimensional representations. Preferably, the bitstream comprises points and / or movement vectors for a plurality of three-dimensional representations. Preferably, the bitstream comprises a plurality of segments, where each segment is associated with a different three-dimensional representation and wherein each segment comprises one or more of: a point section comprising definitions of one or more points in a corresponding three-dimensional representation; a movement vector section comprising definitions of one or more movement vectors associated with a corresponding three-dimensional representation; and a texture atlas section comprising definitions of one or more texture atlases. Preferably, the bitstream comprises a texture atlas segment comprising one or more texture atlases, preferably a plurality of texture atlases, preferably wherein the texture atlases are associated with a plurality of three-dimensional representations. Preferably, the bitstream comprises a header section comprising information about one or three-dimensional representations defined by the bitstream. Preferably, the bitstream comprises an indication of one or more of: a number of points defined in the bitstream; a number of movement vectors defined in the bitstream; a number of texture atlases defined in the bitstream; a correspondence between one or more texture atlases and one or more three-dimensional representations; a time step between a plurality of three-dimensional representations defined in the bitstream; and a time period covered by one or more three-dimensional representations defined in the bitstream. Preferably, the points are arranged to be rendered using a cubemap rendering process. Preferably, the points are arranged to be rendered using an ecospherical rendering process. Preferably, the bitstream is a bitstream for signalling one or more points of a three-dimensional representation. Preferably, the bitstream is a bitstream for signalling one or more points of a plurality of three-dimensional representations, the plurality of representations being arranged to provide a video of a scene. Preferably, the bitstream is a bitstream for signalling one or more points of a three-dimensional representation for providing a virtual reality and / or an augmented reality scene. According to another aspect of the present disclosure, there is described a device for generating the bitstream of any preceding claim, preferably wherein the device comprises an encoder. According to another aspect of the present disclosure, there is described a device for parsing the bitstream of any preceding claim, preferably wherein the device comprises a decoder. Preferably, the device is arranged to determine whether a point defined in the bitstream is a texture point and to determine a value of the point in dependence on whether the point is a texture point. According to another aspect of the present disclosure, there is described a method of generating a bitstream, the method comprising generating a bitstream for defining a point of a three-dimensional representation, the method comprising defining a point in the bitstream, the point comprising: an attribute field indicating an attribute of the point; a capture device field indicating a capture device associated with the point; and a distance field indicating a distance of the point from the capture device. According to another aspect of the present disclosure, there is described a method of parsing a bitstream, the method comprising parsing a bitstream defining a point of a three-dimensional representation, the method comprising: identifying, in the bitstream, an attribute field indicating an attribute of the point; identifying, in the bitstream, a capture device field indicating a capture device associated with the point; and identifying, in the bitstream, a distance field indicating a distance of the point from the capture device. Preferably, the method comprises determining whether the point is a texture point and determining the attribute field based on the point being a texture point. Preferably, the method comprises determining an attribute value for the point based on an external source based on the point being a texture point. According to another aspect of the present disclosure, there is described a method of generating a bitstream, the method comprising generating a bitstream for defining a movement vector of a three-dimensional representation, the movement vector comprising: a direction field indicating of the movement vector; and a magnitude field indicating a magnitude of the movement vector. According to another aspect of the present disclosure, there is described a method of parsing a bitstream, the method comprising parsing a bitstream defining a movement vector of a three-dimensional representation, the method comprising: identifying, in the bitstream, a direction field indicating of the movement vector; and a magnitude field indicating a magnitude of the movement vector. According to another aspect of the present disclosure, there is described a method of generating a bitstream, the method comprising generating a bitstream for defining a texture atlas for a three-dimensional representation, the texture atlas comprising one or more texture patches that are referenced by one or more points of the three-dimensional representation. According to another aspect of the present disclosure, there is described a method of parsing a bitstream, the method comprising parsing a bitstream defining a texture atlas of a three-dimensional representation, the method comprising: identifying a point in the three-dimensional representation that references a texture patch of the texture atlas; determining the texture patch from the texture atlas; and determining an attribute value for the point based on the texture patch. According to another aspect of the present disclosure, there is described a method of generating a bitstream, the method comprising generating a bitstream for defining a viewing zone for a three-dimensional representation. According to another aspect of the present disclosure, there is described a method of parsing a bitstream, the method comprising parsing a bitstream defining a viewing of a three-dimensional representation, the method comprising: extracting one or more parameters of the viewing zone from the bitstream. According to another aspect of the present disclosure, there is described a method for encoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described a method for decoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described a method for transmitting the aforesaid bitstream. According to another aspect of the present disclosure, there is described a method for transcoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described a method for receiving the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus and / or a device for encoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus and / or a device for decoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus and / or a device for transmitting the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus and / or a device for transcoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus and / or a device for receiving the aforesaid bitstream. According to another aspect of the present disclosure, there is described a computer-readable medium for storing the aforesaid bitstream. According to another aspect of the present disclosure, there is described a computer programme product for storing the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus (e.g. an encoder) for forming and / or encoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus (e.g. a decoder) for receiving and / or decoding the aforesaid bitstream. Any feature in one aspect of the disclosure may be applied to other aspects of the invention, in any appropriate combination. In particular, method aspects may be applied to apparatus aspects, and vice versa. Furthermore, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be construed accordingly. Any apparatus feature as described herein may also be provided as a method feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure, such as a suitably programmed processor and associated memory. It should also be appreciated that particular combinations of the various features described and defined in any aspects of the disclosure can be implemented and / or supplied and / or used independently. The disclosure also provides a computer program and a computer program product comprising software code adapted, when executed on a data processing apparatus, to perform any of the methods described herein, including any or all of their component steps. The disclosure also provides a computer program and a computer program product comprising software code which, when executed on a data processing apparatus, comprises any of the apparatus features described herein. The disclosure also provides a computer program and a computer program product having an operating system which supports a computer program for carrying out any of the methods described herein and / or for embodying any of the apparatus features described herein. The disclosure also provides a computer readable medium having stored thereon the computer program as aforesaid. The disclosure also provides a signal carrying the computer program as aforesaid, and a method of transmitting such a signal. The disclosure extends to methods and / or apparatus substantially as herein described with reference to the accompanying drawings. The disclosure will now be described, by way of example, with reference to the accompanying drawings. Description of the Drawings Figure 1 shows a system for generating a sequence of images. Figure 2 shows a computer device on which components of the system of Figure 1 may be implemented. Figure 3 shows a method of determining a three-dimensional representation of a scene. Figures 4a and 4b show method of determining a point based on a plurality of sub-points. Figure 5 shows a scene comprising a viewing zone. Figures 6a and 6b show arrangements of capture devices for determining points of the three-dimensional representation. Figure 7 shows a point that can be captured by a plurality of capture devices. Figures 8a and 8b show grids formed by the different capture devices. Figure 9 describes a method of determining a location of a point of the three-dimensional representation. Figure 10 shows a method of determining an angle of a point from a capture device used to capture the point. Figure 11 shows a method of determining a texture patch associated with a point of the three-dimensional representation. Figures 12a, 12b, 12c, and 12d illustrate the determination and use of a texture patch. Figures 13 and 14 show detailed methods of determining texture patches. Figures 15a, 15b, and 15c show a series of images. Figures 16a, 16b, and 16c show a method of rendering a two-dimensional image based on a predicted location of a point of a three-dimensional representation. Figure 17 shows a method for determining one or more movement vectors for a texture point. Figure 18 shows a texture patch. Figure 19 shows a bitstream. Figures 20a, 20b, and 20c show, respectively, fields of a point, a movement vector, and a texture point that may be signalled in a bitstream. Figures 21a - 21 h show exemplary structures of points, movement vectors, and texture points. Figure 22 shows an exemplary architecture for, and method for, generating a bitstream. Figures 23a and 23b show an exemplary architecture for, and method for, parsing a bitstream. Description of the Preferred Embodiments Referring to Figure 1, there is shown a system for generating a sequence of images. This system can be used to generate, and then display, a representation of an environment, which may comprise a VR environment (or an XR environment). The system comprises an image generator 11, an encoder 12, a transmitter 13, a network 14, a receiver 15, a decoder 16 and a display device 17. These components may each be implemented on separate apparatuses. Equally, various combinations of these components may be implemented on a shared apparatus; for example, the image generator 11, the encoder 12, and the transmitter 13 may all be part of a single image data generation device. Similarly, the receiver 15, the decoder 16, and the display device 17 may all be a part of a single image rendering device. Typically, the system comprises at least one encoding computer device (e.g. a server of a content provider) and at least one rendering computer device (e.g. a VR headset). Referring to Figure 2, each of the components, and in particular the image generator 11, the encoder 12, the transmitter 13, the receiver 15, the decoder 16 and the display device 17 is typically implemented on a computer device 20, where, as described above, a plurality of these components may be implemented on a shared computer device. Each computer device comprises one or more of: a processor 21 for executing instructions (e.g. so as to perform one or more of the steps of the various methods described below), a communication interface 22 for facilitating communication between computer devices (e.g. an ethernet interface, a Bluetooth® interface, or a universal serial bus (UBS) interface, a memory 23 and / or storage 24 for storing information and instructions (e.g. a random access memory (RAM), a read only memory (ROM), a hard drive disk (HDD) a solid state drive (SSD), and / or a flash memory, and a user interface 25 (e.g. a display, a mouse, and / or a keyboard) for enabling a user to interact with the computer device. These components may be coupled to one another by a bus 25 of the computer device. The computer device 20 may comprise further (or fewer) components. In particular, the computer device (e.g. the display device 17) may comprise one or more sensors, such as an accelerometer, a GPS sensor, or a light sensor. These sensors typically enable the computer device to identify an environmental condition and / or an action of wearer of the display device. Turning back to Figure 1, the image generator 11 is configured to generate a sequence of image data (e.g. a sequence of image frames) to enable the display device 17 to use this image data to display a plurality of images. The image data may comprise one or more digital objects and the image data may be generated or encoded in any format. For example, the image data may comprise point cloud data, where each point has a 3D position and one or more attributes. These attributes may, for example, include, a surface colour, a transparency value, an object size and a surface normal direction. Each attribute may have a value chosen from a continuous range or may have a value chosen from a discrete set. The image data enables the later rendering of images. This image data may enable a direct rendering (e.g. the image data may directly represent an image). Equally, the image data may require further processing in order to enable rendering. For example, the image data may comprise three-dimensional point cloud data, where rendering a two-dimensional image using this data requires processing based on a viewpoint of this two-dimensional image. The image data may comprise depth map data, where one or more pixels or objects in the image is associated with a depth that is specified by the depth map data. The depth map data may be provided as a depth map layer, separate from an image layer. In some contexts, such as MPEG Immersive Video (MIV), the image layer may instead be described as a texture layer. Similarly, in some contexts, the depth map layer may instead be described as a geometry layer. The image data may include a predicted display window location. The predicted display window location may indicate a portion of an image that is likely to be displayed by the display device 17. The predicted display window location may be based on a viewing position (such as a virtual position and / or orientation of the user in a 3D environment) of the user, where this viewing position may be obtained from the display device. The predicted display window location may be defined using one or more coordinates. For example, the predicted display window location may be defined using the coordinates of a corner or center of a predicted display window, and may be defined using a size of the predicted display window. The predicted display window location may be encoded as part of metadata included with the frame. The image data for each image (e.g. each frame) may include further information, which may be provided as a part of an image, e.g. as part of the point cloud data, or as separate layers. In particular, the image data may include audio information or haptic feedback information indicating audio or haptics which can accompany displayed visual data. An audio layer or haptic layer may accompany each image, and may be omitted for images where no accompanying audio or haptics are required. Similarly, the image data may comprise interactivity information, where the image data may contain or indicate elements with which a user can interact. The interactivity information may, for example, define a behaviour of an element, where a user is able to interact with the element based on this behaviour. The behaviour typically defines a change in an element that occurs as a result of a user interaction where this change may comprise a change in the attributes of the element or in the rendering of the element. As an example, where an image contains a target element, the target element may be arranged to disappear when a user interacts with this element, or to provide feedback indicating that the user has interacted with the target. This interactivity data may be provided as part of, or separately to, the image data. The image data may indicate, or may be combinable with, a state of the virtual environment, a position of a user, or a viewing direction of the user. Here, the position and viewing direction may be physical properties of the user in the real-world, or position and viewing direction may also be purely virtual, for example being controlled using a handheld controller. The image generator 11 may, for example, obtain information from the display device 17 that indicates the position, viewing direction, or motion of the user. Equally, the image generator may generate image data such that it can later be combined with this position, viewing direction, or motion, where the image generator may generate a full scene which is only partially viewed by a user depending on the position of that user. In some cases, the generated image may be independent of user position and viewing direction. This type of image generation typically requires significant computer resources such as a powerful GPU, and may be implemented in a cloud service, or on a local but powerful computer. For example, a cloud service (such as a Cloud Rendering Service (CRN)) may reduce the cost per-user and thereby make the image frame generation more accessible to a wider range of users. Here “rendering” refers at least to an initial stage of rendering to generate an image. Further rendering may occur at the display device 17 based on the generated image to produce a final image which is displayed. The image generator 11 may, for example, comprise a rendering engine for initially rendering a virtual environment such as a game or a virtual meeting room. The encoder 12 is configured to encode frames to be transmitted to the display device 17. The encoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. In some embodiments, the image generator 11 may transmit raw, unencoded, data through the network 14. However, such transmission typically leads to a high file size and requires a high bandwidth so that it is typically desirable to encode the data prior to the transmission. The encoder 12 may encode the image data in a lossless manner or may encode the data a lossy manner. The encoder may apply inter-frame or intra-frame compression based on a currently-encoded frame and optionally one or more previously encoded frames. The encoder may be a multi-layer encoder, such as a low complexity enhancement video codec (LCEVC) enabled encoder. Where the generated frames comprise depth map data, the encoder 12 may perform layered encoding on each instance of image data (e.g. each frame) to generate an encoded frame comprising a base depth map layer and an enhancement depth map layer. Encoding a depth map in this way may improve compression. In some applications, such as HDR video, depth maps are desirably highly detailed with a bit depth of up to twelve or fourteen bits, which is a significant increase in the data to be transmitted. As a result, providing ways to improve compression of the depth map can make more realistic depth map-based displays viable when performing rendering or transmission of rendered data in real-time. Furthermore, this type of layered encoding makes it easy to drop (and then pick back up) one or more of the layers, which provides flexibility and tools for bandwidth management. Layered encoding is also helpful as the final decoder / user device (such as a user display device) can choose whether to process these extra layers. For example, in a non-layered approach, the best the end device (i.e. the receiver, decoder or display device associated with a user that will view the images) can do is determine that it does not have enough resources for a given quality (be it resolution, frame rate, inclusion of depth map) and then signal to the controller / renderer / encoder that it does not have enough resources. The controller then will send future images at a lower quality. In that alternative scenario, the end device still unfortunately has to process the higher quality data until the lower quality data arrives, if it can process the received images at all. In some of the described embodiments, this situation is improved upon because when / if the end device determines for example that it does not have the processing capabilities to handle the highest level of quality, then it can drop and / or choose not to process certain layers. The end device may also signal to the controller that it needs a lower level of quality, but in the meantime the end device can only process the number of layers that it can handle. Therefore, the end device can react to conditions much more quickly. In some cases, depth map data may be embedded in image data. In this case, the base depth map layer may be a base image layer with embedded depth map data, and the enhancement depth map layer may be an enhancement image layer with embedded depth map data. Alternatively, when the generated images comprise a depth map layer separate from an image layer and multi-layer encoding is applied, the encoded depth map layers may be separate from the encoded image layers. This has the advantage that the encoded depth map layers can be dropped under some conditions while still retaining image layers that can be displayed (albeit with a lower level of realism). For example, the encoded depth map layers can be dropped by a transmitter or encoder when available communication resources are reduced, or can be dropped by an end device which lacks the processing resources to handle the highest level of quality. Similarly, if some images comprise an audio base layer, a haptic feedback base layer, an audio enhancement layer or a haptic feedback enhancement layer, these can be processed or dropped flexibly. Again similarly, if some images comprise an interactivity data base layer or an interactivity enhancement layer these can be processed or dropped flexibly. For example, certain interactions may only be possible where a threshold bandwidth is available, where complex interactions (e.g. those enabling a conversation with a digital object) may be disabled before less complex interactions (e.g. changing a pixel colour) are disabled. Additionally or alternatively, where the image data comprises point cloud data, the encoder may apply a point cloud data encoding technique such as described in European patent application EP21386059.6, which is incorporated herein by reference. Such a point cloud encoder may act as a base encoder for a layered encoding technique such as LCEVC or VC-6. Notably LCEVC and VC-6 techniques encode and decode a layered signal, but are agnostic about the content type of data encoded in the signal. For example, the signal can include textures, video frames, geometry or depth data, meshes, point clouds, rendering attributes or physics engine attributes. The transmitter 13 may be any known type of transmitter for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The transmitter 13 may be configured to make decisions about how to transmit the image data, and / or may provide feedback to the encoder 12 or the image generator 11. For example, the transmitter may determine available communication resources (e.g. bandwidth) for transmitting image data, and may drop one or more layers from an encoded frame, or indicate to the image generator and / or encoder that image data should be generated and encoded with fewer layers, when insufficient bandwidth is available for transmission of all generated data. As specific examples, the transmitter may be configured to drop a depth map layer, an LCEVC enhancement layer, or a VC-6 enhancement layer from a frame when insufficient communication resources are available. The network 14 provides a channel for communication between the transmitter 13 and the receiver 15, and may be any known type of network such as a WAN or LAN or a wireless Wi-Fi or Bluetooth network. The network may further be a composite of several networks of different types. Many users only have access to a network with a bandwidth of 30MBps which can lead to latency jitter when streaming. The required bandwidth and the observed latency can be reduced by means of tactics such as forward-looking rendering and last-millisecond reprojection, which are enabled by improved compression. The receiver 15 may be any known type of receiver for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The decoder 16 is configured to receive and decode an encoded frame. The decoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. The display device 17 may for example be a television screen or a VR headset. The timing of the display may be linked to a configured frame rate, such that the display device may wait before displaying the image. The display device may be configured to perform warping, that is, to obtain a final display window location, adjust a warpable image to obtain a final image corresponding to a final viewing direction of the user, and display the final image. In this regard, the image data is typically arranged to provide a warpable image for which a portion of the image that is displayed at the display device 17 is dependent on a position or orientation of a viewer. The warpable image may then be rendered before a most up to date viewing direction of the user is known. The warpable image may be transmitted to the display device, or the warpable image may be transmitted to a rendering node which is near to the display device, and the display device or rendering node may perform time warping to generate a displayed image portion based on the warpable image and the most up to date viewing direction of the user. As mentioned above, a single device may provide a plurality of the described components. For example, a first rendering node may comprise the image generator 11, encoder 12 and transmitter 13. Additional similar rendering nodes may be included in the system, and may work together to generate the sequence of frames. In one case, multiple rendering nodes may each provide separate image data to an image data assembling node; for example, each rendering node may provide a part of a sequence of frames to a frame assembling node. For example, the receiver 15, decoder 16 or display device 17 may be configured to assemble parts of image data from multiple sources to generate a sequence of images for display on the display device. Alternatively, the image data assembling node may be separate from the receiver 15, decoder 16 and display device 17. Additionally or alternatively, multiple rendering nodes may be chained. In other words, successive rendering nodes may add to a sequence of image data as it passes from rendering node to rendering node, and eventually a complete sequence of image data is then provided to the receiver 15. Furthermore, each rendering node may obtain components of a render from multiple upstream rendering nodes and / or distribute components of a render to multiple downstream rendering nodes. A chain of rendering nodes may be useful for performing different rendering tasks that require different quantities of processing resources, or different frame rates. For example, a company may provide distributed processing in the form of a centralised hub which has abundant processing resources but is distant from users, and peripheral locations which have more scarce processing resources but are closer to users. Expensive but fairly static rendering features such as background lighting or environmental impact on sound may be generated at the central hub (for example using ray tracing), while features that require fewer resources but faster responses or higher frame rates may be generated closer to the user. In other words, the more responsive a rendering feature needs to be, the lower latency it needs between the rendering node which generates the feature and the user display and, in a chain of rendering nodes, the node which generates each rendering feature can be chosen based on a required maximum latency of that feature. On the other hand, if it is expensive to generate a rendering feature, then it may be preferable to generate the feature less frequency and with a higher maximum latency. For example, a static, high-quality background feature may be generated early in the chain of rendering nodes and a dynamic, but potentially lower-quality, foreground feature may be generated later in the chain of rendering nodes, closer to the user device. Here, environmental impact on sound means, for example, a set of surfaces may be constructed where each surface has different sound reflection and absorption properties depending upon material and shape. The frame rates may be matched by creating multiple frames with features generated at the lower frame rate, and combining them with the frames with features generated at the higher frame rate. In a non limiting embodiment, a preliminary rendering generates volumetric object data including movement vectors at a first (lowest) frame rate, then produces 2D rendered frames plus depth information for a specific user at a second (higher) frame rate, then transmits video plus depth data to the user device, which produces final frames for display via space warping (depth-based reprojections) at a third (highest) frame rate. One or more of these steps may be performed in combination with the other described embodiments. The viewing position of the user may change as additional rendering tasks are performed at different rendering nodes in the chain. Each or any rendering node may obtain an updated viewing position before performing its respective rendering task. Additionally, the system may simultaneously generate multiple sequences of image data for different respective users or different respective display devices. For example, in the context of a VR or AR experience, each user or display device may view a different 3D environment, or may view different parts of a same 3D environment. When using a chain of rendering nodes, each node may serve multiple users or just one user. For example, a starting rendering node (e.g. at a centralised hub) may serve a large group of users. For example, the group of users may be viewing nearby parts of a same 3D environment. In this case, the starting node may render a wide zone of view (“field of view”) which is relevant for all users in the large group. The starting node may send this wide field of view to a first middle rendering node which renders additional aspects of the 3D environment. These additional aspects may for example be aspects which require less processing power to render, or may be aspects which are specific to individual users of the group. Additionally, the middle rendering node may render features in a smaller field ofviewthan the starting node - this smaller field of view may be relevant to each user rather than the group of users. The first middle rendering node may additionally only serve a smaller number of users (e.g. half of the large group of users), with the remaining users being served by a second middle rendering node which also receives the wide field of view from the starting node. The middle rendering node(s) may then send sequences of second partially or fully rendered frames to an end device for each user. The end device may perform further processes such as warping or focal distance adjustments, optionally using depth map data. Preferably, each rendering node encodes the partially or fully rendered frames before transmitting them on to a next rendering node or to the receiver 15. This means that the required communication resources can be reduced when the rendering nodes are separated by one or more networks, or more generally are implemented in a distributed system such as a cloud. However, each rendering node in a chain is encoding a different partially or fully rendered frame, with different data. Therefore, it may be advantageous for different rendering nodes to use different rendering formats and / or encoding formats. For example, the output from a first rendering node may be point cloud data which logically describes a 3D scene. This point cloud data can be encoded using the techniques of EP21386059.6. A second rendering node may then operate on the point cloud data to generate image data that is more readily displayed by a generic display device, without requiring the display device to model the 3D environment. This image data may be encoded using video coding techniques. The chaining of rendering nodes may be extended to arbitrary tree structures, where a rendering node obtains partially rendered frames from more than one preceding rendering node, and generates further partially or fully rendered frames based on the multiple obtained sequences of partially rendered frames. For example, a content rendering network (CRN) comprising numerous rendering nodes may be used to serve a volumetric event to a large number of same-time users, such as users participating in a shared virtual environment. Rendering the same event for each user is far more expensive in terms of computation time and power consumption than rendering the volumetric effect once and performing the rendering equivalent of multicasting the volumetric effect for multiple users. For example, each user may have a second rendering node (such as a VR headset), and the network may comprise a central first rendering node. The first rendering node may render the volumetric event, and distribute partially rendered frames depicting the volumetric event to the different second rendering nodes. The second rendering node for each user may then integrate the partially rendered frames depicting the volumetric event into a view of the virtual environment which is currently being shown to each user, based on parameters such as the user’s virtual position. The receiver 15, decoder 16 and display device 17 may be consolidated into a single device, or may be separated into two or more devices. For example, some VR headset systems comprise a base unit and a headset unit which communicate with each other. The receiver 15 and decoder 16 may be incorporated into such a base unit. In some embodiments, the network 14 may be omitted. For example, a home display system may comprise a base unit configured as an image source, and a portable display unit comprising the display device 17. In the event that the decoder 16 or the display device 17 does not or cannot handle one or more layers, the receiver 15 or another transmitter associated with the decoder or display device may send a corresponding layer drop indication back through the network 14. The layer drop indication may be received by each rendering node. A rendering node which generates partially or fully rendered frames for that specific decoder or display device may cease generating the dropped layer. On the other hand, a rendering node which generates partially or fully rendered frames for multiple end devices may disregard a layer drop indication received from one end device (as the dropped layer is still needed for other devices). Alternatively, rendering nodes which serve multiple end devices may record received layer drop indications, and may cease generating the dropped layer only when all end devices served by the rendering node indicate that the layer is to be dropped. In preferred examples, the encoders or decoders are part of a tier-based hierarchical coding scheme or format. Hierarchical coding enables frames to be communicated with higher resolution and / or higher frame rate than is possible in single-tier coding schemes. In hierarchical coding, one or more enhancement layers is communicated with base data, where the enhancement layers can be used to up-sample the base data at the decoder, for example providing up-sampling in a spatial or temporal dimension. When combined with equivalent down-sampling of the original frames and generation of the enhancement layer at an encoder, hierarchical coding can overall provide lossless compression of data, with higher resolution and / or higher frame rate for a given transmission bit rate. Examples of a tier-based hierarchical coding scheme include LCEVC: MPEG-5 Part 2 LCEVC (“Low Complexity Enhancement Video Coding”) and VC-6: SMPTE VC-6 ST-2117, the former being described in PCT / GB2020 / 050695, published as WO 2020 / 188273, (and the associated standard document) and the latter being described in PCT / GB2018 / 053552, published as WO 2019 / 111010, (and the associated standard document), all of which are incorporated by reference herein. However, the concepts illustrated herein need not be limited to these specific hierarchical coding schemes. A further example is described in WO2018 / 046940, which is incorporated by reference herein. In this example, a set of residuals are encoded relative to the residuals stored in a temporal buffer. LCEVC (Low-Complexity Enhancement Video Coding) is a standardised coding method set out in standard specification documents including the Text of ISO / IEC 23094-2 Ed 1 Low Complexity Enhancement Video Coding published in November 2021, which is incorporated by reference herein. The system describes above is suitable for generating and presenting a representation of a scene, where this scene displays media content to a user. The scene typically comprises an environment, where the user is able to move (e.g. to move their head or to turn their head) to look around the environment and / or to move around the environment. For example, the scene may be a scene of a room in a building, where the user is able to move around the room (e.g. by moving in the real-world and / or by providing an input to a user interface) in order to inspect various parts of the room. Typically, the scene is a XR (e.g. a VR) scene, where the user is able to move about the scene in three degrees of freedom (3DoF) or six degrees of freedom (6DoF) so as to experience the scene. As has been described with reference to Figure 1, the image generator 11 may be arranged to determine point cloud data, where each point of the point cloud has a 3D position and one or more attributes. More generally, the image generator (or another component) is arranged to determine a three-dimensional representation of a scene, where this three-dimensional representation is thereafter used to generate two-dimensional images that are presented to a user at the display device 17. While the points are typically points of a point cloud, more generally the disclosure extends to any point that is associated with a location and a value. Therefore, the points may, more generally, be considered to be data (or datapoints), which data is associated with a location and a value, and the ‘points’ may comprise polygons, planes (regular or irregular), Gaussian splats, etc. Referring to Figure 3, there is described a method of determining (an attribute for) a point of such a three-dimensional representation. The method comprises determining the attribute using a capture device, such as a camera or a scanner. The scene may comprise a real scene, in which attribute values are captured using a camera, or a virtual scene (e.g. a three-dimensional model of a scene), in which attribute values are captured using a virtual scanner. Where this disclosure describes ‘determining a point’ it will be understood that this generally refers to determining a point that has a location and an attribute value, where determining the point comprises determining the attribute value and / or storing a point that comprises at least an attribute value and a location value (these values may be indirect values, e.g. where the location is identified relative to another point). Once a plurality of points have been captured, these points can be stored as a three-dimensional representation (e.g. a point cloud) so as to enable the reconstruction of the three-dimensional scene based on this representation. Typically, the scene comprises a simulated scene that exists only on a computer. Such a scene may, for example, be generated using software such as the Maya software produced by Autodesk®. The attributes determined using the methods described herein may then depend on virtual objects located within the scene as well as a virtual lighting arrangement used in the scene. In a first step 31, a computer device initiates a capture process for a capture device, the capture process being initiated with an initial azimuth angle (e.g. of 0°) and an initial elevation angle (e.g. of 0°). In a second step 32, the computer device causes a point to be captured using the capture device at the current azimuth angle and current elevation angle. Capturing a point typically comprises assigning an attribute value to the point, which attribute value may, for example, be a color of the point and / or a transparency value of the point. Typically, the point has one or more color values associated with each of a left eye and a right eye of a viewer. Capturing the point may also comprise determining a normal value associated with the point, e.g. a normal of a surface on which the point lies. Typically, capturing the point further comprises determining a location of the point, e.g. by determining a distance of the point from the camera. In practice, determining the point may comprise sending a ‘ray’ from the capture device and then stepping through a computer model to determine which surface of the computer model is impacted by the ray. The color, transparency, and normal of this surface are then recorded alongside the distance of the surface from the capture device. In a third step, 33, the computer device determines whether a point has been captured for the capture device at each azimuth of a range of azimuths and in a fourth step 34, if points have not been captured at each azimuth, then the azimuth angle is incremented and the method returns to the second step 32 and another point is captured. The azimuth angle may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of azimuth angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. Once a point has been captured for each azimuth, in a fifth step 35, the computer device determines whether a point has been captured for the capture device at each elevation of a range of elevations and in a sixth step 36, if points have not been captured at each elevation, then the azimuth angle is reset to the initial value, elevation angle is incremented and the method returns to the second step 32 and another point is captured. The elevation angles may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of elevation angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. In a seventh step 37, once points have been captured for each azimuth angle and each elevation angle, the scanning process ends. This method enables a capture device to capture points at a range of elevation and azimuth angles. This point data is typically stored in a matrix. The point data may then be used to provide a representation of the scene to a user, e.g. the three-dimensional representation formed by the point data may be processed to produce two-dimensional images for each eye of a user, with these images then being shown to a user via the display device 17 to provide a virtual reality experience to the viewer. By using the captured data, a video can be provided to a viewer that enables the viewer to move their head to look around the scene (while remaining at the location of the capture device). It will be appreciated that the capture pattern (or scanning pattern) described with reference to Figure 3 is purely exemplary and that numerous capture patterns are possible. In general, the capture process for each capture device comprises capturing one or more points at one or more azimuth angles and / or one or more elevation angles. The ‘points’ captured by the capture device are typically associated with a size, such as a height, a width, or a depth. That is, the points typically relate to two-dimensional planes / pixels and / or three-dimensional voxels. In this regard, there is necessarily some space between the locations of adjacent points (since if the points had no width, then an infinite number of points would be required to capture points at each angle). The size provides points that depict a non-negligible area of the three-dimensional space so that a plurality of points can be fit together to provide a depiction of the scene to a viewer. The width and height of each point is typically dependent on the distance of that point from the capture device, where more distant points have a larger width / height. The width and height of each point is typically determined so that when each point is displayed, there is no space between adjacent points (indeed, there may be some overlap between points to ensure that no gaps appear between points). This height / width of each point can be determined at the time of capturing the points, or can be determined or defined after the capture of the points. Typically, the points comprise a size value, which is stored as a part of the point data. For example, the points may be stored with a width value and / or a height value. Typically, the minimum width and the minimum height of a point are set by the angle increment of the azimuth angle and the elevation angle respectively. The size may be then specified in terms of this angle increment and / or in terms of this minimum width / minimum height (e.g. as being a multiple of the angle increment). In some embodiments, the size value is stored as an index, which index relates to a known list of sizes (e.g. if the size may be any of 1x1, 2x1, 1x2, 2x2, pixels this may be specified by using 3 bits and a list that relates each combination of bits to a size). The size may be stored based on an underscan value. In this regard, where an object is very near to the viewing zone it may be captured using an unnecessarily dense arrangement of points. Therefore, certain surfaces or areas of the representation may be associated with an underscan value, which underscan value defines a reduction in the number of points captured as compared to a representation without underscan. The size of the points may be defined so as to indicate this underscan value. In an exemplary embodiment, the underscan value is an integer value between 0 and 3 and the size is stored as a combination of point dimensions (e.g. a width in the range [0,2]) and a height in the range ([0,2]) and an underscan factor (e.g. an underscan factor in the range [0,3]). In some embodiments, the width and the height are dependent on the underscan factor. For example, when the underscan factor exceeds a threshold value, the possible height and width values may be limited. In a specific example, when the underscan factor is 3, the width and the height may be limited to the range [0,1 ]. The size may then be defined as size = underscan*9 + height*3 + width. Such a method provides efficient storage and indication of width, height, and underscan values. As shown in Figure 4a, typically, for each capture step (e.g. each azimuth angle and / or each elevation angle), a plurality of sub-points SP1, SP2, SP3, SP4, SP5 is determined. For example, where the azimuth angle increment is 0.1° then for an azimuth angle of 0°, sub-points may be determined at azimuth angles of-0.05°, -0.025°, 0, 0.025°, and 0.05° (and similar sub-points may be determined for a plurality of elevation angles). Attribute values of these sub-points may then be combined to obtain an attribute value for the point. For example, a maximum attribute value of the sub-points may be used as the value for the point, an average attribute value of the sub-points may be used as the value for the point, and / or a weighted average ofthe sub-points may be used as the value forthe point. It will be appreciated that numerous other methods for combining the attribute values ofthe sub-points are possible. By determining the attribute of a point based on the attributes of sub-points, the accuracy ofthe capture process can be increased. While it would be possible to simply reduce the increment ofthe angle steps to provide a higher resolution scene, by considering sub-points but only storing attributes for points, a balance can be struck between accuracy and file size (since storing every sub-point would lead to a substantial increase in the amount of data that needs storing). With the example of Figure 4a, for each point ofthe three-dimensional representation that is captured by a capture device, this capture device may obtain attributes associated with each ofthe sub-points SP1, SP2, SP3, SP4, SP5, combine these attributes to obtain a point attribute, and then store a point with a distance that is an average (e.g. a weighted average) ofthe distances ofthe sub-points from the capture device, at the nominal angle ofthe point, with the point attribute. As shown in Figure 4b, where a plurality of sub-points SP1, SP2, SP3, SP4, SP5 are considered, these points may have different distances from the location ofthe capture device. In some embodiments, the attributes ofthe sub-points may be combined in dependence on this distance, e.g. so that sub-points nearer to the capture device have higher weightings. However, the possibility of sub-points with substantially different distances raises a potential problem. Typically, in order to determine a distance for a point, the distances forthe sub-points are averaged. But where the sub-points have substantially different distances and / or are related to different surfaces in the scene, this may result in the point having a distance that does not correspond to any actual surface in the scene. Therefore, the point may seem to hang in space (e.g. to hang between the front and rear surfaces shown in Figure 4b. Similarly, where the attribute values ofthe sub-points greatly differ, e.g. if the sub-points SP1 and SP2 are white in colour and the sub-points SP3 and SP4 are black in colour, then the attribute value ofthe point may be substantially different to the attribute value of other points in the scene. In an example, if the scene were composed of black and white objects, the point may appear as a grey point hanging in space between these objects. In some embodiments, the computer device is arranged to aggregate sub-points so as not to create any floating points. For example, the computer device may determine whether the sub-points are spatially coherent by employing a clustering algorithm (e.g. a k-means clustering algorithm). Where the sub-points are spatially coherent (e.g. where a difference in the distance of the sub-points is below a threshold value), these distances may be averaged to obtain a distance for the point. Where the sub-points are not spatially coherent, the sub-points may be processed to ensure that the distance of any point places it upon a surface; for example, in the system of Figure 4b, sub-points SP1, SP2, and SP3 may be grouped into a first point and sub-points SP4 and SP5 may be grouped into a second point. Since each sub-point is associated with the same capture device and capture angle (all of these sub-points being associated with a capture step that has a particular azimuth angle and elevation angle), these points may be located at the same angle with respect to a capture device. Therefore, to ensure that each sub-point affects the representation considered, the first point (made up of sub-points SP1, SP2, and SP3) may have a smaller distance value than the second point (made up of sub-points SP4 and SP5) and the first point may be assigned a nonzero transparency value so that the second point can be seen through the first point. By capturing points at a plurality of azimuth angles and elevation angles, e.g. using the method described with reference to Figure 3, it is possible to provide a three-dimensional representation of the scene that can later be used to enable a viewer to view the scene from a plurality of angles. More specifically, given the three-dimensional points captured by the capture device, a computer device is able to render a two-dimensional representation (e.g. a two-dimensional image) of the scene for each eye of a viewer so as to provide a representation with an impression of depth. The computer device may render a series of two-dimensional representations to enable the viewer to look around the scene, where the two-dimensional representations are rendered based on an orientation of the viewer’s head. In this way, the determined representation is useable to provide, for example, a virtual reality (VR), mixed reality (MR), augmented reality (AR), and / or extended reality (XR) experience to the viewer. To enable such a display, the display device 17 is typically a virtual reality headset, that comprises a plurality of sensors to track a head movement of the user. By tracking this head movement, the display device is able to update the images being displayed to the viewer as the viewer moves their head to look about the scene. Typically, this involves the display device sensing the sensor data to an external computer device (e.g. a computer connected to the display device via a wire). The external computer device may comprise powerful graphical processing units (GPUs) and / or computer processing units (CPUs) so that the external computer device is able to rapidly render appropriate two-dimensional images for the viewer based on the three-dimensional images and the sensor data. In some embodiments, the external computer device may comprise a server device, where the display device 17 may be connected to this server device wirelessly. This enables the two-dimensional images to be streamed from the serverto the display device so as to enable the display of high-quality images without the need for a viewer to purchase expensive computer equipment. In other words, operations that require large amounts of computing power, such as the rendering of two-dimensional images based on the three-dimensional representation, may be performed by the server, so that the display device is only required to perform relatively simple operations. This enables the experience to be provided to a wide range of viewers. In some embodiments, a first two-dimensional image is provided to the display device 17 (and / or a connected device) and this first image is ‘warped’ in order to provide an image for viewing at the display device. The warping of the image comprises processing the image based on the sensor data in order to provide an image that matches a current viewpoint of the viewer. By performing the warping at the display device or another local device, the lag between a head movement of the user and an updating of the two-dimensional representation of the scene can be reduced. One issue with the above-described method of capturing a three-dimensional representation is that it only enables a viewer to make rotational movements. That is, since the points are captured using a single capture device at a single capture location, there is no possibility of enabling translational movements of a viewer through a scene. This inability to move translationally can induce motion sickness within a viewer, can reduce a degree of immersion of the viewer, and can reduce the viewer’s enjoyment of the scene. Therefore, it is desirable to enable translational movements through the scene. To enable such movements, the three-dimensional representation of the scene may be captured using a plurality of capture devices placed at different locations (or the same capture device placed at different locations). A viewer is then able to move around the scene translationally (e.g. by moving between these locations). More generally, by capturing points for every possible surface that might be viewed by a viewer, a three-dimensional representation of a scene may be captured that allows a suitable two-dimensional representation ofthis scene to be rendered regardless of a location of a viewer(e.g. regardless of where a user is standing within a virtual room). This need to capture points for every possible surface (so as to enable movement about a scene) greatly increases the amount of data that needs to be stored to form the three-dimensional representation. Therefore, as has been described in the application WO 2016 / 061640 A1, which is hereby incorporated by reference, the three-dimensional representation may be associated with a viewing zone, or a zone of viewpoints (ZVP), where the three-dimensional representation is arranged to enable a user to move about the viewing zone so as to view the scene. Figure 5 illustrates such a viewing zone 1 and illustrates how the use of a viewing zone limits the amount of image data that needs to be stored to provide a three-dimensional representation of the scene. With the scene shown in this figure, and the viewing zone 1 shown in this figure, it is not necessary to determine attribute data for the occluded surface 2 since this occluded surface cannot be viewed from any point in the viewing zone. Therefore, by enabling the user to only move within the viewing zone (as opposed to around the whole scene) the amount of data needed to depict the scene is greatly reduced. While Figure 5 shows a two-dimensional viewing zone, it will be appreciated that in practice the viewing zone 1 is typically a three-dimensional zone or volume. The viewing zone 1 may, for example, comprise a rectangular volume, or a rectangular parallelepiped, and the viewing zone may have a height of at least 30 cm, a depth of at least 30 cm, and / or a width of at least 30 cm, where these dimensions enable a userto move their head while remaining in the viewing zone. This is merely an exemplary arrangement of the viewing zone; it will be appreciated that viewing zones of various shapes and sizes may be used (e.g. spherical viewing zones). That being said, it is preferable that the viewing zone is limited so as to cover only a part of the volume of the scene, e.g. no more than 50% of the scene no more than 25% of the scene, and / or no more than 10% of the scene. In this regard, if the viewing zone is the same size as the scene, then the three-dimensional representation will simply be a standard representation for virtual reality (that enables a user to move freely about the scene) - and so the use of the viewing zone will not provide any reduction in file size. The viewing zone 1 enables movement of a viewer around (a portion of) the scene. For example, where the scene is a room, the base representation may enable a user to walk around the room so as to view the room from different angles. In particular, the viewing zone enables a user to move through the scene with six degrees-of-freedom (6DoF) movement through the scene, where this aids in the provision of an immersive experience. In some embodiments, the viewing zone 1 may be four-dimensional, where a three-dimensional location of the viewing zone changes over time - and in such embodiments the size and location of the occluded surface 2 may also change over time. More generally, it will be appreciated that viewing zones may be formed in any size or shape, with different sizes and shapes being suitable for different scenes. The volume of the viewing zone 1 is typically selected so that a user is able to move to a degree sufficient to avoid motion sickness and to provide an immersive sensation, while still only enabling a limited amount of movement (where this leads to a smaller file size as compared to an implementation where a user is able to fully move about the scene). Typically, the viewing zone is arranged to enable a user to move their head while they are sitting or standing, but not to freely roam around a room. The viewing zone 1 may have a (e.g. real-world) volume of less than five cubic metres (5m3), less than one cubic metre (1m3), less than one-tenth of a cubic metre (0.1m3) and / or less than one-hundredth of a cubic metre (0.01m3). The viewing zone 1 may also have a minimum size, e.g. the viewing zone may have a volume of at least 1% of the volume of the scene, at least 5% of the volume of the scene, and / or at least than 10% of the volume of the scene. Similarly, the viewing zone may have a volume of at least one-thousandth of a cubic metre (0.01m3); at least one-hundredth of a cubic metre (0.01 m3); and / or at least one cubic metre (1m3). The ‘size’ of the viewing zone 1 typically relates to a size in the real world, where if the viewing zone has a length of one metre this means that a user is able to move one metre in the real world while staying within the viewing zone. The size of the viewing zone in the scene may be greater than, equal to, or less than the size of the viewing zone in the real world. For example, the viewing zone may scale a real-world distance so that moving one metre in the real world moves the user less than (or more than) one metre in the scene. This enables the scene to provide different perceptions to the user (e.g. to make the user feel larger or smaller than they are in real life). Similarly, the viewing zone may scale a real-world angle so that rotating one degree in the real world rotates the user less than (or more than) one degree in the scene. Therefore, a viewing zone with a volume of one cubic metre typically connotes a viewing zone in which the user is able to move about a one cubic metre volume in the real world while remaining in the viewing zone. And this may cause the user to move about a volume that is more than, or less than, one metre in the scene. Referring to Figure 6a, in order to capture points for each surface and location that is visible from the viewing zone 1, a plurality of capture devices C1, C2, ..., C9 may be used (e.g. a plurality of virtual scanners and / or a plurality of cameras). Each capture device is typically arranged to perform a capture process, e.g. as described with reference to Figure 3, in which the capture device captures points at a plurality of azimuth angles and elevation angles. By locating the capture devices appropriately, e.g. by locating a capture device at each corner of the viewing zone, it can be ensured that most (or all) points of a scene are captured. Typically, a first capture device C1 is located at a centrepoint of the viewing zone 1. In various embodiments, one or more capture devices C2, C3, C4, C5 may be located at the centre of faces of the viewing zone; and / or one or more capture devices C6, C7, C8, C9 may be located at edges of and / or corners of the viewing zone. Figure 6a shows a two-dimensional view (e.g. a plan view) of a rectangular viewing zone. It will be appreciated that within this viewing zone each capture device may be located on a shared plane. Equally, the various capture devices may be located on different planes. Referring, for example, to Figure 6b, there is shown a three-dimensional view of a cuboid viewing zone, where there is a capture device located: at the centre of the viewing zone; at the centre of each face of the viewing zone; and at each corner of the viewing zone. With this arrangement, many locations in the scene (e.g. specific surfaces) will be captured by a plurality of capture devices so that there will be overlapping points relating to different capture devices. This is shown in Figure 7, which shows a first point P1 being captured by each of a first capture device C1, a sixth capture device C6, and a seventh capture device C7. Each capture device captures this point at a different angle and distance and may be considered to capture a different ‘version’ of the point. Typically, only a single version of the point is stored, where this version may be the highest quality version of the point and / or may be the version of the point associated with the nearest and / or least angled capture device. In this regard, the highest ‘quality’ version of the point is captured by the capture device with the smallest distance and smallest angle to the point (e.g. the smallest solid angle). In this regard, as described with reference to Figures 4a and 4b, capturing a point for a given azimuth angle and elevation angle typically comprises capturing a plurality of sub-points at varying sub-point azimuth and elevation angles spread around the point azimuth and elevation angles. Due to the different spreads of sub-points, each capture device will capture a different version of the point (that has a different attribute) even when the points are at the same location. Capture devices that are close to the point and less angled with respect to the point typically have a smaller spread of sub-points and so typically obtain a version of a point that is sharper than a version of that point captured by more distant capture devices. In some embodiments, a quality value of a version of the point is determined based on the spread of subpoints associated with this version (e.g. based on the perimeter formed by these sub-points and / or based on a surface area or volume bounded by these sub-points). The version of the point that is stored may depend on the respective quality values of possible versions of the points. Regarding the ‘versions’ of the points, it will be appreciated that two ‘points’ in approximately the same location captured by each capture device may not have exactly the same location in the three-dimensional representation. More specifically, since each capture device typically projects a ‘ray’ at a given angle, the rays of differing capture devices may contact the surface at different locations for each capture device. Two points may be considered to be two ‘versions’ of a single point when they are within a certain proximity, e.g. a threshold proximity. For example, where the first capture device C1 captures a first point and a second point at subsequent azimuth angles, and the sixth capture device C6 captures a further point that is in between the locations of the first point and the second point, this further point may be considered to be a ‘version’ of one of the first point and the second point. This difference in the points captured by different capture devices is illustrated by Figures 8a and 8b, which show the separate captured grids that are formed by two different capture devices. As shown by these figures, each capture device will capture a slightly different ‘version’ of a point at a given location and these captured points will have different sizes. Each capture step is associated with a particular range of angles (e.g. a nominal capture angle of 1° might encompass angles from 0.9° to 1.1°), and therefore capture devices that are far from a point to be captured represent a wider region at the capture distance than capture devices closer to that point to be captured. As shown in Figure 8a, the capture device C1 would capture the points P1 and P2 in separate brackets, whereas for the capture device C2 these points are in the same bracket. Therefore, the capture device C2 might determine a single point that encompasses both points P1 and P2, whereas the capture device C1 would determine separate points for these two points. Considering then a situation in which points P1 and P2 are captured separately, and capture device C1 is used to capture point P1 while capture device C2 being used to capture point P2, it should be apparent that the ‘sizes’ of these captured points, and the locations in space that are encompassed by the captured points will be based on different grids. For example, the width of the captured point P2 captured by the capture device C2 will be larger than the width of the captured point P1 captured by the capture device C1. The capture process may be determined based on the existence of these different grids, and on the different bracket widths that occur at different distances from a capture device. Figure 8a shows an exaggerated difference between grids for the sake of illustration. Figure 8b shows a more realistic embodiment in which the three-dimensional representation comprises a plurality of points associated with different capture devices, where these points lie on different grids associated with these different capture devices. In order to store the points of the three-dimensional representation, the points may be stored as a string of bits, where a first portion of the string indicates a location of the point (e.g. using x, y, z coordinates) and a second portion of the string locates an attribute of the point. In various embodiments, further portions of the string may be used to indicate, for example, a transparency of the point, a size of the point, and / or a shape of the point. A computer device that processes the three-dimensional representation after the generation of this representation is then able to determine the location and attribute of each point so as to recreate the scene. This location and attribute may then be used to render a two-dimensional representation of the scene that can be displayed to a viewer wearing the display device 17. Specifically, the locations and attributes of the points of the three-dimensional representation can be used to render a two-dimensional image for each of the left eye of the viewer and the right eye of the viewer so as to provide an immersive extended reality (XR) experience to the viewer. The present disclosure considers an efficient method of storing the locations of the points (e.g. at an encoder) and of determining the locations of the points (e.g. at a decoder). As has been described with reference to Figures 5a and 5b, the points of the three-dimensional representation are determined using a set of capture devices placed at locations about the viewing zone, where these capture devices are arranged to capture points at a series of azimuth angles and elevation angles. Typically, each of the capture devices is arranged to use the same capture process (e.g. the same series of azimuth angles and elevation angles), though it will be appreciated that different series of capture angles are possible. For example, there may be a plurality of possible series of capture angles, where different capture devices use different capture angles. In general, the present disclosure considers a method in which points are stored based on a capture device identifier and an indication of a distance of the point from the capture device associated with this capture device identifier. Typically, the point is also associated with an angular indicator, which indicates an azimuth angle and / or an elevation angle of the point relative to the identified capture device. It will be appreciated that the storage of the distance and the angle may take many forms. For example, the distance and the angle of each point may be converted into a universal coordinate system, where each capture device has a different location in this universal coordinate system. In particular, each point may be stored with reference to a centre of this universal coordinate system, which centre may be co-located with a central capture device. Where a point is determined based on a distance and an angle from a capture device of a known location in this universal coordinate system, the coordinates of the point in this universal coordinate system can be determined trivially - and the location of the point may then be stored either relative to the capture device or as a coordinate in the universal coordinate system. The capture device identifier may comprise a location of a capture device (e.g. a location in a co-ordinate system of the three-dimensional representation). Equally, the capture device identifier may comprise an index of a capture device. Similarly, the indication of the azimuth angle and the elevation angle for a point may comprise an angle with reference to a zero-angle of a co-ordinate system of the three-dimensional representation. Equally, the azimuth angle and / or the elevation angle may be indicated using an angle index. In some embodiments, the three-dimensional representation is associated with configuration information, which configuration information comprises one or more of: a set of capture device indexes; locations associated with the capture devices and / or the capture device indexes; a spacing of capture devices (e.g. so that locations of the capture devices can be determined from a location of a first capture device and the spacing); angles associated with a capture process for the capture devices; an azimuth angle increment and / or an elevation angle increment associated with the capture process; and a set of angle indexes (e.g. to match an angle index to an angle). With this configuration information, it is possible to determine a location of each capture device from an index of that capture device and / orto determine a capture angle from a known capture process. Therefore, given two numbers: a capture device index and an angle index (that is associated with a combination of a specific azimuth angle and a specific elevation angle), a location of a capture device and a direction of a point from this capture device can be determined. By also signalling a distance of the point from the signalled capture device, a precise location of the point in the three-dimensional space can be signalled efficiently. Typically, the point is associated with each of: a camera index, a distance, an first angular index (e.g. a first azimuth), and a second angle (e.g. a second elevation) This method of indicating a location of a point enables point locations to be identified using a much smaller number of bits than if each point location is identified using x, y, z coordinates. Referring to Figure 9, there is shown a method of determining a location of a point. This method is carried out by a computer device, e.g. the image generator 11 and / or the decoder 15. In a first step 41, the computer device identifies an indicator of a capture device used to capture the point. Typically, this comprises identifying a portion of a string of bits associated with a capture device index. In a second step 42, the computer device identifies an indicator of an angle of the point from the capture device. Typically, this comprises identifying an angle index, e.g. an azimuth index and / or an elevation index and / or a combined azimuth / elevation index, which index(es) identifies a step of the capture process during which the point was captured. In a third step 43, based on the identifiers, the computer device determines the location of the capture device and the angle of the point from the capture device. The capture device identifier is typically a capture device index, which is related to a capture device location based on configuration information that has been sent before, or along with, the point data. For example, the configuration information may specify: Location of first capture device is (0,0,0). Step between capture devices is (0,0,1) along the grid, then across the grid, then up the grid. - The grid is (10,10,10). With this information, a capture device with an index of 1 can be determined to be located at (0,0,0); a capture device with an index of 5 can be determined to be located at (0,0,4); a capture device with an index of 12 can be determined to be located at (0,1,0), and so on. Equally, the configuration information may specify a list of camera indexes and locations associated with these indexes, where this enables the use of a wide range of setups of capture devices. Typically, the three-dimensional representation is associated with a frame of video. The configuration information may be constant over the frames of the video so that the configuration information needs to be signalled only once for an entire video. Therefore, the configuration information may be transmitted alongside a three-dimensional representation of a first frame of the video, with this same information being used for any subsequent frames (e.g. until updated configuration information is sent). The angle identifier may similarly be related to an angle by a location and an increment that are signalled in a configuration file. For example, the configuration information may specify: An azimuth increment and an elevation increment are each 1°. There are 359 increments for each angle type. With this information: a capture angle with an index of 1 can be determined to be at an azimuth angle of 0° and an elevation angle of 0°; a capture angle with an index of 10 can be determined to be at an azimuth angle of 10° and an elevation angle of 0°; a capture angle with an index of 360 can be determined to be at an azimuth angle of 0° and an elevation angle of 1°; and a capture angle with an index of 370 can be determined to be at an azimuth angle of 9° and an elevation angle of 1°; etc. In a fourth step 44, based on the determined location of the capture device and the determined angle, a location of the point is determined. Typically, this comprises determining the location of the point based on the location of the capture device, the capture angle, and a distance of the point from the capture device (where this distance is specified in the point data for the point). Determining the location of the point typically comprises determining the location of the point relative to a centrepoint of the three-dimensional representation, this location of the point may then be converted into a desired coordinate system and / orthe point may be processed based on its location (e.g. to stitch together adjacent points). The angular identifier typically comprises a first angular identifier and a second angular identifier, where the first identifier provides the azimuthal angle of the point and the second identifier provides the elevation angle of the point. Referring to Figure 10, each angular identifier may be provided as an index of a segment of the three-dimensional representation, where, for example, an index of 0 may identify the point as being in a first angular bracket 101 and an index of 1 may identify the point as being in a second angular bracket 102. In this regard, the capture devices are arranged to perform a capture process, e.g. as described with reference to Figure 3, with a non-infinite angular resolution. Given this non-infinite resolution, each point is not a one-dimensional point located at a precise angle. Instead, each point is a point for a particular area of space, with the size ofthis area being dependent on the angular resolution as well as the distance of the point from the capture device. In other words, each capture angle determines a point for an angular range (with the range being dependent on the angular resolution). That is, if the capture process leads to points being captured at angles of 10°, 11°, and 12° then this can equally be considered to relate to points being captured at a first range of 9.5°-10.5°, a second range of 10.5°-11.5°, and a third range of 11.5°-12.5°. This is shown in Figure 10, which shows a series of angular brackets, with the size of these angular brackets at a given distance being dependent on the angular resolution. The angular identifier(s) typically comprise a reference to such an angular bracket. Consider, for example, a cube placed with the capture device C1 at the centre ofthis cube. By dividing this cube into x segments at regular azimuth angles and y segments at regular elevation angles, it is possible to identify any angular range of the representation by reference to an x segment and a y segment (and then the space bracketed by this angular range will depend on both the angular resolution (e.g. the angle between adjacent brackets) and the distance of the point from the capture device). Typically, each capture device has the same capture pattern so that the angular bracketing of each device is the same (albeit centred differently at the location of the relevant capture device). For example, in an embodiment with 1000 equal angular brackets, the angle for each bracket may be 360 / 1000. In some embodiments, different capture devices are associated with different capture patterns, where this may be signalled in configuration information relating to the three-dimensional representation. In some embodiments, each capture device is arranged to capture a point for a plurality of angular brackets, where each bracket is associated with a different angle. The angular spread of each bracket (that is, the angle between a first, e.g. left, angular boundary of the bracket and a second, e.g. right, angular boundary of the bracket) may be the same; equally, this angular spread may vary. In particular, the angular spread may vary so as to be smaller for points which are directly in front of (or behind, or to a side of) the capture device. For example, the embodiment shown in Figure 7 shows an angular bracketing system that is based on a cube. With this system, a cube is placed such that a capture device is located at the centre of the cube and the cube is then split into 1000 sections of equal size (it will be appreciated that the use of 1000 sections is exemplary and any number of sections may be used). Each of these sections is then associated with an angular index. With this arrangement, the angular spread of each section (or bracket) varies, as has been described above. Figure 10 shows a two-dimensional square, where each angular bracket of the square is referenced by an index number (between 1 and 100). In a three-dimensional implementation, an angular bracket of a cube could be indicated with two separate numbers (with a first azimuthal indicator that identifies a ‘column’ of the cube and a second elevational indicator that identifies a ‘row’ of the cube). Equally, a singular indicator may be provided that indicates a specific bracket of the cube. Therefore, for a cube that is divided into 1000 elevational sections and 1000 azimuthal sections, the bracket may be indicated with two separate indicators that are each between 0 and 999 or with a single indicator that is between 0 and 999999. It will be appreciated that the use of a cube to define the brackets is exemplary and that other bracketing systems are possible. For example, a spherical bracketing system may be used (where this leads to curve angular brackets). Equally, a lookup table may be provided that relates angular indexes to angles, where this enables irregularly spaced brackets to be used. Typically, determining the location of the point comprises determining the location of the point so as to be at the centre of the angular bracket identified by the angular identifier(s). Texture patches In order to reduce the file size of the three-dimensional representation (and the bandwidth required to transmit the three-dimensional representation) it is desirable to reduce the number of points within the three-dimensional representation. Therefore, referring to Figure 11, there is described a method of determining a texture patch that can replace a plurality of points in the representation. In a first step 71, the computer device identifies a plurality of points of the representation; in a second step 72, the computer device determines that the points lie on a shared plane; in a third step 73, the computer device determines a texture patch based on the attributes of the points; and in a fourth step 74, the computer determines a new point that references the texture patch (this new point may be referred to as a ‘texture point’). The texture patch typically comprises a patch with a plurality of attribute values, which attribute values may be the same as the attribute values of the identified points. Therefore, the texture patch enables the recreation of the plurality of points. A benefit of using the texture patch is that a single point, with a single location value and a (single) reference to the texture point, can replace the plurality of identified points. The attribute values of each point are contained in the texture patch so that little (or no) information is lost from the original representation, but by representing all of these attribute values by reference to the texture patch, only a single location needs to be signalled (saving on the computational cost of signalling locations for a plurality of points). For example, an 8x8 square of identified points that each have separate locations and attribute values may be replaced by a single point with a single location and an attribute value that is a reference to a texture patch (which texture patch comprises the attribute values of the identified points arranged in the relative positions of the identified points); this would reduce the size of the representation by 63 points (where a single point replaces an 8x8 grid of points) at the cost of needing to signal a 8x8 texture patch (that has 64 attribute values and / or transparency values and / or normal values). This is shown in Figures 12a, 12b, and 12c. Figure 12a shows a plurality of points of a three-dimensional representation that lie on a shared plane. Each of these points has a location and an attribute value. Figure 12b shows how these points may be replaced by a single point (e.g. a ‘texture point’) that contains a reference to the texture patch shown in Figure 12c. This texture patch may comprise the attribute values of the plurality of points without separately storing the locations of the attributes (instead, the attribute values are laid out in a predetermined pattern, which is a 5x5 grid in the example of Figure 12c). It will be appreciated that various sizes of texture patch are possible and that the 5x5 grid of Figure 12c is only an example. Another (practical) example of a texture patch is shown in Figure 12d, which shows an 8x8 arrangement of values laid out in the form of a texture patch. As shown in Figure 12d, typically the texture patch provides a continuous grid of pixel values (e.g. that can be used to form a continuous image) - in this regard, the points shown in Figures 12a - 12c are shown as separated points. In practice, these ‘points’ are typically abutting points that form a joined arrangement of values. In some embodiments, the method comprises determining the texture patch in dependence on a difference of the attributes of the identified points exceeding a threshold (e.g. in dependence on a variance, a range, or a maximum difference of these attributes exceeding a threshold). In this regard, points that are similar in both location and attribute may be aggregated into a single point with a location and attribute that is based on the initial points and a size that covers both of the initial points (e.g. two adjacent points of the same colour and a size of 1 may be aggregated into a single point of this colour with a size of 2). Such an aggregation does not require any determination of a texture patch. In contrast, a texture patch may be determined where there is a plurality of dissimilar points (e.g. points with dissimilar attributes) that lie on a shared plane, where the use of the texture patch enables the attributes of each of these points to be signalled in an efficient manner. The second step 72 of determining that the points lie on a shared plane may comprise determining that the points are lie on a shared surface (e.g. on the same object), where the method may comprise identifying a surface associated with the identified points. Determining that the points lie on a shared plane may comprise comparing a distance of (each of) the points from this plane and / or surface to a threshold distance and determining that the points lie on the threshold / plane if they are within this threshold distance from the plane. This second step 72 may also, or alternatively, comprise identifying normals for each of the points, which normals may be contained in point data of the points, and determining a similarity of the normals (e.g. determining that each of the normals is within a threshold value of an average normal and / or determining that a variance of the normals is below a threshold value). The texture patch is typically determined based on this determination in the second step 72, where if (e.g. only if) the identified points lie on a shared surface or plane then they may be replaced by a single point that references a texture patch. In some embodiments, the texture patch may be determined for points that lie on a curved plane, where the second step 72 may comprise determining that the points lie on a curved plane or a curved surface. Such a texture patch may be associated with a bend value to enable the reproduction of the identified points. Typically, the texture patch comprises a quadrilateral, where the texture path may be able to bend about a line that is formed between opposite corners of this quadrilateral so as to map the texture patch to a curved surface. The threshold distance (forthe points to be considered co-planar) may depend on the distance of the points from the viewing zone; in particular, points that are located far from the viewing zone may have a higher threshold separation than points that are located nearer to the viewing zone. Typically, users are better able to identify separations between a plurality of surfaces when these surfaces are near to the viewing zone whereas users may not be able to identify separations between surfaces that are distant from the viewing zone. Therefore, the maximum (threshold) acceptable distance between the identified points and a plane passing through the identified points may be dependent on the distance of the identified points from the viewing zone (e.g. the threshold may increase from a first value when the identified points are within 1 km of the viewing zone to a second value when the identified points are more than 1 km from the viewing zone). Typically, the texture patch is associated with a size, where there may also be provided a plurality of texture patches of different sizes. Identifying the plurality of points may then comprise identifying a plurality of points that could potentially be replaced with a single point that references a texture patch. This may involve the computer device iterating through a plurality of pluralities of identified points and then evaluating each of these pluralities of identified points in order to determine whether the points lie on a shared plane. If these identified points are found to lie on such a shared plane, then a texture patch may be determined based on the attributes of these points and this texture patch may be added to a database of texture patches (e.g. a texture ‘atlas’). The texture patch may be associated with one or more of: one or more attribute values; one or more transparency values; one or more normals; etc. The texture patch may comprise a plurality of points that correspond to the points used to form the texture patch (e.g. where each point of the texture patch comprises an attribute, a normal, and / or a transparency of a corresponding point of the three-dimensional representation). The method may comprise determining a multi-layered texture patch and / or a plurality of texture patches. For example, the method may comprise determining a texture patch for each eye of a user (where these texture patches may be located at the same index of separate databases of texture patches so that they can be signalled by a single reference in a point). In some embodiments, texture patches for each of a left eye and a right eye are stored in a shared database. The index for a texture patch for a first eye may then be set as being one greater than the index for a texture patch of a right eye, where this simplifies the signalling of the texture patches. In some situations, e.g. for diffuse non-reflective materials, each eye may be associated with the same texture patch. In these situations, only a single texture patch may be included in the database (with the index that would otherwise contain a second texture patch instead pointing to the single texture patch). Equally, the same texture patch may be stored twice. By storing the texture patches for each eye adjacent to each other, there is an increased ability to benefit from similarities between these texture patches when encoding the texture patch database. Determining the new point may comprise replacing (one or more of, or all of) the identified points with the new point (the texture point). While the identified points each comprise an attribute value (e.g. to indicate the colour of the point), the new point may instead comprise a reference to the texture patch. This reference may be located in an attribute datafield associated with the point so that the new point has the same form as the identified points (and the same form as the other points of the three-dimensional representation). In some embodiments, determining the new points comprises modifying one of the identified points. In particular, the attribute value of one of the identified points may be replaced with a reference to the texture patch. Furthermore, the size of this point may be modified, e.g. increased, to signal that the point is now associated with a texture patch. Typically, the first step 71 of identifying the plurality of points comprises identifying a plurality of points associated with the same capture device (and / or similar, e.g. adjacent capture devices). In this regard, as has been described above, each points is typically associated with a capture device, a distance (from that capture device) and one or more angles (from the capture device), where this method of defining points enables the location of each point to be determined relative to the associated capture device (and then absolute locations of each point can be determined using the locations of the various capture devices). Typically, each point of the representation comprises a size, which size may relate to a number of angular boundaries that is encompassed by the point. Therefore, a point that has a size of 1x1 may relate to a point that covers a single angular bracket, a point that has a size of 2x1 may cover two angular brackets in a row etc. Typically, the size is stored via an index so that, for example, a size value of 0 (which may be signalled by a binary value of 000) may signal a point that covers a 1x1 arrangement of angular brackets, a size value of 1 (which may be signalled by a binary value of 001) may signal a point that covers a 2x1 arrangement of angular brackets, a size value of 2 (which may be signalled by a binary value of 010) may signal a point that covers a 1x2 arrangement of angular brackets, etc. Typically, the texture patch is similarly associated with a size, where the method of determining texture patches may be performed such that every texture patch is of the same size (where this can simplify the storage of the texture patches). In some embodiments, a computer device determining the texture patches may be arranged to determine texture patches of different sizes. The size of the texture patch may refer to a number of angular brackets covered by a texture patch. For example, each texture patch may cover an 8x8 arrangement of angular brackets. The first step 71 of identifying the points may comprise identifying a plurality of contiguous points, where these points may be in a predetermined arrangement. For example, the first step may comprise identifying an 8x8 arrangement of points in adjacent (and contiguous) angular brackets. In some embodiments, the three-dimensional representation comprises a plurality of separately signalled points and / or comprises a plurality of different sections, with each section comprising a different type of points. In particular, the representation may be associated with a file that has at least two sections, where a first section comprises points for which an attribute datafield contains an attribute value and a second section comprises points for which an attribute datafield comprises a reference to an external value (e.g. to a texture patch). Equally, those points for which the attribute value references a texture patch may comprise an identifier that enables these points to be distinguished from points for which the attribute datafield comprises an attribute value (e.g. each attribute datafield that references a texture patch may begin with a recognisable string or pattern of bits). In some embodiments, points that reference a texture patch are identified based on a size of that point, where typically these points have a greater size than other points (equally a size of 0 may be used to signal a texture patch where this size would not otherwise be used). In some embodiments, the three-dimensional representation comprises at least two of the following sections: A first section comprising opaque points with attribute values (e.g. points for which a transparency (a) value of 255). A second section comprising opaque points that reference texture patches (e.g. points that have an attribute datafield that comprises a reference to a texture patch and also a transparency value of 255). A third section comprising transparent points with attribute values and transparent points that reference texture patches. Typically, the transparency values for the texture patch are a part of the texture patch (e.g. these transparency values are stored in the texture atlas). The points with attribute values and the points referencing texture patches may then be distinguished by setting the transparency values of the points referencing texture patches to 255. Since the points are in the third section, a computer device parsing the three-dimensional representation is able to determine that a point with a transparency value of 255 is not opaque (since otherwise it would be located in the first or second section) and therefore the computer device can identify that such points include references to texture patches). It will be appreciated that the use of a value of ‘255’ to denote an opaque point is purely exemplary. More generally, points referencing transparent texture patches may be signalled by including these points in a section associated with transparent points and transparent texture patches and setting a transparency value of the points referencing transparent texture patches to a value that indicates an opaque point. Such a method of signalling points that reference texture patches enables an increase in the efficiency of the storage of the three-dimensional representation. In some embodiments, one or more ofthe points of the three-dimensional representation is associated with a movement vector, where this vector indicates a motion of that point that occurs between frames of a video associated with the three-dimensional representation. A texture patch may similarly be associated with a single vector (e.g. where the texture patch is a rigid structure and so moves as a single structure). Equally, a texture patch may be associated with a plurality of movement vectors, e.g. each corner ofthe texture patch may be associated with a different movement vector, where this enables the movement vector to flex and to change in size during the course of a video (and enables this flexing to be signalled via movement vectors). Referring to Figure 13, there is described a method for determining one or more texture patches for a three-dimensional representation. This method is typically performed by a computer device as a post-processing step, after the three-dimensional representation has been generated (e.g. after each point ofthe three-dimensional representation has been captured). This method of Figure 13 may then be used to replace a number of points within this three-dimensional representation with one or more new points that reference texture patches so as to reduce the number of points in the three-dimensional representation. In a first step 81, the computer device identifies a plurality of points associated with a given location in the representation. Typically, this step comprises identifying a plurality of points in adjacent angular brackets. For example, the computer device may identify an 8x8 arrangement of points (it will be appreciated that various other arrangements may be identified). In a second step 82, the computer device determines whether these identified points lie on a shared plane. This determination may comprise identifying a plane that passes through these identified points and determining that each ofthe identified points is within a threshold distance of this plane. The plane may comprise a curved plane. Additionally, or alternatively, the second step may comprise comparing the normals ofthe identified points to determine that the identified points have similar directional values (e.g. face in the same direction). In a third step 83, if the points lie on a shared plane, the computer device determines a texture patch based on the identified points; e.g. where the texture patch comprises a two-dimensional plane that has attribute values corresponding to the attribute values ofthe identified points. The computer device may then replace the identified points with a new point that has a location related to the identified points and that references the texture patch. In a fourth step 83, the computer device considers a next location in the representation, and the method then returns to the first step 81 so that a next plurality of points can be evaluated. This method may be performed so as to move through a plurality of pluralities of points ofthe representation and to determine, for each of these pluralities of points, whether the points lie on a shared plane (and so can be replaced with a texture quad). For example, the fourth step 84 of considering a next location may comprise incrementing an angular identifier so as to move incrementally through the angular brackets for a computer device and, at each stage, to compare the points in a number of adjacent angular brackets. In this way, the computer device essentially initialises a window of points and then moves the window so as to evaluate a moving window of points that passes about the entirety ofthe representation. Typically, the method of Figure 13 is performed separately for each capture device so that, for one or more capture devices, one or more pluralities of points are evaluated to determine whether these points lie on a shared plane and to determine a texture patch if the points do lie on a shared plane. In some embodiments, points captured by separate capture device may be considered together (e.g. points captured by adjacent capture devices may be considered together for the purposes of determining texture patches). The iterative process described with reference to Figure 13 is capable of obtaining a three-dimensional representation that efficiently represents a space. Typically, the three-dimensional representation is associated with a video, e.g. a VR video, and so the three-dimensional representation (and the texture atlas associated with this three-dimensional representation) may relate to a single frame of the video. The video may be composed of a plurality of frames, with each frame relating to a different three-dimensional representation. By iterating separately over each three-dimensional representation using the method of Figure 11 it is possible to obtain efficient representations for each frame of the video. However, this method of determining texture patches can result in similar features being sorted into different representations in each frame, which may prevent the use of temporal encoding methods. In this regard, the texture patch may represent a point that is moving throughout a scene. And so the same (or a similar) texture patch may be identified in a plurality of three-dimensional representations, e.g. where a movement of this texture patch is signalled by one or more movement vectors of this texture patch. If each representation is considered separately, then a texture patch that could be present in each of a first and second three-dimensional representation may not be identified in the second three-dimensional representation. For example, in the second three-dimensional representation, a number of the constituent points of this texture patch may instead be included in a different texture patch. As a result, the processing of the entire representation might change based on a change to only a few points (since any formation of a new texture patch will have a knock-on effect throughout the three-dimensional representation). Therefore, referring to Figure 14, there is envisaged a grid-based method of processing the representation that can be used to ensure isolated changes in a first area of the scene do not cause wholesale changes in the processing of a three-dimensional representation. While this grid-based system is useable for the aggregation process, it will be appreciated that more generally the grid system may be used as the basis for any processing procedure. In a first step 91, a grid system of a representation is determined. For example, where 8x8 arrangements of points may be evaluated and potentially replaced with a point referencing a texture patch, a grid may be determined that has grid sizes of 8x8. Equally, each grid square may, for example have a size of 16x16 (where the size indicates a number of angular brackets covered by the grid square). Typically, the method comprises determining a grid system such that each of a plurality of three-dimensional representations (e.g. relating to frames of a video) has a similar grid system. This may comprise determining the grid system such that a first grid square in a first three-dimensional representation is located at the same space in the scene as a first grid square in a second three-dimensional representation. A grid system is typically determined for each of the capture devices, where the grid system may be based on the angular brackets of that capture device. Therefore, for each of the three-dimensional representations (e.g. for each of a plurality of three-dimensional representations associated with a certain viewing zone), a grid system may be determined for each of the capture devices such that the grid squares of each grid system cover the same angular brackets for each three-dimensional representation. These grid squares can then be processed separately. If an object is moving in the first grid square, then the process of determining texture patches that occurs for this first grid square will likely differ in successive three-dimensional representations (relating to successive frames), but with this grid based system, if there are no moving objects in the second square then the determination of texture patches in the second grid square will likely be the same in these successive frames (and will not be affected by the object moving in the first grid square). For each grid square the process of determining texture patches mirrors that described with reference to Figure 13. That is, for each grid square, the method involves identifying a plurality of points in a second step 92, determining whether these points lie on a shared plane in a third step 93, if the points do lie on a shared plane then determining a texture patch in a fourth step 94, and then considering a next location in the grid square in a fifth step 95. Once, in a sixth step 96, the computer device determines that all possible pluralities of points in the grid square have been identified, then in an eighth step 98, a next grid square is considered and this process is repeated. As described above, the grid squares typically cover the same angular brackets in a plurality of three-dimensional representations that are associated with the same video and / or the same viewing zone. Equally, in some embodiments, the grid squares may be associated with a movement vector so that the grid system may differ (in a calculable way) between the three-dimensional representations. Such embodiments can enable the efficient encoding of three-dimensional representations that are associated with consistent movements. While the above method has referred to ‘grid squares’, it will be appreciated that the grid system may be associated with subdivisions of any shape or size (these subdivisions being consistent through a plurality of three-dimensional representations). The determination of the texture patches and of an aggregate point may be a part of the same process. For example, the computer device may identify a plurality of points with a similar location (e.g. being captured by the same capture device and being located in adjacent angular brackets) and the computer device may: in dependence on the points having similar attribute values, determine an aggregate point based on these points; in dependence on the points having different attribute values, determine a texture point (and a texture patch) based on the points. Movement vectors The three-dimensional representation is typically used to render one or more two-dimensional images in order to form a video that can be presented on the display device 17. In particular, a computer device (e.g. the display device 17) may render a series of two-dimensional images for each eye of a user, where a user viewing these two series of images simultaneously is provided with the impression of a three-dimensional scene. Referring to Figure 15a, the series of images typically comprises a plurality of frames of a video, e.g. first to sixth frames F1 ... F6. By showing these frames in succession, the display device is able to provide a viewer with the impression of an unbroken video. Achieving this impression of an unbroken video requires the viewer to be shown a number of frames every second. The exact rate / density at which the frames F1 ... F6 are shown typically depends on a frame rate (or a refresh rate) of the display device 17 and / or a frame rate defined by a creator of a video. Typically, the display device is arranged to display video with a frame rate of at least 60 frames per second (60 ‘fps’), e.g. the frame rate of the display device may be 72 fps. In some embodiments, the image generator 11 is arranged to generate three-dimensional representations at a corresponding rate so that each three-dimensional representation is associated with a single frame of a video. Therefore, if the refresh rate of the display device 17 is 72 fps then the image generator may be arranged to generate 72 three-dimensional representations for every second of video (e.g. to generate three-dimensional representations at 72 fps). However, each three-dimensional representation is typically large in size so that a file that contains this rate of three-dimensional representations might be prohibitively large (leading to difficulties in storing and / or transmitting the file). Furthermore, since the scene may be viewed on a plurality of different display devices with different refresh rates, generating a set of three-dimensional representations that enables each three-dimensional representation to be associated with a single frame of a video for each of the plurality of display devices can require generating a prohibitively large number of sets of three-dimensional representations. Therefore, the present disclosure considers methods of rendering a two-dimensional video, where a number of frames of the two-dimensional video is different to (e.g. greater than or less than) a number of three-dimensional representations used to render these frames. This may be considered as the frame rate of the two-dimensional video being different to the frame rate of the three-dimensional representations. Referring to Figure 15b, the three-dimensional representations may be generated and / or rendered so as to provide a two-dimensional video that has a frame rate that is a multiple of a frame rate of the three-dimensional representations, e.g. where the video comprises six frames F1 ... F6, the image generator 11 may be arranged to generate three three-dimensional representations R1, R2, R3 where each three-dimensional representation is associated with a pair of two-dimensional images. For example, the three-dimensional representations may have a frame rate of 36 fps (i.e. 36 three-dimensional representations may be generated for each second of the scene) and these three-dimensional representations may be used to render a video with a frame rate of 72 fps. More generally, referring to Figure 15c, the present disclosure envisages methods by a plurality of three-dimensional representations of a scene may be used to render a two-dimensional video of any frame rate. This enables the three-dimensional representations to be used with various devices and for various situations. With the example of Figure 15c, the first second and third video frame F1, F2, F3 may be rendered based on the first three-dimensional representation R1, the fourth and fifth video frame F4, F5 may be rendered based on the second three-dimensional representation R2 and the sixth video frame F6 may be rendered based on the third three-dimensional representation R3. For example, the three-dimensional representations may have a frame rate of 24 fps (i.e. 24 three-dimensional representations may be generated for each second of the scene) and these three-dimensional representations may be used to render a video with a frame rate of 70 fps. More generally, it will be appreciated that any pairing of frame rates may be provided using the methods described herein. A simple solution to this problem of differing frame rates is to repeat a frame of the video. For example, the first frame F1 that is rendered based on the first three-dimensional representation R1 could also be used as the second frame F2 and the third frame F3. However, rendering a video in this way can leads to stuttering and ghosting. Regarding ghosting, the brain of a viewer expects an object to move in a consistent manner so that copying frames in the above-described manner can result in a viewer seeing objects in the second and third frames at both expected locations (the locations that would be expected from a previously movement of the objects) and also at the displayed location (the location of the object in the repeated first frame). Therefore, simply repeating frames in order to match the frame rate of the three-dimensional representations to the frame rate of a rendered two-dimensional video is typically unsatisfactory (unless there is very little movement between successive representations). Referring to Figures 16a - 16c, there is described a method of rendering a two-dimensional image from a three-dimensional representation that avoids the problems of stuttering and ghosting. This method is performed by a computer device, for example the image generator 11 or the display device 17. In a first step 101, the computer device identifies a point of a three-dimensional representation. As described above, the points typically comprise a location and at least one attribute value (e.g. a colour). Furthermore, the points may comprise a movement vector that identifies a motion of the point (or a ‘movement vector’, these terms are used interchangeably in this document). This is shown in Figure 16b, which figure shows a point that has a first location L1 and a first movement vector MV1. In a second step 102, the computer device identifies this first (initial) location L1 and identifies the first movement vector MV1. The initial location is typically a location defined in point data of the point. In a third step 103 that is shown in Figure 16c, the computer device identifies a second, predicted, location L2 of the point based on the first location L1 and the first movement vector MV1. In a fourth step 104, the computer device renders a two-dimensional image based on the predicted location L2 ofthe point. In some embodiments, the first movement vector MV1 indicates an end location ofthe point, that is the movement vector indicates where the point will be at the end of a time period associated with the three-dimensional representation. The predicted location L2 can then be determined by interpolating between the first location L1 and the end location. In practice, the three-dimensional representation is typically associated with a first time period, e.g. t=10 to t=11, with the first location L1 being at t= 10 and the first movement vector MV1 indicating the end location ofthe point at t=11. In order to find the location of the point at t=10.25, the computer device may interpolate between the first location and the end location so that the predicted location L2 is between these locations and is 25% ofthe way along a line drawn between these locations. The computer device may then consider a further three-dimensional representation that is associated with a second time period, e.g. t=11 to t=12 so as to form a continuous video. This method enables the rendering of a two-dimensional video with a first frame rate based on three-dimensional representations with a second frame rate (which second frame rate is typically smaller than a frame rate ofthe two-dimensional video). Typically, this involves rendering a plurality of frames ofthe two-dimensional video based on a single frame ofthe three-dimensional representation. In a simple example, the first frame F1 may be rendered based on the first (e.g. defined) location of each point in the first three-dimensional representation R1 - e.g. each of F1 and R1 may relate to the same time t=0. Thereafter, for example at time t=0.1, the second frame F2 may be rendered based on the three-dimensional representation. For this second frame F2, the predicted location L2 of each point may be determined based on the first locations L1 and the first movement vectors MV1 of those points. As described in Figure 16a, the method may comprise rendering a two-dimensional image based on the predicted locations L2 of points in a first three-dimensional representation. Additionally, or alternatively, the method may comprise determining a further (e.g. intermediate) three-dimensional representation based on the predicted locations ofthe points ofthe three-dimensional representation. For example, based on the movement vectors, a three-dimensional representation may be formed for each frame of a video. This enables a set of three-dimensional representations (and / or a set of two-dimensional images) to be generated for a video with any frame rate. It will be appreciated that in the description below any features described with reference to the generation of an intermediate three-dimensional representation may equally be applied to the rendering of a two-dimensional image (and vice versa). Typically, the first frame F1 is associated with an earlier time than the second frame F2 and the first movement vector MV1 define a ‘forwards’ movement ofthe point (e.g. defines a direction in which the point is moving). Additionally, or alternatively, the movement vector MV2 may comprise a ‘backwards’ movement vector that indicates a backwards movement (e.g. that defines a direction from which the point has come). Therefore, referring to Figure 15c, the third frame F3 may be determined based on locations and movement vectors from either of, or both of, the first three-dimensional representation R1 and the second three-dimensional representation R2. Any combination of forwards and backwards movement vectors may be used to determine the locations of points in the frames of the video. For example, a first location L1 and a first, forwards, movement vector MV1 from the first three-dimensional representation R1 may be combined with a second location and a second, backwards, movement vector from the second three-dimensional representation in order to determine the predicted location of a point in the frames F2 and F3. This determination may involve a weighted combination of locations and / or movement vectors from a plurality of frames. The backwards movement vector for each point may be the inverse of the forwards movement vector (this implementation is generally used only when the velocity of the point is constant across a time period associated with a plurality of three-dimensional representations). Equally, the backwards movement vector may be a different vector where each point may be associated with one or more of a backwards movement vector and a forwards movement vector. Therefore, the location of the point in the second frame F2, the ‘predicted’ location, may be determined as any one or more of: Lpred = T1 + MV1 * Lpred = ^2 — ^^2 * ^2 ^pred = fl(^i + * tj) + b(L2 — MV2 * t2), where a + b = 1 Where: Lpred = the predicted location of a point in an intermediate three-dimensional representation and / or in a two-dimensional image (e.g. a frame of a video). Lt = the location of the point in a first three-dimensional representation before the intermediate three-dimensional representation. L2 = the location of the point in a second three-dimensional representation after the intermediate three-dimensional representation. MVt = the movement vector of the point in a first three-dimensional representation. MV2 = the movement vector of the point in a first three-dimensional representation (it will be appreciated that the sign of the movement vector, i.e. whether the movement vector is positive or negative, is not fixed and is selected in dependence on the exact implementation). is the time between the first three-dimensional representation and the intermediate three-dimensional representation. t2 is the time between the second three-dimensional representation and the intermediate three-dimensional representation. a and b are (optional) weighting factors. Typically, the movement vector has units of velocity (e.g. m / s) so that this movement vector can be combined with a time in order to obtain a (vector) movement of a point. Typically, each three-dimensional representation and each frame is associated with a time, so that the method of Figure 16a may involve determining a time difference between a three-dimensional representation and a frame and determining the predicted location L2 in dependence on this time difference. Typically, the movement vector is a single value (e.g. 5m / s). In some embodiments, the movement vector comprises a range of values, a formula, and / or a reference to a range of values. For example, the movement vector may relate to a speed value of: 5m / s + 0.1t. Such embodiments enable the signalling of movement vectors that change over time (e.g. to indicate that an object is accelerating). In practice, the frame rate of the three-dimensional representations is typically high enough that changes in the velocity of a point in between representations are insufficient to cause problems with video quality. Therefore, using movement vectors that are a single (constant) value typically provides sufficient quality while maintaining an efficient three-dimensional representation. In some embodiments, e.g. those with many erratically moving objects, the use of variable movement vectors (e.g. that have an acceleration and / or a dependence on time) can cause a sufficient improvement in video quality to justify the increase in data that is required to signal these movement vectors. As described above, the location of each point is typically stored in dependence on a capture device used to capture that point. The movement vectors may similarly be stored with reference to a capture device, where, for example, the movement vectors may identify a number of angular brackets through which a point moves in a unit time. Typically, the movement vectors are stored in absolute values. For example, the movement vectors may be stored with x, y, z values based on a cartesian grid or as a speed, an elevation angle, and an azimuthal angle based on a grid. In this regard, typically the scene is associated with a grid that, for example, has a center at the center of the viewing zone and that is aligned with the viewing zone. Therefore, determining the predicted location L2 of the point may comprise converting the first location L1 into an absolute value (e.g. determining a coordinate value for L1 based on a location of the capture device used to capture the point and the distance / angle of the point from this capture device) and then combining this absolute location value with the determined motion of the point (this motion being determined based on the movement vector and the time difference between the first three-dimensional representation and the intermediate three-dimensional representation). The movement vectors of each point of a three-dimensional representation are typically determined during the formation of that three-dimensional representation. In this regard, the three-dimensional representation may be determined based on a computer-generated model (e.g. an animated scene), in which case the objects in this model may be associated with movements. The movement vectors may then be determined directly from the model. For example, the description above has described a method of capturing point data based on scanning rays that are sent by a capture device. These scanning rays may be used to determine each of: a location; an attribute; and a movement vector of a surface impacted by the scanning rays. In some embodiments, the scene represents a real scene where the point data may be captured using (real) cameras. In such embodiments, the movement vectors may be determined based on an analysis of the frames captured by the cameras (e.g. a movement of an object between subsequent frames may be determined and this movement may be used to determine a movement vector). In some embodiments, the determination of the movement vectors uses artificial intelligence or machine learning algorithms. In some embodiments, the movement vectors are determined or modified (e.g. refined) based on a postprocessing process in which similar points in different three-dimensional representations are processed in order to determine a movement vector for these points. For example, the locations of corresponding points in each of the first three-dimensional representation R1 and the second three-dimensional representation R2 may be compared to determine a suitable movement vector (e.g. to determine a forwards movement vector for inclusion a point of the first three-dimensional representation). Such embodiments typically require the representations to be evaluated to determine corresponding points in successive representations; this may involve determining points with similar attributes in successive representations. Viewing zone movement In some embodiments, the viewing zone 1 is arranged to move and so the viewing zone may be associated with a movement matrix or vector. For example, the viewing zone may be moving forward and / or rotating within the scene, where this causes a relative movement between each of the points of the three dimensional representation and the viewing zone (even points that would otherwise be stationary in the three-dimensional representation).In some embodiments, the viewing zone movement matrix is signalled using one or more (or all) of the following components: a translation movement vector, a rotation element (e.g. in the form of a quaternion), and a scale change. Such a representation enables a computer device to determine appropriate relative movements for each point of the representation based on the viewing zone movement matrix. In this regard, in these embodiments, the locations of one or more points in the three-dimensional representation at a second time may be predicted based on the locations of the points at a first time and based on a movement matrix associated with the viewing zone. The prediction of these points may occur using a process similar to that described above, with the viewing zone movement matrix being inverted to determine a suitable movement vector for applying to each points. In some embodiments, the movement of the viewing zone is instead effected by adding a viewing zone movement matrix to each of the movement vectors of the points in the scene (and then storing these points with this movement vector). However, such an implementation can result in every point in the scene needing a movement vector and this is typically less efficient than associating the viewing zone with a viewing zone movement matrix that can be applied to the points at the time of rendering the scene. In some embodiments, one or more points of the three-dimensional representation is arranged to move with the viewing zone. For example, the viewing zone may be located inside a virtual cockpit, where certain points surrounding the viewing zone are also part of the cockpit (and these points move with the viewing zone). In orderto identify these points, one or more (or each) points of the three-dimensional representation may be associated with a flag (e.g. a ‘cockpit flag’), which flag defines whether a viewing zone movement vector should be applied to the points. Equally, these points may be associated with movement vectors that are defined so as to counteract the viewing zone movement matrix where this precludes the need for the cockpit flag, but requires numerous movement vectors to be added to the three-dimensional representation that would not be needed otherwise. Furthermore, due for example to rounding errors that may occur during the determination of the predicted locations of each point, embodiments that use counteracting movement vectors may result in some relative movement between points. Therefore, the cockpit flag typically provides a benefit in embodiments where there is a viewing zone movement matrix. The viewing zone movement matrix may comprise a single vector so that the viewing zone moves through the scene as a fixed shape. Equally, the viewing zone movement matrix may comprise a rotational component and / or a deforming component. Movement vectors - texture patches As described above, the three-dimensional representation may comprise one or more texture points, which texture points reference a texture patch that is stored in a texture atlas. The texture points may be associated with a single movement vector, where this enables a predicted location of the texture points to be determined based on an assumption that the texture patch moves as a solid object. In such embodiments, predicted locations of each pixel of the texture patch may be determined based on this movement vector, where each pixel moves in the same way. In some embodiments, the texture points (and / or texture patches) are associated with a plurality of movement vectors where this enables the texture patches to deform between the first three-dimensional representation and the intermediate three-dimensional representation. In particular, the texture point and / or the texture patches may be associated with (e.g. the texture point may include) movement vectors for each of the (e.g. four) corners ofthe texture point / texture patch. This enables the texture point / texture patch to deform, e.g. to stretch or shear, between the first three-dimensional representation and the intermediate three-dimensional representation. The method of determining the texture point may then comprise determining one or more movement vectors for the texture point. Referring to Figure 17, there is described a method for determining one or more movement vectors for a texture point. In a first step 111 (e.g. the first step 71 of Figure 11), the computer device identifies a plurality of points. In a second step 112, the computer device determines movement vectors for one or more of the plurality of points. In a third step 113, the computer device determines a texture point and a texture patch based on the identified points (e.g. as described with reference to Figure 11). In a fourth step 114, the computer device determines one or more movement vectors for the texture point (and / or texture patch) based on the determined movement vectors. In some embodiments, determining the movement vector(s) for the texture point (the ‘texture vector(s)’) comprises combining the movement vectors of the constituent points of the texture point (the ‘constituent vectors’). For example, the texture vector may be determined as an average, or a weighted average, of the constituent vectors. In some embodiments, movement vectors are determined for a plurality of sections of the texture point, e.g. for the corners of the texture point. This may comprise dividing the texture patch into sections and combining the movement vectors of the points in each section. Referring to Figure 18, there is shown an exemplary texture patch. Each point that is associated with the texture patch (e.g. each point that is replaced by the texture point) is typically associated with a movement vector. In order to obtain a movement vector for the texture patch, the computer device may be arranged to combine (e.g. take an average of) the movement vectors of each constituent point. As shown in Figure 18, the texture patch may be divided into four quadrants (e.g. quarters), where a movement vector may be determined for each quadrant. This movement vector for each quadrant may be determined based on a (e.g. weighted) average of the movement vectors of the points within the quadrant. The weighting for each movement vector may, for example, depend on how far an associated point is from a corner of the movement vector. More generally, the texture point / texture patch may be divided into a plurality of sections, where each section is associated with a respective movement vector and the rendering of the texture point is dependent on these respective movement vectors. In order to render the two-dimensional image, the computer device may determine a change in the shape of the texture point and / or texture patch based on the movement vectors of the texture patch. For example, the texture point may grow, shrink, deform, shear etc. Typically, this movement is worked out following a transformation of the location of the texture point into absolute coordinates. In particular, the texture point may be converted from relative coordinates (relative to a capture device) into absolute coordinates; the attribute values for each pixel of the texture patch associated with the texture point may be rendered so as to provide a plurality of derived points that are located in the representation based on the location of the texture point; and then a deformation of these derived points may be determined. In some embodiments, the determination of the texture patch is dependent on the movement vectors of the identified points (e.g. the points identified in the first step 111 of Figure 17 or the first step 71 of Figure 11). In particular, the determination of the texture patch may depend on a similarity of the movement vectors of the identified points, where the texture patch (and texture point) are only generated if this similarity of movement vectors is below a threshold value (e.g. if a variance of the movement vectors is beneath the threshold value). Such embodiments can be used to ensure that only related points are combined into a texture patch (e.g. to avoid situations where a plurality of points lying on a shared plane begin moving in different directions that cannot reasonably be modelled via a deformable texture patch). Formation of a bitstream The processing of a three-dimensional representation, and the rendering of two-dimensional images, that has been described above results in a three-dimensional representation and / or a two-dimensional image that may then be stored and / or transmitted by a computer device. In particular, the three-dimensional representation and / or the two-dimensional image may be encoded in a bitstream that is transmitted to another device and / or may be used to render one or more two-dimensional images that can be encoded in a bitstream that is transmitted to another device. This bitstream can then be decoded by this other device (e.g. the display device 17) in order to extract the three-dimensional representation(s) and / or the two-dimensional image(s) from the bitstream. The present disclosure envisages a bitstream that contains or is determined based on an intermediate three-dimensional representation that comprises one or more points with predicted locations, the predicted locations being determined based on movement vectors of those points. Figure 19 shows a schematic of such a bitstream. Specifically, Figure 19 shows a bitstream that comprises a plurality of bits Bit-a to Bit-d. These bits signal one or more points of a three-dimensional representation and / or that defines one or more immersive two-dimensional images. This bitstream can be decoded by a decoding device and this may allow the three-dimensional representation and / or the two-dimensional images to be extracted from the bitstream. In some embodiments, the bitstream comprises one or more flags that indicate features of the bitstream and / or of the three-dimensional representations or two-dimensional images signalled by the bitstream. For example, the bitstream may comprise one or more flags that indicate: whether movement vectors are present in the three-dimensional representation; whether a viewing zone movement vector is present in the three-dimensional representation; whether (for each point), a point is arranged to move with the viewing zone (e.g. cockpit flags); whether (for each point) that point should be affected by a viewing zone movement and, e.g. to what extent (e.g. there may be a scaling factor applied so that the viewing zone movement vector is scaled before being combined with a point, and this scaling factor may be different for different points); an accuracy of the movement vectors and / or a quantisation curve associated with the movement vectors; a time spacing of three-dimensional representations in the bitstream (e.g. so that a time period for each three-dimensional representation can be determined); and a unit type of the movement vectors signalled in the bitstream. Bitstream As described above, in order to store or transmit the three-dimensional representation, a computer device may generate a bitstream that comprises one or more points of the three-dimensional representation. This bitstream can then be received and parsed by another computer device. The bitstream may additionally, or alternatively, comprise one or more texture patches and / or one or more movement vectors associated with a point or a texture patch. Such a bitstream is shown in Figure 19, which shows a bitstream that comprises a plurality of bits Bit-a, Bit-b, Bit-c, Bit-d. It is desirable for this bitstream to be accurate (e.g. to enable the points / texture patches / movement vectors to be accurately received and regenerated) and also to be efficient (e.g. to be small in size). In some embodiments, a computer device is arranged to encode and / or decode the bitstream, where this may comprise compressing the bitstream (e.g. using entropy encoding). The bitstream, and the various structures / components defined in the bitstream may be arranged to improve the efficiency of this encoding. A number of options for defining components (e.g. points, movement vectors, texture points, texture atlases) of the three-dimensional in a bitstream are described below. It will be appreciated that throughout the below description, each disclosed field is optional and that implementations that exclude any of the disclosed fields are also possible (e.g. with this information being signalled separately or not being used in order to reconstruct the three-dimensional representation). For example, a texture point is described below that comprises a diagonal orientation field and a bend field that enable the texture point to bend around a bending line - such bending is not a requirement of the implementation of a texture point and so texture points may be implemented without such a bend field. Equally, a computer device may be arranged to parse a texture point (with or without a bend field) while not having the functionality to implement any defined bending. Such a computer device may simply ignore any fields in the texture point that relate to the bending of the point. Typically, the points (and the texture points) of the three-dimensional representation are arranged to be rendered using a cube mapping process in which views from a viewpoint onto six faces of a surrounding cube are determined. With such representations, each point may cover a rectangular point of space and / or an angular bracket (that is, e.g. perpendicularto a ray from the viewing zone). Therefore, a face of the cube can be determined by combining rectangular points from various angular brackets. It will be appreciated that other methods of rendering are possible, e.g. ecospherical renderings. The signalling of the points may depend on the intended rendering method; for example, where the intended rendering method is a cubemap, a direction of movement vectors may be signalled using x and y coordinates and a magnitude / length; where the intended rendering method is an ecospherical method, then the direction of movement vectors may be signalled using a phi and theta coordinates and a magnitude / radius. Points Referring to Figure 20a, there is shown an exemplary arrangement of fields that may be used to signal (e.g. define) a point (or a ‘particle’) of the three-dimensional representation. It will be appreciated that this arrangement of fields is purely exemplary and that the definition of a point in the bitstream may comprise any one or more of these fields (and the bitstream may comprise further fields). Typically, the point comprises at least (an indication of) a location and (an indication of) an attribute value. Typically, each point in the bitstream comprises one or more of: A transparency, or ‘alpha’ field 1002. The transparency field indicates a transparency (or an opacity) of the point and is typically arranged to be a value between 0 and 255. The transparency field may comprise an 8 bit field. An attribute field 1004. The attribute field indicates an attribute value of the point. Typically, this attribute is a colour of the point. Various arrangements are possible for indicating the attribute; for example, the attribute field may indicate a chrominance, luminance, and luma of the point or the attribute field may indicate a red value, a green value, and a blue value of the point or the attribute field may identify a location of a point in a colour space such as an International Commission on Illumination (CIE) colour space, e.g. the CIE 1931 colour space. Typically, the attribute field comprises a plurality of sub-fields, e.g. a ‘Y’ field of 10 bits, a ‘y’ field of 7 bits, and a ‘x’ field of 7 bits. Typically, the attribute field 1004 indicates a single attribute value for the point (e.g. a single colour), where this single value may be formed from a plurality of sub-fields (e.g. ‘Y’, ‘y’, and ‘x’ fields) that can be combined to obtain the value. In some embodiments, the attribute field indicates a plurality of attribute values and / or indicates an attribute function. For example, the attribute field may indicate a value that changes over time or that depends on a viewing angle of the point. A capture device field 1006. The capture device field indicates a capture device associated with the point. Typically, the capture device field comprise a numerical identifier or capture device index, where this index can be used to identify a capture device from a list of possible capture devices. In this regard, the bitstream may comprise a header section that comprises a list of capture devices (or equally, this information may be sent separately to the bitstream). This list may relate a capture device index to a location in a three-dimensional space so that the index can be used to identify the location of a capture device identified with the point. The capture device field may comprise a field of 8 bits. Typically, the capture device field 1006 identifies the capture device used to capture the point. Equally, the capture device field may identify a different (real or virtual) capture device, so that the location of the point may be defined relative to a capture device that was not used to capture the point. A computer device may be arranged to identify a location of a point based on an initial capture device (e.g. the capture device used to capture the point) as well as the distance and angle of the point from this initial capture device and then to determine the location relative to a different capture device. The capture device field 1006 (and the distance and angle fields) of the point may then be updated so that the location of the point is defined relative to this different capture device. By re-defining one or more points, the number of different capture devices that need to be signalled in the list of capture devices can be reduced. In a simple example, points could be captured using many (e.g. 100 capture devices) arranged around the viewing zone 1 and then each point could be redefined based on a virtual capture device located at the centre of the viewing zone 1. With this example, the list of capture devices would only need to signal information for this single virtual capture device. Typically, a plurality of capture devices are required to accurately store the points (e.g. to account for occlusions in the scene) and so typically the list of capture devices comprises more than one device and the bitstream comprises a plurality of points with different capture device fields. A further attribute field 1008. As with the attribute field 1004, the further attribute field indicates an attribute value of the point. The attribute field is typically associated with a first eye of a viewer and the further attribute field is typically associated with a second eye of a viewer. The further attribute field may be arranged to signal an attribute in the same ways as the attribute field. As with the attribute field, typically the further attribute field comprises a plurality of sub-fields, e.g. a ‘Y’ field of 10 bits, a ‘y’ field of 7 bits, and a ‘x’ field of 7 bits. In some embodiments, the further attribute field 1008 is arranged to signal a difference between a value of a further attribute for a point (e.g. a value of an attribute for a second eye) and a value of an (first) attribute of the point (e.g. a value of an attribute for a first eye), where the first attribute value is signalled in the attribute field 1004. For example, the further attribute field 1008 may comprise sub-fields indicating a difference to ‘Y’, ‘y’, and ‘x’ values defined in the attribute field 1004. In some embodiments, the further attribute field 1008 signals the further attribute value as an exclusive or (XOR) value of the further attribute value and the attribute value signalled in the attribute field 1004. Therefore, if the attribute value and the further attribute value are the same, the value of the further attribute field is 0. In practice, this may comprise filling each sub-field of the further attribute field 1008 with zeros (which can then be efficiently encoded). In some embodiments, the further attribute field 1008 is arranged to be combined with the attribute field 1004 in order to determine an attribute value of a point. In particular, where the attribute is a complex value (e.g. a very precise colour), the attribute field may not contain enough bits to accurately signal this value and so the further attribute field may provide further space. A type field 1010 (e.g. a glass field). The type field may comprise a flag, e.g. a 1 -bit flag and / or may comprise an index. The type field indicates a type of the point, where this type may affect the rendering of the point. In particular, the type field may indicate whether the point comprises a glass point (or more generally the type field may indicate whether the point comprises a transparent point and / or a translucent point and / or a reflective point). As described in GB2412534.6, which is incorporated herein by reference, the rendering of such points may differ from the rendering of opaque points. In particular, in some embodiments, the points are sized so as to cause an overlap between adjacent points. This prevents the presence of any gaps between adjacent points, however where points are translucent this can cause such translucent points to be rendered incorrectly. In this regard, two identical overlapping opaque surfaces will render the same as a single opaque surfaces; however, two identical overlapping translucent surfaces will render differently to a single translucent surface. The glass flag may be used to identify that a surface is translucent, e.g. so that only a single surface of two overlapping surfaces is rendered if the glass flag is set for each of these surfaces. Similarly, the type flag may identify that a surface and / or a point is reflective so that a device that is rendering the point is able to render reflections based on the type flag. It will be appreciated that a translucent point may be identified based on the transparency field 1002. The type field may still provide a quick flag to indicate that a point is non-transparent. Furthermore the type field may indicate that a point is a specific type of non-transparent material (e.g. the type field may be a glass field that indicates whether a point is a point on a glass surface). A distance field 1012. The distance field indicates the distance of the point from the location of the capture device referenced in the capture device identifier field 1006. This distance may be indicated in absolute terms (e.g. as a number of metres). Typically, this distance is signalled as an index, where this index can then be related to a quantisation curve, in particular a 1 / z quantisation curve, or to a table of distances. This curve or table may be signalled elsewhere in the bitstream, e.g. in a header of the bitstream, and / or may be sent separately to the bitstream. Typically, the distance is signalled as an index relating to a quantisation level (of a curve with a plurality of quantisation levels). In this regard, the distances of points are typically quantised based on a curve, so that the distances of points close to the viewing zone 1 can be signalled more accurately than the distances of points further from the viewing zone. A normal field 1014. The normal field indicates a value of a normal of a surface defined by the point. The normal value is typically signalled in the coordinates of a space defined in a header of the bitstream. In some embodiments, the normal value is signalled with reference to a capture device (e.g. the normal field may indicate a difference between the normal of a point and the angle of the point from the capture device). The normal is typically stored as two components using octahedron normal encoding so that the normal field 1014 typically comprises two (or more) sub-fields, e.g. a normal x sub-field and a normal y sub-field. The normal field may comprise two 8 bit sub-fields to enable such encoding. Equally, the normal may be stored as an index that can be used to identify a normal from a table (e.g. from a table defined in a header of the bitstream), a curve, and / or a function. An angle field 1016. The angle field indicates an angle of the point from a capture device indicated by the capture device field 1006. Therefore, by combining the distance from the distance field 1012 and the angle from the angle field point with a location of the capture device identified in the capture device field, it is possible to identify the location of the point. The angle field 1016 typically comprises an index that identifies an angle based on an angular step. For example, a header of the bitstream may identify an angular step size or may contain a table that relates angular indexes to angular values so that a computer parsing the bitstream is able to identify an angular value from a signalled angular index (e.g. with a starting angle of 0° and an angular step of 10°, an angle field value of 7 would indicate an angle of 60°). This angular step and / or table may also be signalled separately to the bitstream. Typically, the angle field 1016 comprises a plurality of sub-fields. For example, the angle field may comprise one an X sub-field and a Y sub-field to indicate, respectively an azimuthal angle and an elevational angle. The angle field may comprise a 12 bit Y field and a 13 bit X field (or vice versa). Typically, the angle field 1016 identifies an angular bracket covered by a point. In this regard, each point typically covers an area in two-dimensional space (e.g. the point is not a one-dimensional point, but instead the point covers an area of an angular bracket). The angle values defined by the angular field may then identify a centre (or a leftmost, rightmost, topmost, and / or or bottommost point) of the angular bracket. Equally, the angle values may indicate an index of the angular bracket. A relative movement field 1018 (that may be referred to as a ‘reff field). The reff field may comprise a flag, e.g. a 1 -bit flag and / or an index. As described above, in some embodiments the viewing zone is arranged to move between successive three-dimensional representations (this may equally be considered as a plurality of points moving relative to the viewing zone between these representations). In some embodiments, the viewing zone may be associated with a viewing zone movement vector (and / or a ‘cockpit’ flag). The relative movement, or ‘reff, field 1018 indicates whether a point moves with the viewing zone. Therefore, the reff field may indicate whether a movement vector of a point is defined with reference to a viewing zone movement vector (or is defined in absolute terms). Where the viewing zone is associated with a movement vector, often there are many points in the scene that are stationary relative to the viewing zone (e.g. points that are within a ‘cockpit’). These points may have the reff field set to 1 (e.g. the reff flag set) and then have a movement vector flag set to 0 to indicate that these points are stationary relative to the viewing zone. Equally, these points within the cockpit may have the reff flag set and also have a motion vector flag set to indicate that these points are moving within the cockpit (where any movement of points within the cockpit is typically less than a movement of points outside of the cockpit). Similarly, a plurality of points that are outside of the cockpit may have a reff field 1018 value of 0 (e.g. the reff flag is not set) and a movement vector flag that is also not set, where these points may be moving substantially relative to the viewing zone. Therefore, movements for a plurality of points can be efficiently signalled via a movement vector of the viewing zone and via reff flags. Typically, the viewing zone movement vector is signalled in a header of the bitstream. The bitstream may indicate a plurality of viewing zones and / or a plurality of viewing zone movement vectors. In some embodiments, the viewing zone movement vector is signalled via a point of a three-dimensional representation. For example, the points may be arranged in the bitstream so that a first point comprises a viewing zone movement vector. Equally, the reff field 1018 may comprise a plurality of bits so that a point that indicates a viewing zone movement vector can be signalled by the reff field (e.g. by having a value of 2 in the reff field). In some embodiments, the relative movement field 1018 indicates a type of a movement of the point relative to the viewing zone. For example, the relative movement field may indicate that the point moves in the same way as the viewing zone and / or that the point moves in an opposite direction to the viewing zone. A motion vector (or movement vector) flag 1020. The movement vector flag identifies whether the point is associated with a movement vector. Typically, a value of 0 indicates that the point is not associated with a movement vector (e.g. that the point is stationary) and a value of 1 indicates that the point is associated with a movement vector (e.g. that the point is moving). A size field 1022. The size field indicates a size of the point. Typically, the size field 1022 indicates a number of angular brackets covered by the point. The size field may, for example, indicate a number of azimuthal (e.g. horizontal) brackets covered by a point and / or a number of elevational (e.g. vertical) brackets covered by a point. Typically, the size indicates a size from a list of possible sizes, which list of possible sizes typically contains sizes that are rectangles (and / or rows and / or squares) - such sizes can be defined by a width and a height. In some embodiments, the size indicates an absolute size of the field (e.g. a size in metres). In some embodiments, the size field 1022 comprises one or more sub-fields. For example, the size field may comprise one or more of: a width sub-field, a height sub-field, and an underscan subfield. Equally, the size field may comprise a single value that can be used to derive widths and heights of the point (e.g. by using this value as an index for a table of possible widths and heights). Where the size field comprises a height sub-field and / or a width sub-field, these sub-fields are typically arranged to provide available values of 1 - 3 (e.g. so that the maximum size of a point is 3x3 (where the point covers 9 angular brackets in a 3x3 arrangement). In some embodiments, the arrangement of brackets that is covered by a point and / or the direction of an area covered by the point is inferred from a value of the size field 1022. Equally, this information may be signalled separately. For example, the size may indicate a number of brackets covered by the point that are to the right of and below the bracket signalled by the angle field 1016. This arrangement may depend on a value of the size field. For example, a width of two may signal that the point covers two brackets including a signalled bracket and a bracket to the right of the signalled bracket and a width of three may signal that the point covers three brackets including a signalled bracket and two brackets to the right of the signalled bracket. Equally, a width of two may signal that the point covers two brackets including a signalled bracket and a bracket to the right of the signalled bracket and a width of three may signal that the point covers three brackets including a signalled bracket, a bracket to the left of the signalled bracket, and a bracket to the right of the signalled bracket. The arrangements covered by different sizes of points may be signalled elsewhere in the bitstream (e.g. in a header of the bitstream) and / or may be signalled separately to the bitstream. For example, the arrangements covered by the different sizes of points may be defined by a file type associated with the bitstream. The size value may be stored as an index, which index relates to a possible plurality of sizes and / or a list of sizes (e.g. if the size may be any of 1x1, 2x1, 1x2, 2x2, pixels this may be specified by using 3 bits and a list that relates each combination of bits to a size). In some embodiments, the size field 1022 and / or the value in the size field comprises an underscan value, which underscan value defines a reduction in the number of points captured as compared to a representation without underscan. The size of the points may be defined so as to indicate this underscan value. In an exemplary embodiment, the underscan value is an integer value between 0 and 3 and the size is stored as a combination of point dimensions (e.g. a width in the range [0,2]) and a height in the range ([0,2]) and an underscan factor (e.g. an underscan factor in the range [0,3]). In some embodiments, the width and the height are dependent on the underscan factor. For example, when the underscan factor exceeds a threshold value, the possible height and width values may be limited. In a specific example, when the underscan factor is 3, the width and the height may be limited to the range [0,1], The size may then be defined as size = underscan*9 + height*3 + width. Such a method provides efficient storage and indication of width, height, and underscan values. In some embodiments, the use of underscan may be identified based on the value in the size field exceeding a threshold value. In some embodiments, each point has a fixed size (e.g. 1x1) and so the size field 1022 can be excluded. In some embodiments, the bitstream comprises a plurality of different types of points (e.g. texture points and other points) where these points may have different sizes - where each type of point has a fixed size, these sizes may be inferred from a type of the points so that the size field may still be excluded. It will be appreciated that the fields described above are exemplary and that variations of these fields may be provided in order to define a point. For example, the location of the point is typically defined in relation to a capture device such that an absolute location of the point can be determined using an identifier of this capture device, a distance of the point from the capture device, and an angle of the point for this capture device. The location of each potential capture device, and an angular step associated with the capture devices, may be defined in a header of the bitstream. A device parsing the bitstream is then able to: identify a capture device associated with a point based on the capture device field 1006; determine a distance of the point from this capture device based on the distance field 1012; identify an angle of the point from the capture device by determining an angular bracket of the point based on a value in the angle field 1016; and then determine a location of the point by combining the location of the capture device; the angle of the point to the capture device; and the distance of the point to the capture device. However, it should be appreciated that other methods of signalling the location of the point are possible. For example, the point may comprise coordinates (e.g. x, y, z coordinates) that can be used to locate the point in a three-dimensional space. The capture device field, the distance field, and the angle field may then be replaced with coordinate fields that define the point in coordinates (e.g. cartesian or polar coordinates). Equally, the points may be arranged in the bitstream so that the bitstream contains a sequence of points in known locations (e.g. the bitstream may comprise a series of adjacent points in order moving from left to right and top to bottom); in such embodiments, the points may comprise no location data, since these locations can be inferred. In some embodiments, the points are arranged in the bitstream based on the capture device associated with those points. For example, the bitstream may first define points associated with a capture device with an index of ‘0’ and then points associated with a capture device with an index of ‘1 ’, etc.. The bitstream may then comprise a flag between each point or between each sets of points that indicates a change (e.g. an increment) in a capture device. For example, each point may be followed with a zero to indicate that the following point is associated with the same capture device or a one to indicate that the following point is associated with a different capture device (e.g. in order to increment a capture device index). Equally, further bits may be used to indicate a specific change in a capture device identifier. In some embodiments, a specific string of bits is used to signal a change in a capture device identifier so that it may be assumed that the capture device identifier for a current point is the same as for a preceding point unless this string of bits is present in the bitstream in between the points. One or more of the values defined in the point, e.g. the transparency value 1002, the attribute value 1004 (and / or sub-values of the attribute value), the further attribute value 1008 (and / or sub-values of the further attribute value, the distance value 1012, the normal value 1014, and / or the angle value 1016, may be quantised whereby one uniform more of these values may be mapped onto a set of discrete values. The quantisation may be a regular quantisation (e.g. where the step between the values in the set is constant) or a non-uniform quantisation (e.g. where the step between the values in the set changes). For example, the quantisation may be based on a quantisation curve. The number of levels in the set, e.g. a degree of quantisation, may be the same for each point. Equally this number of levels may change. For example, one or more values for points within a threshold distance of the viewing zone 1 may be quantised based on a first set of quantisation levels and one or more values for points beyond the threshold distance may be quantised based on a second set of quantisation levels, where the second set of levels may be smaller than the second set of levels. This enables points near to the viewing zone 1 to be defined in a greater level of detail than points further from the viewing zone. In particular, a value of a distance defined in the distance field 1012 may be quantised so that a step between quantisation levels increases with distance from the viewing zone 1. For example, the distance may be quantised based on a quantisation curve, such as a 1 / z curve. It will be appreciated that various methods of quantisation are possible, e.g. using a quantisation curve, using a table of possible quantisation values, and / or using a function. Therefore, these signalled values, e.g. the distance signalled in the distance field may be compared to a curve, a table, or a value, in order to extract an actual value from a signalled value. In a very simple example, the possible distances for a point may be either 0m, 50m, 100m, or 200m. These distances can then be signalled using a two-bit distance field that can contain a value of 00 (for 0m), 01 (for 50m), 10 (for 100m), or 11 (for 200m). In some embodiments, a computer device is arranged to divide the three-dimensional representation into a plurality of containers, or‘froxels’ (frustrum voxels), where the dimensions - in particular the depths and / or volumes - of each froxel may depend on a distance of that froxel to the viewing zone. Typically, the three-dimensional representation is divided into froxels using a coordinate system that is: based on a location of / in the viewing zone; based on a centrepoint of the viewing zone; and / or based on a capture device. A point may then be defined based on a froxel that encompasses that point. For example, each froxel may be associated with a set of quantisation levels, where a distance of a point can then be defined based on a froxel identifier and a quantisation level. A first froxel may cover a distance from 0m - 10m and a second froxel may cover a distance from 10m - 30m. An identifier of the first froxel and a second quantisation level may then indicate a distance of 12.5m. An identifier of the first froxel and a second quantisation level may then indicate a distance of 15m. In some embodiments, the distance field 1012 is defined with reference to a container (e.g. a froxel) and a quantisation level, where the quantisation level defines a position of the point in the container. Typically, each container is associated with the same quantisation levels and / or quantisation curve. Typically, a depth and / or volume of each container depends on the distance of the container from the viewing zone 1. The volumes / distances associated with each froxel may, for example, be defined in a header of the bitstream. A detailed implementation of the use of froxels is described in GB2405098.1, which is incorporated herein by reference. Movement vectors Referring to Figure 20b, there is shown an exemplary range of fields that can be used to signal a movement vector. Typically, each movement vector is associated with a point (or a plurality of points). As described above, each point may comprise a movement vector flag 1020 that indicates whether this point is associated with a movement vector. The movement vectors may then be associated with points based on these flags and / or an ordering of the bitstream. In some embodiments, points and the movement vectors associated with these points are signalled sequentially in the bitstream, e.g. if the movement vector flag 1020 of a point is set to T then a movement vector may be expected after this point. Equally, the points and movement vectors may be signalled in different parts or sections of the bitstream. In some embodiments, each movement vector is associated with an identifier, which identifier may be signalled in the movement vector flag field of a corresponding point (e.g. this ‘flag’ field may comprise an identifier and not simply a binary flag). In some embodiments, the movement vectors in the movement vector section may be arranged in order so that a first point in the point section of the bitstream that has a movement vector flag value of T is associated with the first movement vector in the movement vector section of the bitstream, a second point in the point section of the bitstream that has a movement vector flag value of T is associated with the second movement vector in the movement vector section of the bitstream, and so on. Each movement vector in the bitstream typically comprises one or more of: A direction field 1102 that indicates a direction of the movement vector. The direction may be defined in absolute coordinates. Equally, the direction may be defined in relative coordinates, for example the direction of movement may be defined relative to an angle of a point associated with the movement vector (with this angle of the point being defined in the angle field 1016 of that point). The direction field may comprise a plurality of sub-fields that indicate components of a direction (e.g. ‘theta’ and ‘phi’ fields, or'x’, ‘y’, and ‘z’ fields). A magnitude field 1104 that indicates a magnitude of the movement vector. In some embodiments, a movement vector comprises an 8-bitx-direction field relating to a direction of the movement vector relative to an x-axis, an 8-bit y-direction field relating to a direction of the movement vector relative to a y-axis, and a 16-bit magnitude field relating to a magnitude of the movement vector. In some embodiments, a movement vector comprises an 8-bit phi field and an 8-bit theta field that together provide a direction of a movement vector as well as an 8 bit radius field that indicates a magnitude of the movement vector. Typically, each of the direction field 1102 and the magnitude field 1104 defines a single number (which number may be derived from a plurality of sub-fields). In some embodiments, one or more of the direction field and the magnitude field defines a plurality of numbers, and / or a function by which a number varies. In particular, these fields may define values that change with time and / or a function of the direction / magnitude of the movement vector that depends on a time variable. Such an implementation can be used to define movement of a point that is accelerating or decelerating in the period between successive three-dimensional representations. In some embodiments, one or more parts of the movement vector information is incorporated into the point date in the bitstream. That is, the movement vector fields may be included in a point in the bitstream (e.g. in place of, or in addition to, the movement vector flag 1020). An example of this is shown in Figure 20a in which a phi component of a movement vector for a point is incorporated into a definition of a point. The movement vectors enable a device parsing the bitstream to determine locations of points in between signalled three-dimensional representations. Typically, this comprises providing a movement vector that indicates a movement of a point within a time period associated with a three-dimensional representation (e.g. the location of the point is signalled at the start of the time period and then the movement vector indicates the way in which the point moves during the time period). In some embodiments, movement during a time period associated with a three-dimensional representation is determined by interpolating between a plurality of (e.g. successive) three-dimensional representations or extrapolating based on a plurality of three-dimensional representations. This enables intermediate frames or intermediate three-dimensional representations without the need to signal movement vectors so that the movement vector flags and fields may be excluded. However, the use of interpolation typically prevents the use of complex movements (e.g. this typically precludes the use of points that accelerate over the time period associated with a three-dimensional representation). Furthermore, in order to enable such interpolation the computer device must be able to identify related points in the plurality of three-dimensional representations. In some embodiments, the points comprise identifiers in order to enable such identification of related points. In some embodiments, the movement vector flag 1020 of a point indicates whether an interpolation is used to determine a movement vector and / or the fields of the movement vector indicate a difference of a location of a point from an interpolated location. For example, the computer device parsing a movement vector may be arranged to interpolate a location of a point based on the locations of this point in successive three-dimensional representations and then to combine this interpolated location with a value signalled in a movement vector. In some embodiments, the bitstream is associated with one or more possible movement vectors (e.g. there are only a limited number of available movement vectors of specified directions and / or sizes). These possible movement vectors may, for example, be signalled in a header of the bitstream. The movement vector component may then comprise a single index that indicates a movement vector from these possible movement vectors. In such embodiments, the movement vector flag 1020 of a point may indicate this single index. Such implementations may be suitable for situations where the movement of objects is constrained (e.g. where the three-dimensional representations show objects that only move at right angles). Texture points Referring to Figure 20c, there is shown an exemplary arrangement of fields that can be used to signal a texture point. This texture point references (via an index) a texture patch in a texture atlas, where this texture patch comprises attribute data for the texture point. Typically, the texture patches are signalled separately to the texture points. In particular, the texture patches may be collated into a texture atlas (that contains a plurality of texture patches) and this texture atlas may then be transmitted in a section of the bitstream. The texture atlas may relate to a single three-dimensional representation or equally the texture atlas may relate to a plurality of three-dimensional representations. Typically, each texture atlas is signalled as a two-dimensional image (with each texture patch being a patch of this two-dimensional image. More generally, the texture points comprise points that reference a separate source, where a computer device is able to identify one or more attributes of the texture points based on this separate source. While the separate source is typically a texture patch, this separate source may be another source. For example, the separate source could be a video (or another complex form of media), where texture points are then useable to include videos in the three-dimensional representation. As described above, the three-dimensional representation may comprise one or more texture points which reference texture patches (which texture patches contain a plurality of attribute values, e.g. an 8x8 square of pixel values). Typically, texture points comprise a texture patch identifier in the attribute field 1004 and / or the further attribute field 1008. In some embodiments, a flag is used to identify that a point is a texture point (e.g. each point may comprise a texture patch flag). In some embodiments, the presence of a certain sequence of bits in the attribute field or the further attribute field may be used to identify that a point is a texture point. For example, a certain value in the attribute field may be used to identify that the point is a texture point and that the further attribute field contains a reference to a texture patch in a texture atlas. In some embodiments, texture points are identified by a transparency value of the texture points. In particular, the bitstream may be divided into different segments of points, where a first segment contain opaque points and a second segment contains non-opaque points. Texture points may then be located in the second segment with these texture points being identified by having an alpha field 1002 that indicates the point is opaque (e.g. that has a value of ‘255’). In some embodiments, the texture points are different (in form or size) to other points. These texture points may then be signalled separately to the other points. However, typically, the texture points have the same size (e.g. the same number of bits) as other points. Therefore, the texture points can be included alongside the other points. A computer device parsing the bitstream is then able to identify that a point is a texture point (e.g. based on a texture point flag, based on a transparency value of the point, and / or based on an identification of a reference in the texture point to a texture patch) and to parse this point separately to other points. In some embodiments, different capture devices are used to signal texture points and other points so that a texture point can be identified from a capture device field 1006 of a point (where this capture device field may then be located at the same place in texture points and other points). Typically, each texture point comprises one or more of: A transparency or ‘alpha’ field 1202. As with the alpha field 1002 of ‘typical’ points, the alpha field of a texture point typically indicates a transparency of the texture point. The transparency value of the texture point may be a transparency value forthe entirety of a texture patch associated with the texture point or may be a multi-dimensional transparency value that enables different parts of the texture point to have different transparencies. A texture index field 1204. The texture index field contains a reference to a texture patch. Typically, this reference comprises an index value that enables a texture patch to be determined from a texture atlas comprising a plurality of texture patches. Typically, the texture atlas is defined separately in the bitstream such that the relevant texture atlas can be identified from the location of the texture point in the bitstream (e.g. each three-dimensional representation may be associated with a corresponding texture atlas, which texture atlas may be signalled in a header of the bitstream). In some embodiments, the texture index field comprises an identifier of a texture atlas as well as an identifier of a texture patch. The texture index field may, more generally, comprise an identifier of an outside source, where the attribute values forthe texture point are obtained from this outside source. For example, the texture index field may comprise a hyperlink. Typically, the texture atlas is arranged so that texture patches for a left eye and a right eye of a point are stored adjacently. Therefore, the texture index field 1204 typically references a single index for a single texture patch for a single eye and a computer device parsing the texture point is able to identify a further texture patch for another eye based on the single index. For example, where the texture index has a value of n, a first texture patch for a left eye may be extracted from an nth index of a texture atlas and a second texture patch for a right eye may be extracted from an (n+1)th index. Equally, the texture index field may comprise a plurality of indexes, with each index indicating a texture patch of that texture atlas. In this way, the texture point may reference a plurality of texture patches (e.g. a texture patch for each eye). In some embodiments, the same texture patch may be used for each eye for a texture point. In some embodiments, the bitstream comprises a plurality of texture points that each referto the same texture patch and / or same plurality of texture patches, where this can enable efficient encoding of a scene with repeated motifs. In these embodiments, the texture patches may be signalled in the texture atlas based on a number of points that reference each texture patch (e.g. so that texture patches that are referenced by multiple points are located at the start of a texture atlas section of the bitstream). The texture index field 1204 may, for example, comprise a 2D tile coordinate on a texture atlas (e.g. a top-left coordinate of the texture patch). Typically, the texture index field comprises an index value, where the computer device parsing the texture point is able to identify a coordinate value in the texture atlas based on this texture point due to the fixed size of the texture patches. For example, where the texture patches are each of a size of 8x8, a texture index of (1,1) may signal a texture patch with a top-left corner at pixel (1,1) of a texture atlas, a texture index of (2,1) may signal a texture patch with a top-left corner at pixel (9,1) of a texture atlas, and so on. Typically, the texture index is specified as a row-major tile index into the texture atlas. The texture point index may be considered to comprise an attribute field that is comparable to the attribute field 1004 of a ‘conventional’ point. In other words, each point may be considered to comprise an attribute field where the attribute field for a ‘conventional’ point may comprise a colour value and the attribute field for a texture point may comprise a texture index. A stereo flag 1206. The stereo flag indicates whether the same texture patch should be used for each attribute (e.g. each eye) of a texture point. Where the stereo flag is ‘O’, the same texture patch (that is referenced by the texture index in the texture index field) may be used for each eye. Where the stereo flag is T, different texture patches may be used for each eye. For example a first texture patch for a left eye may be extracted from an nth index of a texture atlas and a second texture patch for a right eye may be extracted from an (n+1)th index of that texture atlas. Equally, a first texture patch for a left eye may be extracted from an nth index of a first (e.g. left eye) texture atlas and a second texture patch for a right eye may be extracted from an nth index of a second (e.g. right eye) texture atlas. More generally, in some embodiments there is a known (e.g. preset) relationship between the texture patches for a left eye and a right eye of a user. For example, these texture patches may be adjacent to each other. This enables the texture patches for each eye to be signalled with a single index. Such embodiments may be implemented with or without the stereo flag 1204. In some embodiments, the stereo flag 1204 comprises value that identifies a relationship between the (indexes and / or locations of the) texture patches fora left eye and a right eye of a user- and a plurality of relationship options may be possible. For example, the stereo flag may indicate that the relationship is one of: the texture patches for the left and right eyes being the same; the index of the texture patch for a second eye being one greater than the index of a texture patch for a first eye; the texture patch for a second eye being contained in a different texture atlas to the texture patch for a first eye; the texture patch for a second eye being defined as a difference to the texture patch for a first eye. Such embodiments may require the stereo ‘flag’ to comprise a plurality of bits. A capture device field 1208, a distance field 1218, and an angle field 1222. As with the capture device field 1006, the distance field 1012, and the angle field 1016 for other (‘typical’), these fields identify a capture device associated with the texture point, identify a distance of the texture point from that capture device, and identify an angle of the texture point from that capture device. These fields thereby enable a location of the texture point to be determined. A movement vector flag and / or a movement vector field 1210. Typically, the texture points are associated (or can be associated) with a plurality of movement vectors. For example, the texture points may be associated with up to four movement vectors so that each corner of the texture point is associated with a separate movement vector. In some embodiments, the texture points are associated with a plurality of different movement vectors that are signalled elsewhere in the bitstream. In some embodiments, one or more of (and typically a plurality of) movement vectors are signalled directly in the texture point via the movement vector field 1210. In this regard, since the texture point does not need to signal attribute data, the texture point can use a number of bits for the signalling of a movement vector. The movement vector field 1210 may comprise a plurality of sub-fields, e.g. direction sub-fields and magnitude sub-fields. Where the texture point is associated with a plurality of movement vectors, each movement vector may be signal a corner (or a part) of the texture point that is acted upon by the movement vector. Equally, the movement vector information may be arranged in a known sequence so that the corner (or part) can be determined from the location of a movement vector information in the texture point data (e.g. a first signalled movement vector may be applied to a top-left corner, a second signalled movement vector may be applied to a top-right vector, a third signalled movement vector may be applied to a bottom-left vector, and a fourth signalled movement vector may be applied to a bottom-right vector). In some embodiments, the movement vector flag and / or movement vector field 1210 may signal that the same movement vector is used for each corner of the texture point. For example, a value of 0x1 FFFFF may signal that the same movement vector is used for each field. While the example of Figure 20c shows a single contiguous movement vector field 1210, in some embodiments the movement vector field is distributed within the texture point. More specifically, the texture point may comprise a plurality of movement vector sub-fields that can be combined by the computer parsing the texture point in order to identify one or more movement vectors for the texture point. A diagonal orientation (or bending) field 1212. The diagonal orientation field indicates a diagonal of the texture point about which the texture point is able to bend. The diagonal orientation field typically comprises a flag that indicates whether a texture quad has a diagonal orientation of bottom-left to topright or top-left to-bottom right. More generally, the diagonal orientation field may indicate a line of bending or flexure of the texture point. Typically, the texture point has a default bending line - e.g. a bottom-right to top-left diagonal of the texture point - where the diagonal orientation field 1212 indicates whether or not this bending line is used. Typically, there are only two possible bending lines for the texture point so that the flag indicates which of these bending lines should be used for the texture point. A bend field 1214. The bend field indicates a bending of the texture point, which bending occurs about the bending line indicated by the diagonal orientation field. The bend field typically indicates a bend value that is an 8-bit signed integer. During bending, the corners of the texture point that are on the bending line typically remain stationary. The corners not on the bending line are taken from a hypothetical texture patch with the same normal as the texture patch but with a different distance from the capture point associated with the texture point. This new distance may be calculated as new distance = original distance * (1 + bend value * 16 / cubemap resolution). The cubemap resolution is typically signalled in a header of the bitstream. A normal field 1220. As with the normal field 1012 for other (‘typical’) points, this normal field indicates a value of a normal of a surface defined by the texture point. A reff field 1224. As described above, the reff field indicates whether the texture point moves with a viewing zone and / or whether a movement vector of the texture point is defined relative to a movement vector of a viewing zone. A motion vector (or movement vector) flag 1220. The movement vector flag identifies whether the texture point is associated with a movement vector. The movement vector flag may also indicate a number of movement vectors associated with the texture point. An underscan field 1228. Typically, the texture point has a fixed size (e.g. the texture point may have a fixed size of 8x8). Therefore, the width and height of the texture point do not need to be signalled. The texture point may still comprise an underscan field that Indicates whether the texture point should be scaled (e.g. an underscan value of 2 may indicate that the width and height of the texture point should be doubled). Equally, the texture point may have a various size and so may still comprise size fields or comprise any of a height, a width, or an underscan field. It will be appreciated that any features described above with reference to the points can equally be applied to the texture points. For example, the locations of the texture points may be signalled either with reference to capture devices or using absolute coordinates. It will be appreciated that many of the above-described fields are optional and can be excluded from implementations of texture points. For example, the bend field may be excluded to provide texture points that do not bend about a bending line. Deformation of the texture points can still be provided via movement vectors. Equally, the texture points may be rigid points in which case each of the diagonal orientation field 1212, the bend field 1214, orthe movement vector 1210 may be excluded. While the texture point structure has been described separately to the point structure, it will be appreciated that a texture point is a type of a point and that typically the texture points are transmitted alongside other points (where the computer device then identifies that a point is a texture point based on, for example, a value ofthe transparency field of the point). Texture atlas One or more texture atlases may be signalled in the bitstream, e.g. in a header ofthe bitstream and / or alongside one or more points ofthe bitstream. In some embodiments, a plurality of texture atlases (e.g. relating to a plurality of different three-dimensional representations) are defined in a first section of the bitstream so that they can later be referenced by one or more texture points in a second section of the bitstream. Each texture atlas may relate to a single three-dimensional representation. Equally, one or more texture atlases may relate to a plurality of three-dimensional representations. In some embodiments, a second texture atlas is defined based on a first texture atlas where, in particular, one or more texture patches in the second texture atlas may be defined as a difference to one or more texture patches in the first texture atlas. The texture atlas is typically defined as a two-dimensional (2D) image, where each texture patch in the texture atlas is formed by a subset of this two-dimensional image. A texture patch can then be extracted from the texture atlas based on a reference to a pixel or an area ofthe texture atlas image. Equally, the texture atlas may be defined as a database where each texture patch is associated with a different index of the database. Typically, the texture patches are of a fixed size, so that a texture patch can be identified within a 2D image that forms a texture atlas based on a provided index value and the fixed size of the texture patches. In some embodiments, the definition of a texture atlas comprises: a definition of a first ‘representative’ texture patch (e.g. a first two-dimensional image); and a definition of one or more second texture patches based on this firsttexture patch. More specifically, the second texture patches may be signalled by including ‘deltas’ in the bitstream, where a set of deltas for a specific second texture patch indicates a difference between the firsttexture patches and this specific second texture patch. In some embodiments, the deltas comprise a two-dimensional image that indicates the differences between the texture patches. In some embodiments, the deltas comprise one or more numbers that indicate a difference between the texture patches (e.g. a vector difference of the texture patches when converted into an n-dimensional space). In some embodiments, a plurality of texture atlases are signalled in the form of a video, where this video can then be encoded using video-encoding techniques such as MP4, HEVC, VVC, or LCEVC codecs. In some embodiments, the encoding and / or decoding of the texture atlases are performed by a hardware encoderand / or decoder and / or are performed by an encoder / decoderthat are configured to encode / decode videos. In some embodiments, the bitstream is arranged to enable separate and / or independent encoding and decoding of the texture atlases (and of other parts of the bitstream). Therefore, for example, a first decoding device may decode point data while a second decoding device simultaneously decodes a texture atlas. By encoding / decoding the texture atlases using an encoder / decoder that is configured for video encoding / decoding, other computer componentry can be freed up for other (different) encoding / decoding tasks. In some embodiments, the computer device is arranged to reorder or rearrange one or more texture atlases and / or one or more texture patches within one or more texture atlases in order to enable more efficient encoding of the texture atlases. For example, the computer device may rearrange a texture atlas to reduce differences between points within this texture atlas and / or to reduce differences between a plurality of texture atlases. In some embodiments, the computer device is arranged to rearrange a second texture atlas, e.g. to update the indexes of texture patches within this second texture atlas based on a first texture atlas. This may comprise rearranging the second texture atlas so as to be similar to the first texture atlas to enable more efficient encoding of a plurality of texture atlases. The computer device may also be arranged to update the values of the texture index 1204 fields within one or more points associated with the (rearranged, e.g. second) texture atlas. In particular, the computer device may be arranged to update the index of one or more texture patches within a texture atlas and then to update the texture index fields of one or more points associated with these texture patches (so as to maintain a correct association between the texture points and the texture patches). Other points The description above has described texture points as being a different type of point to other points. It will be appreciated that other types of points (e.g. with different structures) may also be used. For example, as described in GB2410602.3, which is incorporated herein by reference, a point may be defined that is associated with a two-dimensional image and / or a video. Such a point may be formatted as a texture point and / or as another point. For example, the bitstream may comprise a video point that identifies a video that can be used to render one or more two-dimensional images. Various types of points may be indicated using a flag or a field or using an ordering of points in the bitstream. For example, each point may comprise a type field that indicates a type of that point. Bitstream ordering In some embodiments, the bitstream is separated into a plurality of different parts, portions, sections, or segments, where each part of the bitstream defines a different component of the three-dimensional representation. For example, the bitstream may comprise one or more of the following parts: A header part comprising information about the three-dimensional representations such as locations of capture devices and information (e.g. flags) indicating features of the components encoded in the bitstream. For example, the header may identify a number of points in the three-dimensional representation and / or the bitstream and / or the header part may define one or more viewing zones associated with the three-dimensional representation. A texture atlas part comprising one or more texture atlases. A point part comprising one or more points and / or one or more texture points. A movement vector part comprising one or more movement vectors. Equally, these components may be interspersed in the bitstream. For example, each movement vector (or set of movement vectors for a texture patch) may follow the point to which the movement vector(s) relates. Typically, the bitstream signals a plurality of three-dimensional representations. In some embodiments, each three-dimensional representation is signalled separately (e.g. with a first section of the bitstream containing a texture atlas, one or more points, and one or more movement vectors for a first three-dimensional representation and a second section of the bitstream containing a texture atlas, one or more points, and one or more movement vectors for a second three-dimensional representation). Equally, the components of a plurality of three-dimensional representations may be interspersed so that, for example, texture atlases for different three-dimensional representations are located adjacent in the bitstream. Such an arrangement can enable the efficient encoding of these components (in particular, such an arrangement enables a plurality of texture atlases to be encoded as a video). In some embodiments, the components are signalled in a plurality of sections so that: A first section signals one or more opaque points and / or texture points; and A second section signals one or more non-opaque (e.g. translucent or transparent) points and / or texture points. With such an arrangement, texture points may be identified (and distinguished from other points) by interrogating a transparency value of the signalled points. In particular, a point in the second section that has an alpha value indicating that this point is opaque (e.g. a value of 255) may be identified as a texture point. Similarly, the components may be signalled using a first section that signals one or more non-transparent (e.g. translucent or opaque) points and a second section that signals one or more transparent points. Texture points may then be identified in the first section based on a transparency value of that indicates the point is transparent (e.g. by having an alpha value of 0). In general, the present disclosure envisages a method of identifying a texture point (and / or distinguishing a texture point from another point) based on a transparency value of a signalled point. In embodiments where points are separated into a first section of non-opaque points and a second section of opaque points, this separation may enable the computer device parsing the bitstream to more efficiently render the three-dimensional representation. In particular, the computer device may be able to first render points that can be seen through and then to render points that cannot be seen through (in this regard, the rendering of a non-opaque point in a particular location in space requires the computer device to still consider points behind this location; in contrast, the rendering of an opaque point at this location enables the computer device to ignore any points behind this location). By rendering the non-opaque points first, it can be ensured that the rendered points are thereafter modified based on the opaque points behind the non-opaque points. In some situations, this can be more efficient than rendering the opaque points first and then needing to modify these opaque points based on nearer non-opaque points. Equally, the computer device may be able to render opaque points first so that any other (e.g. non-opaque) points located behind already rendered opaque points can be ignored. In some embodiments, points (and / or texture points) are arranged in the bitstream based on a capture device associated with the points. For example, the bitstream may first define points associated with a capture device with an index of ‘0’ and then points associated with a capture device with an index of ‘1 ’, etc.. The bitstream may then comprise a flag between each point or between each sets of points that indicates a change (e.g. an increment) in a capture device. For example, each point may be followed with a zero to indicate that the following point is associated with the same capture device or a one to indicate that the following point is associated with a different capture device (e.g. in order to increment a capture device index). Equally, further bits may be used to indicate a specific change in a capture device identifier. In some embodiments, a specific string of bits is used to signal a change in a capture device identifier so that it may be assumed that the capture device identifier for a current point is the same as for a preceding point unless this string of bits is present in the bitstream in between the points. This enables the signalling of points without requiring the points to explicitly indicate capture devices associated with the points. Typically, texture points are a subset of points - that is, texture points and other points typically have the same structure and / or the same size (e.g. all points, including texture points, may have a size of 16 bytes), with a parsing computer device being able to distinguish between texture points and other points based on, for example, a transparency value of a point (as described above). Equally, each point (or each texture point) may comprise a flag to indicate whether it is a texture point and / or whether it is not a texture point. In some embodiments, the texture points and the other points are defined in different parts of the bitstream so that a computer device parsing the bitstream is able to distinguish between texture points and other points. In some embodiments, texture points and other points are different in structure and / or size. Typically, points (both texture points and other points) have a size of 16 bytes and / or movement vectors have a size of 4 bytes. In some embodiments, movement vectors have a size of 2 bytes. It will be appreciated that in various embodiments, different sizes and / or arrangements of each component (e.g. points, texture points, movement vectors, texture atlases etc.) may be used. In some embodiments, the scene signalled by the bitstream is associated with a plurality of viewing zones. Typically, each three-dimensional representation signalled in the bitstream comprises one or more viewing zones, where these viewing zones may be the same viewing zones or may be different viewing zones. In some embodiments, one or more viewing zones are arranged to move, where these viewing zones may be associated with movement vectors that indicate a movement of the viewing zone within a time period associated with a three-dimensional representation. The bitstream may comprise information about one or more viewing zones, for example a location and / or a movement vector of one or more viewing zones. In some embodiments, each three-dimensional representation is associated with one or more viewing zones, where the viewing zones for a three-dimensional representation may be defined in a first portion of the bitstream and the points of the three-dimensional representation may be defined in a second portion of the bitstream. In some embodiments, the bitstream comprises a viewing zone part (e.g. within the header part) that defines one or more viewing zones. The viewing zone part may, for example, define locations and / or sizes of one or more viewing zones and / or the viewing zone part may identify an arrangement of capture devices associated with each viewing zone. In some embodiments, the viewing zone part of the bitstream may be associated with a plurality of three-dimensional representations, where one or more of the viewing zones persist throughout this plurality of representations. For example, a viewing zone may be defined a single time and may then be used for each of a plurality of three-dimensional representations. Where a three-dimensional representation is associated with a plurality of viewing zones, the bitstream may be ordered so as to associate each point with a related viewing zone. For example, the bitstream may be separated into a plurality of segments, where each segment of the bitstream comprises one or more points and where each segment of the bitstream is associated with a different viewing zone. This viewing zone may be signalled at the start of the segment. In some embodiments, the bitstream comprises a flag that indicates a change in viewing zone (e.g. the start of a new segment), where this flag may comprise a specific stream of bits. Equally, a number of points and / or movement vectors in each segment may be signalled in a header of that segment so that the end of a segment can be identified based on this signalled number of points. In some embodiments, one or more (or each) point (e.g. the points and the texture points) comprises a viewing zone identifier field that identifies a viewing zone associated with that point. For example, the bitstream may indicate a list of viewing zones and the viewing zone identifier field may indicate an index that identifies a viewing zone associated with a point. The viewing zone identifier field may be associated with (e.g. a sub-field of) the capture device identifier field, where this field then indicates a viewing zone associated with a point and a capture device of that viewing zone that is associated with the point. In some embodiments, that capture device identifier may identify a viewing zone associated with the point, where each viewing zone may be associated with a different set of capture devices. For example, capture devices 1 - 20 may be associated with a first viewing zone and capture device 21 - 40 may be associated with a second viewing zone. Flags and sicinallinci The bitstream, e.g. a header of the bitstream, may contain flags that indicate capabilities of the bitstream and / or that indicate options that are used by the bitstream. The bitstream, e.g. a header of the bitstream, may also signal information that can be used to parse the components defined in the bitstream. For example, the bitstream and / or the header may comprise: Information about one or more capture devices, for example an indication of one or more locations of capture devices. For example, the indication may define coordinates of the capture devices or the indication may indicate a configurations of capture devices that are used for the representation. On this latter point, the three-dimensional representations may be associated with plurality of possible configurations of capture devices where a selected configuration (e.g. that are defined by a file type) is signalled in the bitstream. The information about the capture devices may comprise a table of capture devices so that an index in a capture device field 1006, 1208 of a point can be associated with a specific capture device. The information about the capture devices may indicate a capture device step, where the capture devices may be arranged in a regular pattern aboutthe viewing zone 1 so that locations of each capture device can be determined from a location of a first capture device and a location of this capture device step. The information about capture devices may indicate relative priorities of the capture devices and / or may indicate a number of points associated with each capture device. The points may be arranged so that more points are associated with higher priority capture devices (e.g. the three-dimensional representation may be processed so that points associated with low priority capture devices are redefined so as to be associated with higher priority capture devices, where this may comprise altering a location, angle, and capture device index ofthese points). The signalling of the capture devices may then comprise a method of entropy encoding, e.g. Huffman coding, where the higher priority capture devices are signalled using an index with fewer bits than an index for lower priority capture devices. Where the number of points associated with each capture device is signalled in the bitstream, this may enable a capture device for each point to be inferred. More specifically, the points may be arranged in the bitstream based on the capture devices used to capture the points - e.g. so that the bitstream first defines a first plurality of points associated with a first capture device, then defines a second plurality of points associated with a second capture device, etc. The computer device parsing the bitstream may be able to determine a switch from the first plurality of points to the second plurality of points based on the number of points captured by each capture device (this number and / or the order of the capture devices, e.g. the order of the pluralities of points in the bitstream, being signalled in the bitstream). The capture device information may comprise information that links two capture devices so that this information enables an identity of a second capture to be determined given an identity of a first capture device and a value in the capture device field 1006, 1208 of a point. Typically, adjacent points in the three-dimensional representation are associated with either the same capture device or by adjacent capture devices. Therefore, a capture device for a second point can be efficiently signalled (e.g. by entropy encoding) as a difference since an expected difference is known. Furthermore, such embodiments may be combined with a priority list, where the expected capture device that is used to capture the second point may be one or more of: a capture device used to capture the first point; a capture device adjacent the capture device used to capture the first point; and a high priority capture device. Information about an angular step and / or an arrangement of angular brackets associated with one or more capture devices. Typically, the angle fields 1016, 1222 of the points contain one or more index values that can be related to an angular bracket associated with a capture device. The (e.g. header of the) bitstream may comprise an angular step, e.g. an azimuthal step and / or an elevational step, so that the index values can be multiplied by these steps to identify an angle of a signalled bracket. The capture devices may be associated with irregular steps to provide increased resolution in certain parts of the three-dimensional representations where the information may comprise a table that relates indexes to a particular angle or angular brackets or the information may contain a function and / or a table that enables an index to be associated with a particular angle or angular bracket (e.g. the information may indicate a plurality of angular steps as well as the angular ranges within which each angular step is used. Typically, each capture device is associated with the same arrangement of angular brackets. Equally, different capture devices may be associated with different arrangements of angular brackets, e.g. so that different capture devices provide different resolutions or so that different capture devices provide different resolutions in different parts of the three-dimensional representation (e.g. where a first capture device provides a high resolution in a first direction and a low resolution in a second direction and a second capture device provides a low resolution in the first direction and a high resolution in the second direction. Information about an origin point and / or a grid arrangement of a coordinate system used by one or more three-dimensional representations. Information about possible sizes of points and / or information that relates sizes of points to the coverage of these points. In particular, the (e.g. header of the) bitstream may comprise information, e.g. a table, that indicates the coverage that should be used for different sizes of points. Where a 2x2 point is signalled with reference to a particular angular bracket, the 2x2 point may cover: the signalled bracket and brackets to the right of and below this bracket; the signalled bracket and brackets to the right of and above this bracket; the signalled bracket and brackets to the left of and below this bracket; the signalled bracket and brackets to the left of and above this bracket. This coverage can be defined in the (e.g. header of the) bitstream. In some embodiments, default coverages are defined, e.g. by a file type, and the bitstream may then signal any variation from these default coverages. Typically, the size information identifies a table of possible sizes so that an indexes contained in the size field 1022 of a point can be relates to an actual size and / or coverage of the point. Information about one or more external sources referenced by points of the bitstream. In this regard, one or more points (e.g. the texture points) of the bitstream may reference an external source that defines attributes for these points. For example, the points may reference a database that contains images, where these images can be used to determine attribute values for the points. The (e.g. header of the) bitstream may indicate a database, e.g. by containing a hyperlink to a database, that enables a computer device parsing the bitstream to identify this external source and to retrieve information to this external source. Information about the timing of one or more three-dimensional representations and / or a time step associated with these three-dimensional representations. This may enable a computer device parsing the bitstream to determine or render one or more intermediate representations based on the points and the movement vectors signalled in the bitstream. The three-dimensional representations may be regularly spaced or the representations may be irregularly spaced, e.g. to provide more accuracy during an important time period. Information about a resolution, quality, or frequency of one or more three-dimensional representations; information about a suitable rendering process; and / or information about a device that is suitable for viewing a scene associated with the three-dimensional representations. This information may enable a computer device evaluating the bitstream to render suitable two-dimensional images based on the three-dimensional representations so as to present the scene to a viewer in a manner desired by the party that initially formed the bitstream. For example, this information may indicate a number of two-dimensional images that should be rendered to adequately represent the scene associated with the three-dimensional representations. Information about any processing that has been applied to the three-dimensional representations. For example, the information may indicate that one or more points of the three-dimensional representation have been processed so as to be associated with a capture device that was not used to capture the points. Or the information may indicate that one or more new points have been added to a three-dimensional representation, e.g. to cover a hole in the three-dimensional representation. This information may be used by the computer device parsing the bitstream to determine whether certain inferences can be made about the points in the representation. Information about a configuration of movement vectors. For example, information about a number of movement vectors used for texture points. Furthermore, the bitstream may indicate whether the viewing zone 1 is associated with one or more movement vectors. Additionally or alternatively, the bitstream may comprise one or more flags that indicate: whether movement vectors are present in the three-dimensional representation; whether a viewing zone movement vector is present in the three-dimensional representation; whether (for each point), a point is arranged to move with the viewing zone (e.g. cockpit flags); whether (for each point) that point should be affected by a viewing zone movement and, e.g. to what extent (e.g. there may be a scaling factor applied so that the viewing zone movement vector is scaled before being combined with a point, and this scaling factor may be different for different points); an accuracy of the movement vectors and / or a quantisation curve associated with the movement vectors; a time spacing of three-dimensional representations in the bitstream (e.g. so that a time period for each three-dimensional representation can be determined); and a unit type of the movement vectors signalled in the bitstream. Information about one or more viewing zones associated with one or more three-dimensional representation and / or a scene associated with the three-dimensional representations. For example, the viewing zone information may indicate locations, dimensions, and / or shapes of viewing zones where this information may be used to render a scene defined by the representations. The information may indicate a change in viewing zone locations and / or dimensions between three-dimensional representations. For example, the information may indicate a movement vector that is associated with a viewing zone so that a location of the viewing zone changes between successive three-dimensional representations. Other information, such as locations of capture devices, may be defined with reference to the viewing zones. For example, capture devices may be located in a regular arrangement about the perimeter of each viewing zone and / or in the centres of the sides of the viewing zone and / or in the centre of the viewing zone. Therefore, the locations of the capture devices may be determined based on the viewing zone information and in particular the viewing zone dimensions. Example components Referring to Figures 21a - 21 e, various exemplary components are shown. It should be appreciated that the various exemplary components and data fields described below may be used in any combination and that the data fields described below with reference to any specific figure may be used with or without other data fields described with reference to that figure. For example, while Figure 21a describes a transparency field that is used in combination with a number of other fields, it will be appreciated that this transparency field could also be used in other contexts (e.g. with a different arrangement of other data fields than is provided by Figure 21a). Referring to Figure 21a, there is shown an embodiment of a point. In this embodiment, the point comprises: A transparency (alpha) field 1302 that indicates a transparency of the point. This transparency field may have a size of 8 bits. A hdr24.Y field 1304, a hdr24.y field 1306, and a hdr24.x field 1308 that each define a separate component of (the colour of) an attribute of the point (e.g. an attribute for a left eye). These fields may have sizes of, respectively, 10 bits, 7 bits, and 7 bits. A capture device field 1310 that indicates a capture device associated with the point. This capture device field may have a size of 8 bits. A xor_hdr24.Y field 1312, a xor_hdr24.y field 1314, and a xor_hdr24.x field 1316 that each define a separate component of (the colour of) a further attribute (e.g. an attribute for a right eye). These fields typically define the further attribute as the XOR of the value of this further attribute and the value of the attribute defined by the hdr24.Y field 1304, the hdr24.y field 1306, and the hdr24.x field 1308. These xor_hdr24.Y, xor_hdr24.y, and xor_hdr24.x fields may have sizes of, respectively, 10 bits, 7 bits, and 7 bits. A type field and / or a glass field 1318 that Indicates a type of the point and / or indicates whether the point is a point on a translucent and / or reflective surface. The glass flag field may have a size of 1 bit. A distance field 1320 that indicates a distance of the point from the capture device identified in the capture device field 1310. The distance field may have a size of 15 bits. A normal x field 1322 and a normal y field 1324 that together indicate a normal of a surface on which the point lies using, e.g., octahedron normal encoding. The normal x and the normal y fields may each have a size of 8 bits. A Y field 1326 and an X field 1328 that together define an angle of the point from the capture device identified in the capture device field 1310. Typically, these fields define an angular bracket covered by the point. The Y field and the X field may have, respectively, sizes of 12 bits and 13 bits. A reff field 1330 that Indicates whether the point moves with a viewing zone of the three-dimensional representation. The reff field may have a size of 1 bit. A motion vector (or movement vector) flag field 1332 that indicates whether the point is associated with a movement vector. The movement vector flag field may have a size of 1 bit. A size field 1334 that indicates a size of the point. The size field may comprise an index that identifies one or more of: a height of the point, a width of the point, and / or an underscan or overscan value of the point. The size field may have a size of 5 bits. Referring to Figure 21b, there is shown another embodiment of a point. In this embodiment, the point comprises: An unused or available field 1402. This available bit provides options for later updates to the point. For example, if a new feature is added to the three-dimensional representation then the available bit may be used to signal (e.g. to flag) this new feature. This available bit may have a size of 1 bit. A first underscan size (size.ul) field 1404 and a second underscan size (size.uO) field 1410 that together indicate the underscan factor of the point, e.g. as (size.ul « 1 | size.uO). Each of the first underscan size and the second underscan size fields may have a size of 1 bit. A height size (size.h) field 1406 that indicates a height of the point. The height size field may have a size of 2 bits. A width size (size.w) field 1412 that indicates a width of the point. The width size field may have a size of 2 bits. A type field and / or a glass field 1408 that Indicates a type of the point and / or indicates whether the point is a point on a translucent and / or reflective surface. The glass flag field may have a size of 1 bit.. A XY 1414 field that defines an angle of the point from a capture device identified in a capture device field 1418. The value in this XY field may be decomposed into an X value and a Y value by applying respectively modulo and division by the particle stride. The particle stride may, for example, be defined in a header of the bitstream or may be defined by a file type associated with the bitstream. The particle stride may, for example, be 4000. The XY field may have a size of 24 bits. A normal field 1416 that together indicate a normal of a surface on which the point lies. The normal field may be an index that relates to a table (e.g. a 256-value table). This table may be in a header of the bitstream or may be defined by a file type associated with the bitstream. The normal field may have a size of 8 bits. A capture device field 1418 that indicates a capture device associated with the point. This capture device field may have a size of 8 bits. A motion vector (or movement vector)flag field 1420 that indicates whether the point is associated with a movement vector. The movement vector flag field may have a size of 1 bit. A distance field 1422 that indicates a distance of the point from the capture device identified in the capture device field 1418. The distance field may have a size of 15 bits. A transparency (alpha) field 1424 that indicates a transparency of the point. This transparency field may have a size of 8 bits. A hdr24.Y field 1426, a hdr24.y field 1428, and a hdr24.x field 1430 that each define a separate component of (the colour of) an attribute of the point (e.g. an attribute for a left eye). These fields may have sizes of, respectively, 10 bits, 7 bits, and 7 bits. A motion vector (or movement vector) phi field 1432 that defines a phi value for a movement vector associated with the point (if the point is associated with the movement vector). Typically, if the movement vector flag 1420 is set to 0, then the movement vector phi field can be defined freely (e.g. it may be defined as a string of zeros that can be simply encoded). A xor_hdr24.Y field 1434, a xor_hdr24.y field 1436, and a xor_hdr24.x field 1438 that each define a separate component of (the colour of) a further attribute (e.g. an attribute for a right eye). These fields typically define the further attribute as the XOR of the value of this further attribute and the value of the attribute defined by the hdr24.Y field 1426, the hdr24.y field 1428, and the hdr24.x field 1430. These xor_hdr24.Y, xor_hdr24.y, and xor_hdr24.x fields may have sizes of, respectively, 10 bits, 7 bits, and 7 bits. The examples provided by Figures 21a and 21b contain a number of different features. For example, the example of Figure 21a defines a normal using two normal fields and octahedron normal encoding; the example of Figure 21 b defines a normal using an index value and a table of available normal values. The method used to signal the normal in a given implementation may depend, for example, on the amount of accuracy needed for normal values and may depend on the number of other features that are signalled by the point (e.g. the space available). Referring to Figure 21c, there is shown another embodiment of a point. In this embodiment, the point comprises: An unused or available field 1502. This available bit provides options for later updates to the point. For example, if a new feature is added to the three-dimensional representation then the available bit may be used to signal (e.g. to flag) this new feature. This available bit may have a size of 1 bit. A first underscan size (size.ul) field 1504 and a second underscan size (size.uO) field 1510 that together indicate the underscan factor of the point, e.g. as (size.ul « 1 | size.uO). Each of the first underscan size and the second underscan size fields may have a size of 1 bit. A height size (size.h) field 1506 that indicates a height of the point. The height size field may have a size of 2 bits. A width size (size.w) field 1512 that indicates a width of the point. The width size field may have a size of 2 bits. A type field and / or a glass field 1508 that Indicates a type of the point and / or indicates whether the point is a point on a translucent and / or reflective surface. The glass flag field may have a size of 1 bit. An XY 1514 field that defines an angle of the point from a capture device identified in a capture device field 1418. The value in this XY field may be decomposed into an X value and a Y value by applying respectively modulo and division by the particle stride. The particle stride may, for example, be defined in a header of the bitstream or may be defined by a file type associated with the bitstream. The particle stride may, for example, be 4000. The XY field may have a size of 24 bits. A rgb.g field 1516, a rgb.rfield 1518, and a rgb.b field 1530 that each define a separate component of (the colour of) an attribute of the point (e.g. an attribute for a left eye). This attribute may be defined as a location in a red-green-blue (rgb) space. These fields may each have a size of 8 bits. A normal field 1524 that together indicate a normal of a surface on which the point lies. The normal field may be an index that relates to a table (e.g. a 256-value table). This table may be in a header of the bitstream or may be defined by a file type associated with the bitstream. The normal field may have a size of 8 bits. A capture device field 1526 that indicates a capture device associated with the point. This capture device field may have a size of 8 bits. A motion vector (or movement vector) flag field 1520 that indicates whether the point is associated with a movement vector. The movement vector flag field may have a size of 1 bit. A distance field 1522 that indicates a distance of the point from the capture device identified in the capture device field 1418. The distance field may have a size of 15 bits. A transparency (alpha) field 1528 that indicates a transparency of the point. This transparency field may have a size of 8 bits. A motion vector (or movement vector) phi field 1532 that defines a phi value for a movement vector associated with the point (if the point is associated with the movement vector). Typically, if the movement vector flag 1520 is set to 0, then the movement vector phi field can be defined freely (e.g. it may be defined as a string of zeros that can be simply encoded). A delta_ rgb.b field 1534, a delta_rgb.g field 1536, and a delta_rgb.r field 1538 that each define a separate component of (the colour of) a further attribute (e.g. an attribute for a right eye). These fields typically define the further attribute as a difference between the value of this further attribute and the value of the attribute defined by the rgb.g field 1516, the rgb.rfield 1518, and the rgb.b field 1530. These delta_rgb.b, delta_rgb.g, and delta_rgb.r fields may each have a size of 8 bits. The example of Figures 21a may be particularly suitable for defining points for a cubemap representation. The examples of Figures 21b and 21c may be particularly suitable for defining points for ecospherical representations. Referring to Figure 21 d, there is shown an embodiment of a movement vector. In this embodiment, the movement vector comprises: A first direction (direction x) field 1602 and a second direction (direction y) field 1604 that together define a direction of the movement vector. The first direction field and the second direction field may each have a size of 8 bits. A length field 1606 that defines a magnitude of the movement vector. The length field may have a size of 16 bits. Referring to Figure 21 e, there is shown another embodiment of a movement vector. In this embodiment, the movement vector comprises: A radius field 1702 that defines a magnitude of the movement vector. The radius field may have a size of 8 bits. A first angle (theta) field 1704 and a second angle (phi) field 1706 that together define a direction of the movement vector. The first angle field and the second angle field may each have a size of 12 bits. Referring to Figure 21 f, there is shown another embodiment of a movement vector. In this embodiment, the movement vector comprises: A radius field 1802 that defines a magnitude of the movement vector. The radius field may have a size of 8 bits. An angle (theta) field 1804 that defines a component of a direction of the movement vector. The angle field may have a size of 8 bits. This embodiment of Figure 21 f is suitable for use with points that define a component of a direction of the movement vector such as those points described in the examples of Figures 20b and 20c. Referring to Figure 21g, there is shown an embodiment of a texture point. In this embodiment, the texture point comprises: A transparency (alpha) field 1902 that indicates a transparency of the texture point. This transparency field may have a size of 8 bits. A plurality of motion vector (or movement vector) fields 1904, 1914, 1924. These movement vector fields may be non-contiguous and may be combined by a computer device parsing the texture point in order to determine movement vector data for the texture point. The movement vector data may relate to a plurality of movement vectors (e.g. one for each corner of the texture point). These movement vector fields may sum to 21 bits and may be, respectively, 2 bits, 16 bits, and 3 bits. A diagonal orientation field 1906. that indicates a diagonal of the texture point about which the texture point is able to bend. This diagonal orientation field may be a flag that indicates whether a bending line of the texture point is a default value and / or that indicates that a bending line is a specific one of two possible options. The diagonal orientation field may have a size of 1 bit. A stereo flag 1908 that indicates whether the same texture patch should be used for each attribute (e.g. each eye) of the texture point. The stereo flag may have a size of 1 bit. A texture index field 1910 that identifies a texture patch that is associated with the texture point (e.g. by identifying an index in a texture atlas). The texture index field may have a size of 20 bits. A capture device field 1912 that indicates a capture device associated with the texture point. This capture device field may have a size of 8 bits. A bend field 1916 that indicates an amount of bending of the texture point about a bending line indicated by the diagonal orientation field 1906. The bend field may have a size of 8 bits. A type field and / or a glass field 1918 that Indicates a type of the point and / or indicates whether the point is a point on a translucent and / or reflective surface. The glass flag field may have a size of 1 bit. A distance field 1920 that indicates a distance of the texture point from the capture device identified in the capture device field 1912. The distance field may have a size of 15 bits. A normal x field 1922 and a normal y field 1924 that together indicate a normal of a surface on which the texture point lies using, e.g., octahedron normal encoding. The normal x and the normal y fields may each have a size of 8 bits. A Y field 1926 and an X field 1928 that together define an angle of the texture point from the capture device identified in the capture device field 1912. Typically, these fields define an angular bracket covered by the point. The Y field and the X field may have, respectively, sizes of 12 bits and 13 bits. A reff field 1930 that Indicates whether the point moves with a viewing zone of the three-dimensional representation. The reff field may have a size of 1 bit.. A motion vector (or movement vector) flag field 1932 that indicates whether the texture point is associated with a movement vector. The movement vector flag field may have a size of 1 bit. An underscan field 1934 that Indicates whether a width and / or a height of the point should be scaled based on a value in the underscan field. The underscan field may have a size of 2 bits. Referring to Figure 21 h, there is shown another embodiment of a texture point. In this embodiment, the texture point comprises: A transparency (alpha) field 2002 that indicates a transparency of the texture point. This transparency field may have a size of 8 bits. One or more unused or available fields 2004, 2012. These fields may be used to add functionality to the texture points following an initial implementation of the texture points. These available fields may each have a size of 4 bits. A first left eye texture index ‘LY’ field 2006 and a second left eye texture index ‘LX’ field 2008 that together provide an index for a texture patch for a left eye. The LY field and the LX field may each have a size of 10 bits. A capture device field 2010 that indicates a capture device associated with the texture point. This capture device field may have a size of 8 bits. A first right eye texture index ‘RY’ field 2014 and a second right eye texture index ‘RX’ field 2016 that together provide an index for a texture patch for a right eye. The RY field and the RX field may each have a size of 10 bits. A type field and / or a glass field 2018 that Indicates a type of the point and / or indicates whether the point is a point on a translucent and / or reflective surface. The glass flag field may have a size of 1 bit. A distance field 2020 that indicates a distance of the texture point from the capture device identified in the capture device field 2010. The distance field may have a size of 15 bits. A normal x field 2022 and a normal y field 2024 that together indicate a normal of a surface on which the texture point lies using, e.g., octahedron normal encoding. The normal x and the normal y fields may each have a size of 8 bits. A Y field 2026 and an X field 2028 that together define an angle of the texture point from the capture device identified in the capture device field 2010. Typically, these fields define an angular bracket covered by the point. The Y field and the X field may have, respectively, sizes of 12 bits and 13 bits. A reff field 2030 that Indicates whether the point moves with a viewing zone of the three-dimensional representation. The reff field may have a size of 1 bit. A motion vector (or movement vector) flag field 2032 that indicates whether the texture point is associated with a movement vector. The movement vector flag field may have a size of 1 bit. A size field 2034 that indicates a size of the point. The size field may comprise an index that identifies one or more of: a height of the point, a width of the point, and / or an underscan or overscan value of the point. The size field may have a size of 5 bits. Figures 21 a - 21 e show exemplary (detailed) implementations for signalling points, movement vectors, and texture points. The features of any of these components may be used with any other embodiments of the components. For example, Figure 20c indicates an implementation where an attribute of a point is signalled as a colour in a red-green-blue colour space. It will be appreciated that such a method of signalling the attribute may be implemented in other embodiments of points. It will be appreciated that the order of the fields in each component may be varied (e.g. there is no need to signal the transparency value as the first field of a point). This variability of order is demonstrated in the examples below. In general, the fields of the various components (and the locations of these field) are predetermined and are known by the computer device parsing the bitstream. For example, the structure of a point may be defined as part of a file extension so that the computer device parsing the bitstream can immediately identify the location of the alpha field 1002 of a point based on this known definition. Equally, the structures of any components signalled in the bitstream may be defined in a header of the bitstream where this enables computer devices that have not pre-installed software to parse this bitstream. Typically, integers in the bitstream have a little-endian byte order. Equally, integers may have a big-endian byte order. Typically, the bitstream is encoded using, e.g. entropy encoding techniques in order to reduce the size of the bitstream. It will be appreciated that various methods of encoding are useable with the bitstream. An exemplary arrangement of the bitstream may comprise components with the form struct Particle (or struct Point) { uint32_t a,b,c,d; }; sizeof(Particle) == 16 struct MV2020 { uint32_t mv; } sizeof(MV2020) == 4 struct MV4 { uint32_t mv; } sizeof(MV4) == 4 struct MV2 { uint16_t mv; } sizeof(MV2) == 2 Bitstream qeneration / parsinci architectures Referring to Figure 22, there is shown an architecture for, and a method of, generating a bitstream, e.g. a bitstream comprising a point as describes above. The architecture of Figure 22 may be provided by the image generator 11 and / or the encoder 12. The architecture for generating the bitstream is arranged to identify an input model 201 that comprises point data. The input model is typically a three-dimensional computer model where attribute data can be extracted from the input model using virtual capture devices. Equally, the input model may comprise a real environment, where attribute data is extracted from this environment using real capture devices, such as cameras and / or sensors. Typically, the architecture comprises a capture device identifier module, which module is arranged to identify (e.g. the locations of) one or more capture devices, which capture devices can be associated with points of the three-dimensional representation. Typically, the architecture comprises a viewing zone module for identifying one or more viewing zones associated with the three-dimensional representation. The capture devices may be defined with reference to a viewing zone. For example, each viewing zone may be associated with one or more capture devices located about this viewing zone. The architecture comprises a point determination module 202 for determining point locations of the three-dimensional representation and a point attribute determination module 203 for determining point attributes (and, e.g. further attributes and / or normals) associated with these points. This typically comprises capturing point locations and attributes as described with Figures 3, 4a, and 4b. In short, one or more capture devices is arranged to capture a point location and / or attribute at a given angle. This enables a point to be defined with a location, an attribute, a further attribute, and / or a normal value. The location of the point may be defined with reference to a capture device, a distance of the point to the capture device, and / or an angle of the point from the capture device. The architecture comprises a motion vector determination module 204 for determining motion vectors (or ‘movement’ vectors) for one or more of the points. The motion vectors may be determined from movement information associated with surfaces relating to the points. For example, where the input model comprises a computer-generated model, each surface in the model may be associated with a velocity. The motion vectors can then be determined from these surfaces. Where the input model is a real environment, the motion vectors may be generated using sensors (e.g. cameras that determine velocities of surfaces in the environment). In some embodiments, motion vectors are determined based on a comparison of points captured at different times. With the aforementioned modules, the architecture is able to generate a bitstream that provides information about the three-dimensional representation. In particular, the architecture may comprise a bitstream generation module 208 that generates this bitstream. The architecture may comprise an encoder, where the architecture is arranged to encode the bitstream (e.g. the points of the bitstream) in order to reduce a size of the bitstream. In some embodiments, the architecture (and / or a separate architecture) is arranged to process the point data and / or the motion vector data in order to improve the efficiency of the bitstream. In particular, the architecture may comprise: A point location update module 205 for updating the location information of one or more points of the three-dimensional representation. In particular, the point location update module may be arranged to: identify a point that is associated with a first capture device and update the point so as to be associated with a second capture device. This process typically comprises determining an updated distance value and an updated angle value for the point, the updated distance and angle values defining the location of the point in relation to the second capture device. A detailed method of updating location information is described in the application GB2405096.5, which is incorporated herein by reference. An aggregate point detection module 206 for determining one or more aggregate points. In particular, the aggregate point detection module may be arranged to determine one or more aggregate points based on a similarity of attribute values and / or normal values and / or capture devices of points in the three-dimensional representation. More specifically, the aggregate point detection module 206 may be arranged to identify a plurality of points in adjacent angular brackets that area associated with the same capture device, the points having attribute values and / or normal values that are within threshold levels of similarity. This plurality of points can then be replaced with an aggregate point with a size greater than 1. A detailed method of determining such aggregate points is described in the application GB2406499.0, which is incorporated herein by reference. A texture point determination module 207 for determining one or more texture points. Typically, the texture point determination module is arranged to process the three-dimensional representation after the aggregate point detection module. As described above, the texture point determination module 207 may be arranged to determine one or more texture points based on a similarity of normal values and / or capture devices of points in the three-dimensional representation. More specifically, the texture point determination module 207 may be arranged to identify a plurality of points in adjacent angular brackets that area associated with the same capture device, the points having normal values that are within threshold levels of similarity. Typically, the texture point determination module is arranged to identify the plurality of points based on a predetermined size requirement (e.g. the texture point determination module may be arranged to identify one or more 8x8 pluralities of points). This plurality of points can then be replaced with a texture point with a size greater than 1, the texture point referencing a texture patch in a texture atlas. Determining the texture points may also comprise determining motion vectors for the texture points, e.g. determining motion vectors for each corner of each texture point. The architecture may be arranged to compile such a texture atlas, which texture atlas may be a two-dimensional image composed of one or more texture patches. A detailed method of determining texture patches is described above and is described in the application GB2406502.1, which is incorporated herein by reference. These various modules may be implemented on one or more computer devices, for example an initial three-dimensional representation may be determined on a first device, a bitstream may be generated based on this initial three-dimensional representation and this bitstream may be transmitted to a second device, the points defined in the bitstream may be processed on the second device so as to determine one or more aggregate points and / or texture points, and a processed bitstream may then be generated at the second device. The encoding and / or decoding of the bitstream may, for example, comprise encoding the bitstream using entropy coding techniques and / or video encoding techniques (e.g. HEVC, VVC, LCEVC, and / or VC-6 techniques). Referring to Figures 23a and 23b, there is described an architecture for parsing the bitstream. This architecture may comprise a decoder that is arranged to receive an encoded bitstream, to decode the bitstream, and to parse the bitstream so as to determine point information and / or motion vector information from the bitstream. The architecture comprises a receiving module 211 for receiving a bitstream (and, e.g. decoding the bitstream). The architecture comprises a point determining module 212 for determining points from the bitstream. The points may, for example, be determined as being in one or more parts of the bitstream. The points may be parsed so as to extract the fields described in Figures 20a - 21 h. The architecture comprises a motion vector determining module 213 that is arranged to determine motion vectors. The motion vectors may then be associated with points (e.g. based on the location of the motion vectors in the bitstream and / or based on flags of the points and / or the motion vectors. The architecture may comprise a rendering module 214 for rendering one or more two-dimensional images based on the points and / or the motion vectors extracted from the bitstream. Specifically, the rendering module may render a left eye image and a right eye image at one or more times. Typically, the bitstream defines a plurality of three-dimensional representations associated with different times, where the rendering module may be arranged to determine one or more intermediate two-dimensional representations based on locations of points in these three-dimensional representations and also movement vectors associated with these points (as has been described above). The rendering of the two-dimensional images is based on the fields of the points defined in the bitstream. For example, the images are rendered based on the type fields 1010, 1216 of the points as well as the locations, sizes, movement vectors, and locations of the points. Referring to Figure 23b, the architecture may comprise a texture points identification module 222 for identifying texture points in the bitstream. This module may be arranged to identify the texture points based on a location of these points in the bitstream and / or based on transparency values of points in the bitstream and / or based on a texture point flag of a point. The architecture may comprise a texture atlas identification module 222 for identifying one or more texture atlases associated with the texture points. Typically, the points and / or texture atlases associated with each three-dimensional representation are located in different sections of the bitstream so that a texture atlas associated with a three-dimensional representation can readily be identified. The architecture may then comprise a point attribute identification module 223 that is arranged to determine attributes of a point based on the fields of a point in the bitstream. In particular, the point attribute identification module may be arranged to determine attribute values of a point based on whether that point is a texture point. Alternatives and modifications It will be understood that the present invention has been described above purely by way of example, and modifications of detail can be made within the scope of the invention. The representation is typically arranged to provide an extended reality (XR) experience (e.g. a representation that is useable to render a XR video). The term extended reality (XR) covers each of virtual reality (VR), augmented reality (AR), and mixed reality (MR) and it will be appreciated that the disclosures herein are applicable to any of these technologies. The representation may be encoded into, and / or transmitted using, a bitstream, which bitstream typically comprises point data for one or more points of the three-dimensional representation. The point data may be compressed or encoded to form the bitstream. The bitstream may then be transmitted between devices before being decoded at a receiving device so that this receiving device can determine the point data and reform the three-dimensional representation (or form one or more two-dimensional images based on this three-dimensional representation). In particular, the encoder 12 may be arranged to encode (e.g. one or more points of) the three-dimensional representation in orderto form the bitstream and the decoder 14 may be arranged to decode the bitstream to generate the one or more two-dimensional images. In some embodiments, the scene comprises a static scene; alternatively, in some embodiments the scene comprises a video and / or a moving (e.g. non-static) scene. That is, in some embodiments the scene 5 comprises a static scene, such as a building, where a viewer is able to move through this scene, e.g. to view different rooms of the building, but where the scene itself does not change. In some embodiments, the scene comprises a moving scene, where elements of the scene vary in time even where the viewer remains stationary. It will be appreciated that typically the scene comprises both static and moving elements where, for example, non-static elements move in front of a static background. 10 The detailed description has primarily described a bitstream. The disclosure extends to methods, apparatuses, and systems for encoding, decoding, transmitting, transcoding, and / or receiving a bitstream. The disclosure further extends to computer-readable medium and / or a computer programme product for storing the bitstream as well as a storage and / or memory comprising a rendition of the bitstream. Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect 15 on the scope of the claims.

Claims

1. A bitstream that defines a viewing zone of one or more three-dimensional representations of a scene, the bitstream defining a location and / or a size of the viewing zone, wherein the viewing zone defines a volume of the scene in which a user is able to move through the scene with six degrees-of-freedom.

2. The bitstream of claim 1 wherein the user is not able to move outside of the viewing zone while viewing the scene.

3. The bitstream of claim 1 or 2, comprising a movement vector associated with the viewing zone, the movement vector indicating a movement of the viewing zone over a time period.

4. The bitstream of any preceding claim, comprising a definition of one or more points of the three-dimensional representations.

5. The bitstream of claim 4, wherein the points are defined with reference to the viewing zone.

6. The bitstream of claim 4 or 5, wherein each point comprises a relative movement field that indicates whether said point moves along with a viewing zone of the three-dimensional representation.

7. The bitstream of claim 6, wherein the relative movement field indicates whether a movement vector of the point is defined relative to a movement vector of the viewing zone.

8. The bitstream of claim 7, wherein the relative movement field comprises a flag that indicates that a movement vector of the point is defined relative to a movement vector of the viewing zone.

9. The bitstream of any preceding claim, wherein the viewing zone has a volume of less than five cubic metres (5m3).

10. The bitstream of any preceding claim, wherein the bitstream defines an arrangement of capture devices associated with the viewing zone.

11. The bitstream of any preceding claim, comprising an indication of viewing zone information for a plurality of three-dimensional representations.

12. The bitstream of any preceding claim, comprising a plurality of segments, where each segment is associated with a different viewing zone and wherein each segment defines one or more points associated with that viewing zone.

13. The bitstream of any preceding claim, wherein the bitstream defines a plurality of viewing zones.

14. The bitstream of any preceding claim, wherein the bitstream indicates a change in a viewing zone between a plurality of three-dimensional representations.

15. The bitstream of any preceding claim, wherein the one or more viewing zones are associated with a plurality of three-dimensional representations.

16. The bitstream of any preceding claim, comprising information about one or more of: a resolution; a quality; and a frequency of one or more three-dimensional representations defined in the bitstream.

17. The bitstream of any preceding claim, comprising information about a device arranged to view a threedimensional representation defined in the bitstream and / or a device that is compatible with the bitstream.

18. The bitstream of any preceding claim, comprising a texture atlas that contains attribute values for one or more points, wherein the texture atlas is referenced by the one or more points.

19. The bitstream of claim 18, wherein the texture atlas is associated with a plurality of three-dimensional representations.

20. The bitstream of claim 18 or 19, comprising a plurality of texture atlases, wherein the texture atlases are located in the same part of the bitstream.

21. The bitstream of any of claims 17 to 20, comprising a plurality of texture atlases encoded as a video, wherein each frame of the video comprises a texture atlas.

22. A device for generating the bitstream of any preceding claim.

23. A device for parsing the bitstream of any of claims 1 to 21.

24. A method of generating a bitstream so that the bitstream defines a viewing zone of one or more three-dimensional representations defined in the bitstream, the method comprising:including, in the bitstream, a definition of the viewing zone, the definition indicating a location and / or a size of the viewing zone;wherein the viewing zone defines a volume of the scene in which a user is able to move through the scene with six degrees-of-freedom.

25. A method of parsing a bitstream defining a viewing zone of one or more three-dimensional representations, the method comprising:identifying, in the bitstream, a definition of the viewing zone, the definition indicating a location and / or a size of the viewing zone;wherein the viewing zone defines a volume of the scene in which a user is able to move through the scene with six degrees-of-freedom.A