Processing a three-dimensional representation of a scene
By identifying and filling holes in three-dimensional representations and incorporating occluded objects, the method addresses processing and bandwidth challenges, enhancing efficiency and reducing file sizes in three-dimensional data handling.
Patent Information
- Application Number
- GB2024009913
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-01-14
AI Technical Summary
Three-dimensional representations of environments require substantial processing power, large file sizes, and significant bandwidth, making them challenging to process and transfer efficiently.
A method of processing a three-dimensional representation by identifying holes, determining new points, and adding them to the representation, as well as predicting and incorporating occluded objects from adjacent frames, to enhance data completeness and reduce processing requirements.
This approach reduces processing demands and file sizes, enabling more efficient handling and transfer of three-dimensional data while maintaining data integrity and quality.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field of the Disclosure The present disclosure relates to methods, systems, and apparatuses for processing a three-dimensional representation of a scene. Background to the Disclosure Three-dimensional representations of environments are used in many contexts, including for the generation of virtual reality videos, in which depth information for a plurality of points of the representation is used to generate different images for a left eye and a right eye of a user. Typically, substantial processing power is required to determine such a three-dimensional representation, and the file size of files associated with these representations is typically large so that substantial amounts of storage are needed to keep the files and substantial amounts of bandwidth are required to transfer the files. Summary of the Disclosure According to another aspect of the present disclosure, there is described a method of processing a three-dimensional representation of a scene, the method comprising: identifying a hole in the three-dimensional representation; determining a location of the hole; determining a new point, the new point being associated with the hole location; and adding the new point to the three-dimensional representation. According to another aspect of the present disclosure, there is described a method of processing a frame of data, said data comprising points (e.g. wherein the data is a three-dimensional representation of a scene), the method comprising: determining that a movement of a point, between a current frame and a next frame, will reveal an object occluded by the point in the current frame; adding data associated with the object, from a further frame, into the frame, preferably wherein the further frame is the next frame or a preceding frame. According to another aspect of the present disclosure, there is described a method of processing a frame of data, said data comprising points (e.g. wherein the data is a three-dimensional representation of a scene), the method comprising: determining that a position of a point, at a time between a time of a current frame and a time of a next frame, will reveal an object occluded by the point in the current frame; adding data associated with the object, from a further frame, into the frame, preferably wherein the further frame is the next frame or a preceding frame. Preferably, the method comprises the new point is determined based on a point of a further three-dimensional representation. Preferably, the new point is determined based on existing points of the three-dimensional representation. Preferably, the existing points are adjacent the hole. Preferably, a movement vector of the new point is determined based on existing points of the three-dimensional representation; and an attribute value, e.g. a colour value, of the new point is determined based on a point of a further three-dimensional representation. Preferably, the hole comprises a portion of the scene and / or a surface of the scene for which no attribute data is available. Preferably, identifying the hole comprises identifying an area of the three-dimensional representation that is not associated with a point and / or that is not associated with an attribute value, preferably not associated with an opaque attribute value. Preferably, the method comprises identifying the hole comprises identifying a surface of the scene that is occluded from a viewer in a viewing zone of the three-dimensional representation at a first time and that is visible from the viewing zone of the scene at a second time. Preferably, the first time and the second time are each within a time period associated with the three-dimensional representation. Preferably, identifying the hole comprises comparing one or more points of the three-dimensional representation to one or more points of a further three-dimensional representation. Preferably, identifying the hole comprises determining the hole based on a discrepancy between the points of the three-dimensional representation and the points of the further three-dimensional representation. Preferably, the method comprises: determining one or more predicted points for the first three-dimensional representation, preferably determining the predicted points by combining locations and movement vectors of points of the three-dimensional representation; and comparing the predicted points to the points of the further three-dimensional representation. Preferably, the method comprises determining a coverage of the points of the further three-dimensional representation. Preferably, the method comprises converting points of the further three-dimensional representation to a global coordinate system. Preferably, the method comprises determining a coverage of the points of further three-dimensional representation. Preferably, the method comprises converting points of the three-dimensional representation to a global coordinate system. Preferably, the method comprises comparing a coverage of the points of the further three-dimensional representation to a coverage of the points of the three-dimensional representation, preferably to a coverage of predicted points of the three-dimensional representation. Preferably, the location is associated with an angle. Preferably, each point of the three-dimensional representation is associated with a capture device, and the location is determined in dependence on a capture device. Preferably, the location is associated with an angular bracket of a capture device. Preferably, identifying the hole comprises identifying one or more angular brackets associated with the capture device for which no information is available at the second time. Preferably, the angular brackets are surrounded by angular brackets for which information is available. Preferably, identifying the hole comprises determining a range of surfaces that will be visible during a first time period associated with the three-dimensional representation. Preferably, identifying the hole comprises identifying a section of the three-dimensional information for which no point data is available. Preferably, identifying the hole comprises identifying a movement vector associated with a moving point of the three-dimensional representation, the movement vector identifying a movement of the point during a first time period associated with the three-dimensional representation. Preferably, determining the point of the further three-dimensional representation comprises determining a point on a surface behind the moving point of the three-dimensional representation. Preferably, the point is a point that will be revealed during a time period associated with the three-dimensional representation due to the movement of the moving point. Preferably, the method comprises determining a quantised movement vector, and determining the point on the surface based on the quantised movement vector. Preferably, the further three-dimensional representation is an immediately preceding or an immediately successive three-dimensional representation. Preferably, the method comprises: identifying a location of the determined point of the further three-dimensional representation; identifying a motion vector of the determined point of the further three-dimensional representation; predicting a location of the determined point at the time of the three-dimensional representation based on the identified location and the identified motion vector; and adding the determined point to the three-dimensional representation based on the predicted location. Preferably, the method comprises: identifying a hole in the three-dimensional representation; determining a location of the hole; determining attribute values of one or more points adjacent the hole; and adding one or more points to the three-dimensional representation so as to fill the hole the attribute values of the points being defined based on the determined attribute values. Preferably, the points adjacent the hole are associated with a first capture device; and adding one or more points comprises adding one or more points associated with a second capture device. Preferably, the one or more points are associated with a surface that is expanding; and adding the one or more points comprises adding one or more points so as to fill in gaps formed between the points as the surface expands. Preferably, the one or more points are associated with a surface that is rotating; and adding the one or more points comprises adding one or more points so as to fill in gaps formed between the points as the surface rotates. Preferably, the method comprises: identifying movement vectors associated with the one or more points adjacent the hole; and determining movement vectors for the one or more points based on the identified movement vectors. Preferably, determining the movement vectors for the one or more points comprises averaging the identified movement vectors. Preferably, the method comprises: identifying normals associated with the one or more points adjacent the hole; and determining normals for the one or more points based on the identified movement vectors. Preferably, determining normals for the one or more points comprises averaging the identified normals. Preferably, the three-dimensional representation is associated with a viewing zone, the viewing zone comprising a subset of the scene and / or the viewing zone enabling a user to move through a subset of the scene, preferably wherein the user is able to move within the viewing zone with six degrees of freedom (6DoF). Preferably, the viewing zone has a volume of less than 50% of the volume of the scene, less than 20% of the volume of the scene, and / or less than 10% of the volume of the scene. Preferably, the viewing zone has, or is associated with, a volume, preferably a real-world volume, of less than five cubic metres (5m3), less than one cubic metre (1 m3), less than one-tenth of a cubic metre (0.1 m3) and / or less than onehundredth of a cubic metre (0.01 m3). Preferably, the three-dimensional representation comprises a point cloud. Preferably, the method comprises storing the three-dimensional representation and / or outputting the three-dimensional representation. Preferably, the method comprises outputting the three-dimensional representation to a further computer device. Preferably, the method comprises generating an image and / or a video based on the three-dimensional representation. Preferably, the method comprises forming one or more two-dimensional representations of the scene based on the three-dimensional representation. Preferably, the method comprises forming a two-dimensional representation for each eye of a viewer. Preferably, each point is associated with one or more of: a location; an attribute; a transparency; a colour; and a size. Preferably, each point is associated with an attribute for a right eye and an attribute for a left eye. Preferably, the scene comprises one or more of: an extended reality (XR) scene; a virtual reality (VR) scene; an augmented reality (AR) scene; and a mixed reality (MR) scene. Preferably, the method comprises forming a bitstream that includes the three-dimensional representation and / or a two-dimensional image generated using the three-dimensional representation. According to another aspect of the present disclosure, there is described a system for carrying out the aforesaid method, the system comprising one or more of: a processor; a communication interface; and a display. According to another aspect of the present disclosure, there is described an apparatus for determining a point of a three-dimensional representation of a scene, the apparatus comprising: means for (e.g. a processor for) identifying a hole in the three-dimensional representation; means for (e.g. a processor for) determining a location of the hole; means for (e.g. a processor for) determining a new point, the new point being associated with the hole location; and means for (e.g. a processor for) adding the new point to the three-dimensional representation. According to another aspect of the present disclosure, there is described a bitstream comprising the three-dimensional representation and / or a two-dimensional image generated based on the three-dimensional representation. According to another aspect of the present disclosure, there is described a bitstream associated with a three-dimensional representation, the bitstream comprising one or more flags indicating: whether movement vectors are present in the three-dimensional representation; whether a viewing zone movement vector is present in the three-dimensional representation; whether (for each point), a point is arranged to move with the viewing zone (e.g. cockpit flags); an accuracy of the movement vectors and / or a quantisation curve associated with the movement vectors; a time spacing of three-dimensional representations in the bitstream; a unit type of the movement vectors signalled in the bitstream; and whether new points have been included in a three-dimensional representation, e.g. due to de-occlusion occurring. According to another aspect of the present disclosure, there is described an apparatus (e.g. an encoder) for forming and / or encoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus (e.g. a decoder) for receiving and / or decoding the aforesaid bitstream. Any feature in one aspect of the disclosure may be applied to other aspects of the invention, in any appropriate combination. In particular, method aspects may be applied to apparatus aspects, and vice versa. Furthermore, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be construed accordingly. Any apparatus feature as described herein may also be provided as a method feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure, such as a suitably programmed processor and associated memory. It should also be appreciated that particular combinations of the various features described and defined in any aspects of the disclosure can be implemented and / or supplied and / or used independently. The disclosure also provides a computer program and a computer program product comprising software code adapted, when executed on a data processing apparatus, to perform any of the methods described herein, including any or all of their component steps. The disclosure also provides a computer program and a computer program product comprising software code which, when executed on a data processing apparatus, comprises any of the apparatus features described herein. The disclosure also provides a computer program and a computer program product having an operating system which supports a computer program for carrying out any of the methods described herein and / or for embodying any of the apparatus features described herein. The disclosure also provides a computer readable medium having stored thereon the computer program as aforesaid. The disclosure also provides a signal carrying the computer program as aforesaid, and a method of transmitting such a signal. The disclosure extends to methods and / or apparatus substantially as herein described with reference to the accompanying drawings. The disclosure will now be described, by way of example, with reference to the accompanying drawings. Description of the Drawings Figure 1 shows a system for generating a sequence of images. Figure 2 shows a computer device on which components of the system of Figure 1 may be implemented. Figure 3 shows a method of determining a three-dimensional representation of a scene. Figures 4a and 4b show method of determining a point based on a plurality of sub-points. Figure 5 shows a scene comprising a viewing zone. Figures 6a and 6b show arrangements of capture devices for determining points of the three-dimensional representation. Figure 7 shows a point that can be captured by a plurality of capture devices. Figures 8a and 8b show grids formed by the different capture devices. Figure 9 describes a method of determining a location of a point of the three-dimensional representation. Figure 10 shows a method of determining an angle of a point from a capture device used to capture the point. Figures 11a, 11b, and 11c show a series of images. Figures 12a, 12b, and 12c show a method of rendering a two-dimensional image based on a predicted location of a point of a three-dimensional representation. Figure 13 shows a surface that moves from a first location to a second location. Figures 14a and 14b show a scene at different times. Figures 15 and 16 show methods of processing a three-dimensional representation based on surfaces that will be visible during a time period associated with the three-dimensional representation. Figures 17a, 17b, and 17c, and 18 show methods of processing a three-dimensional representation based on a change in a size of a surface that occurs during a time period associated with the three-dimensional representation. Figure 19, 20a, 20b, 20c, and 20d show methods of filling a hole in a three-dimensional representation. Figure 21 shows a bitstream. Description of the Preferred Embodiments Referring to Figure 1, there is shown a system for generating a sequence of images. This system can be used to generate, and then display, a representation of an environment, which may comprise a VR environment (or an XR environment). The system comprises an image generator 11, an encoder 12, a transmitter 13, a network 14, a receiver 15, a decoder 16 and a display device 17. These components may each be implemented on separate apparatuses. Equally, various combinations of these components may be implemented on a shared apparatus; for example, the image generator 11, the encoder 12, and the transmitter 13 may all be part of a single image data generation device. Similarly, the receiver 15, the decoder 16, and the display device 17 may all be a part of a single image rendering device. Typically, the system comprises at least one encoding computer device (e.g. a server of a content provider) and at least one rendering computer device (e.g. a VR headset). Referring to Figure 2, each of the components, and in particular the image generator 11, the encoder 12, the transmitter 13, the receiver 15, the decoder 16 and the display device 17 is typically implemented on a computer device 20, where, as described above, a plurality of these components may be implemented on a shared computer device. Each computer device comprises one or more of: a processor 21 for executing instructions (e.g. so as to perform one or more of the steps of the various methods described below), a communication interface 22 for facilitating communication between computer devices (e.g. an ethernet interface, a Bluetooth® interface, or a universal serial bus (UBS) interface, a memory 23 and / or storage 24 for storing information and instructions (e.g. a random access memory (RAM), a read only memory (ROM), a hard drive disk (HDD) a solid state drive (SSD), and / or a flash memory, and a user interface 25 (e.g. a display, a mouse, and / or a keyboard) for enabling a user to interact with the computer device. These components may be coupled to one another by a bus 25 of the computer device. The computer device 20 may comprise further (or fewer) components. In particular, the computer device (e.g. the display device 17) may comprise one or more sensors, such as an accelerometer, a GPS sensor, or a light sensor. These sensors typically enable the computer device to identify an environmental condition and / or an action of wearer of the display device. Turning back to Figure 1, the image generator 11 is configured to generate a sequence of image data (e.g. a sequence of image frames) to enable the display device 17 to use this image data to display a plurality of images. The image data may comprise one or more digital objects and the image data may be generated or encoded in any format. For example, the image data may comprise point cloud data, where each point has a 3D position and one or more attributes. These attributes may, for example, include, a surface colour, a transparency value, an object size and a surface normal direction. Each attribute may have a value chosen from a continuous range or may have a value chosen from a discrete set. The image data enables the later rendering of images. This image data may enable a direct rendering (e.g. the image data may directly represent an image). Equally, the image data may require further processing in order to enable rendering. For example, the image data may comprise three-dimensional point cloud data, where rendering a two-dimensional image using this data requires processing based on a viewpoint of this two-dimensional image. The image data may comprise depth map data, where one or more pixels or objects in the image is associated with a depth that is specified by the depth map data. The depth map data may be provided as a depth map layer, separate from an image layer. In some contexts, such as MPEG Immersive Video (MIV), the image layer may instead be described as a texture layer. Similarly, in some contexts, the depth map layer may instead be described as a geometry layer. The image data may include a predicted display window location. The predicted display window location may indicate a portion of an image that is likely to be displayed by the display device 17. The predicted display window location may be based on a viewing position (such as a virtual position and / or orientation of the user in a 3D environment) of the user, where this viewing position may be obtained from the display device. The predicted display window location may be defined using one or more coordinates. For example, the predicted display window location may be defined using the coordinates of a corner or center of a predicted display window, and may be defined using a size of the predicted display window. The predicted display window location may be encoded as part of metadata included with the frame. The image data for each image (e.g. each frame) may include further information, which may be provided as a part of an image, e.g. as part of the point cloud data, or as separate layers. In particular, the image data may include audio information or haptic feedback information indicating audio or haptics which can accompany displayed visual data. An audio layer or haptic layer may accompany each image, and may be omitted for images where no accompanying audio or haptics are required. Similarly, the image data may comprise interactivity information, where the image data may contain or indicate elements with which a user can interact. The interactivity information may, for example, define a behaviour of an element, where a user is able to interact with the element based on this behaviour. The behaviour typically defines a change in an element that occurs as a result of a user interaction where this change may comprise a change in the attributes of the element or in the rendering of the element. As an example, where an image contains a target element, the target element may be arranged to disappear when a user interacts with this element, or to provide feedback indicating that the user has interacted with the target. This interactivity data may be provided as part of, or separately to, the image data. The image data may indicate, or may be combinable with, a state of the virtual environment, a position of a user, or a viewing direction of the user. Here, the position and viewing direction may be physical properties of the user in the real-world, or position and viewing direction may also be purely virtual, for example being controlled using a handheld controller. The image generator 11 may, for example, obtain information from the display device 17 that indicates the position, viewing direction, or motion of the user. Equally, the image generator may generate image data such that it can later be combined with this position, viewing direction, or motion, where the image generator may generate a full scene which is only partially viewed by a user depending on the position of that user. In some cases, the generated image may be independent of user position and viewing direction. This type of image generation typically requires significant computer resources such as a powerful GPU, and may be implemented in a cloud service, or on a local but powerful computer. For example, a cloud service (such as a Cloud Rendering Service (CRN)) may reduce the cost per-user and thereby make the image frame generation more accessible to a wider range of users. Here “rendering” refers at least to an initial stage of rendering to generate an image. Further rendering may occur at the display device 17 based on the generated image to produce a final image which is displayed. The image generator 11 may, for example, comprise a rendering engine for initially rendering a virtual environment such as a game or a virtual meeting room. The encoder 12 is configured to encode frames to be transmitted to the display device 17. The encoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. In some embodiments, the image generator 11 may transmit raw, unencoded, data through the network 14. However, such transmission typically leads to a high file size and requires a high bandwidth so that it is typically desirable to encode the data prior to the transmission. The encoder 12 may encode the image data in a lossless manner or may encode the data a lossy manner. The encoder may apply inter-frame or intra-frame compression based on a currently-encoded frame and optionally one or more previously encoded frames. The encoder may be a multi-layer encoder, such as a low complexity enhancement video codec (LCEVC) enabled encoder. Where the generated frames comprise depth map data, the encoder 12 may perform layered encoding on each instance of image data (e.g. each frame) to generate an encoded frame comprising a base depth map layer and an enhancement depth map layer. Encoding a depth map in this way may improve compression. In some applications, such as HDR video, depth maps are desirably highly detailed with a bit depth of up to twelve or fourteen bits, which is a significant increase in the data to be transmitted. As a result, providing ways to improve compression of the depth map can make more realistic depth map-based displays viable when performing rendering or transmission of rendered data in real-time. Furthermore, this type of layered encoding makes it easy to drop (and then pick back up) one or more of the layers, which provides flexibility and tools for bandwidth management. Layered encoding is also helpful as the final decoder / user device (such as a user display device) can choose whether to process these extra layers. For example, in a non-layered approach, the best the end device (i.e. the receiver, decoder or display device associated with a user that will view the images) can do is determine that it does not have enough resources for a given quality (be it resolution, frame rate, inclusion of depth map) and then signal to the controller / renderer / encoder that it does not have enough resources. The controller then will send future images at a lower quality. In that alternative scenario, the end device still unfortunately has to process the higher quality data until the lower quality data arrives, if it can process the received images at all. In some of the described embodiments, this situation is improved upon because when / if the end device determines for example that it does not have the processing capabilities to handle the highest level of quality, then it can drop and / or choose not to process certain layers. The end device may also signal to the controller that it needs a lower level of quality, but in the meantime the end device can only process the number of layers that it can handle. Therefore, the end device can react to conditions much more quickly. In some cases, depth map data may be embedded in image data. In this case, the base depth map layer may be a base image layer with embedded depth map data, and the enhancement depth map layer may be an enhancement image layer with embedded depth map data. Alternatively, when the generated images comprise a depth map layer separate from an image layer and multi-layer encoding is applied, the encoded depth map layers may be separate from the encoded image layers. This has the advantage that the encoded depth map layers can be dropped under some conditions while still retaining image layers that can be displayed (albeit with a lower level of realism). For example, the encoded depth map layers can be dropped by a transmitter or encoder when available communication resources are reduced, or can be dropped by an end device which lacks the processing resources to handle the highest level of quality. Similarly, if some images comprise an audio base layer, a haptic feedback base layer, an audio enhancement layer or a haptic feedback enhancement layer, these can be processed or dropped flexibly. Again similarly, if some images comprise an interactivity data base layer or an interactivity enhancement layer these can be processed or dropped flexibly. For example, certain interactions may only be possible where a threshold bandwidth is available, where complex interactions (e.g. those enabling a conversation with a digital object) may be disabled before less complex interactions (e.g. changing a pixel colour) are disabled. Additionally or alternatively, where the image data comprises point cloud data, the encoder may apply a point cloud data encoding technique such as described in European patent application EP21386059.6, which is incorporated herein by reference. Such a point cloud encoder may act as a base encoder for a layered encoding technique such as LCEVC or VC-6. Notably LCEVC and VC-6 techniques encode and decode a layered signal, but are agnostic about the content type of data encoded in the signal. For example, the signal can include textures, video frames, geometry or depth data, meshes, point clouds, rendering attributes or physics engine attributes. The transmitter 13 may be any known type of transmitter for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The transmitter 13 may be configured to make decisions about how to transmit the image data, and / or may provide feedback to the encoder 12 or the image generator 11. For example, the transmitter may determine available communication resources (e.g. bandwidth) for transmitting image data, and may drop one or more layers from an encoded frame, or indicate to the image generator and / or encoder that image data should be generated and encoded with fewer layers, when insufficient bandwidth is available for transmission of all generated data. As specific examples, the transmitter may be configured to drop a depth map layer, an LCEVC enhancement layer, or a VC-6 enhancement layer from a frame when insufficient communication resources are available. The network 14 provides a channel for communication between the transmitter 13 and the receiver 15, and may be any known type of network such as a WAN or LAN or a wireless Wi-Fi or Bluetooth network. The network may further be a composite of several networks of different types. Many users only have access to a network with a bandwidth of 30MBps which can lead to latency jitter when streaming. The required bandwidth and the observed latency can be reduced by means of tactics such as forward-looking rendering and last-millisecond reprojection, which are enabled by improved compression. The receiver 15 may be any known type of receiver for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The decoder 16 is configured to receive and decode an encoded frame. The decoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. The display device 17 may for example be a television screen or a VR headset. The timing of the display may be linked to a configured frame rate, such that the display device may wait before displaying the image. The display device may be configured to perform warping, that is, to obtain a final display window location, adjust a warpable image to obtain a final image corresponding to a final viewing direction of the user, and display the final image. In this regard, the image data is typically arranged to provide a warpable image for which a portion of the image that is displayed at the display device 17 is dependent on a position or orientation of a viewer. The warpable image may then be rendered before a most up to date viewing direction of the user is known. The warpable image may be transmitted to the display device, or the warpable image may be transmitted to a rendering node which is near to the display device, and the display device or rendering node may perform time warping to generate a displayed image portion based on the warpable image and the most up to date viewing direction ofthe user. As mentioned above, a single device may provide a plurality ofthe described components. For example, a first rendering node may comprise the image generator 11, encoder 12 and transmitter 13. Additional similar rendering nodes may be included in the system, and may work together to generate the sequence of frames. In one case, multiple rendering nodes may each provide separate image data to an image data assembling node; for example, each rendering node may provide a part of a sequence of frames to a frame assembling node. For example, the receiver 15, decoder 16 or display device 17 may be configured to assemble parts of image data from multiple sources to generate a sequence of images for display on the display device. Alternatively, the image data assembling node may be separate from the receiver 15, decoder 16 and display device 17. Additionally or alternatively, multiple rendering nodes may be chained. In other words, successive rendering nodes may add to a sequence of image data as it passes from rendering node to rendering node, and eventually a complete sequence of image data is then provided to the receiver 15. Furthermore, each rendering node may obtain components of a render from multiple upstream rendering nodes and / or distribute components of a render to multiple downstream rendering nodes. A chain of rendering nodes may be useful for performing different rendering tasks that require different quantities of processing resources, or different frame rates. For example, a company may provide distributed processing in the form of a centralised hub which has abundant processing resources but is distant from users, and peripheral locations which have more scarce processing resources but are closer to users. Expensive but fairly static rendering features such as background lighting or environmental impact on sound may be generated at the central hub (for example using ray tracing), while features that require fewer resources but faster responses or higher frame rates may be generated closer to the user. In other words, the more responsive a rendering feature needs to be, the lower latency it needs between the rendering node which generates the feature and the user display and, in a chain of rendering nodes, the node which generates each rendering feature can be chosen based on a required maximum latency of that feature. On the other hand, if it is expensive to generate a rendering feature, then it may be preferable to generate the feature less frequency and with a higher maximum latency. For example, a static, high-quality background feature may be generated early in the chain of rendering nodes and a dynamic, but potentially lower-quality, foreground feature may be generated later in the chain of rendering nodes, closer to the user device. Here, environmental impact on sound means, for example, a set of surfaces may be constructed where each surface has different sound reflection and absorption properties depending upon material and shape. The frame rates may be matched by creating multiple frames with features generated at the lower frame rate, and combining them with the frames with features generated at the higher frame rate. In a nonlimiting embodiment, a preliminary rendering generates volumetric object data including motion vectors at a first (lowest) frame rate, then produces 2D rendered frames plus depth information for a specific user at a second (higher) frame rate, then transmits video plus depth data to the user device, which produces final frames for display via space warping (depth-based reprojections) at a third (highest) frame rate. One or more of these steps may be performed in combination with the other described embodiments. The viewing position of the user may change as additional rendering tasks are performed at different rendering nodes in the chain. Each or any rendering node may obtain an updated viewing position before performing its respective rendering task. Additionally, the system may simultaneously generate multiple sequences of image data for different respective users or different respective display devices. For example, in the context of a VR or AR experience, each user or display device may view a different 3D environment, or may view different parts of a same 3D environment. When using a chain of rendering nodes, each node may serve multiple users or just one user. For example, a starting rendering node (e.g. at a centralised hub) may serve a large group of users. For example, the group of users may be viewing nearby parts of a same 3D environment. In this case, the starting node may render a wide zone of view (“field of view”) which is relevant for all users in the large group. The starting node may send this wide field of view to a first middle rendering node which renders additional aspects of the 3D environment. These additional aspects may for example be aspects which require less processing power to render, or may be aspects which are specific to individual users of the group. Additionally, the middle rendering node may render features in a smaller field ofviewthan the starting node - this smaller field of view may be relevant to each user rather than the group of users. The first middle rendering node may additionally only serve a smaller number of users (e.g. half of the large group of users), with the remaining users being served by a second middle rendering node which also receives the wide field of view from the starting node. The middle rendering node(s) may then send sequences of second partially or fully rendered frames to an end device for each user. The end device may perform further processes such as warping or focal distance adjustments, optionally using depth map data. Preferably, each rendering node encodes the partially or fully rendered frames before transmitting them on to a next rendering node or to the receiver 15. This means that the required communication resources can be reduced when the rendering nodes are separated by one or more networks, or more generally are implemented in a distributed system such as a cloud. However, each rendering node in a chain is encoding a different partially or fully rendered frame, with different data. Therefore, it may be advantageous for different rendering nodes to use different rendering formats and / or encoding formats. For example, the output from a first rendering node may be point cloud data which logically describes a 3D scene. This point cloud data can be encoded using the techniques of EP21386059.6. A second rendering node may then operate on the point cloud data to generate image data that is more readily displayed by a generic display device, without requiring the display device to model the 3D environment. This image data may be encoded using video coding techniques. The chaining of rendering nodes may be extended to arbitrary tree structures, where a rendering node obtains partially rendered frames from more than one preceding rendering node, and generates further partially or fully rendered frames based on the multiple obtained sequences of partially rendered frames. For example, a content rendering network (CRN) comprising numerous rendering nodes may be used to serve a volumetric event to a large number of same-time users, such as users participating in a shared virtual environment. Rendering the same event for each user is far more expensive in terms of computation time and power consumption than rendering the volumetric effect once and performing the rendering equivalent of multicasting the volumetric effect for multiple users. For example, each user may have a second rendering node (such as a VR headset), and the network may comprise a central first rendering node. The first rendering node may render the volumetric event, and distribute partially rendered frames depicting the volumetric event to the different second rendering nodes. The second rendering node for each user may then integrate the partially rendered frames depicting the volumetric event into a view of the virtual environment which is currently being shown to each user, based on parameters such as the user’s virtual position. The receiver 15, decoder 16 and display device 17 may be consolidated into a single device, or may be separated into two or more devices. For example, some VR headset systems comprise a base unit and a headset unit which communicate with each other. The receiver 15 and decoder 16 may be incorporated into such a base unit. In some embodiments, the network 14 may be omitted. For example, a home display system may comprise a base unit configured as an image source, and a portable display unit comprising the display device 17. In the event that the decoder 16 or the display device 17 does not or cannot handle one or more layers, the receiver 15 or another transmitter associated with the decoder or display device may send a corresponding layer drop indication back through the network 14. The layer drop indication may be received by each rendering node. A rendering node which generates partially or fully rendered frames for that specific decoder or display device may cease generating the dropped layer. On the other hand, a rendering node which generates partially or fully rendered frames for multiple end devices may disregard a layer drop indication received from one end device (as the dropped layer is still needed for other devices). Alternatively, rendering nodes which serve multiple end devices may record received layer drop indications, and may cease generating the dropped layer only when all end devices served by the rendering node indicate that the layer is to be dropped. In preferred examples, the encoders or decoders are part of a tier-based hierarchical coding scheme or format. Hierarchical coding enables frames to be communicated with higher resolution and / or higher frame rate than is possible in single-tier coding schemes. In hierarchical coding, one or more enhancement layers is communicated with base data, where the enhancement layers can be used to up-sample the base data at the decoder, for example providing up-sampling in a spatial or temporal dimension. When combined with equivalent down-sampling of the original frames and generation of the enhancement layer at an encoder, hierarchical coding can overall provide lossless compression of data, with higher resolution and / or higher frame rate for a given transmission bit rate. Examples of a tier-based hierarchical coding scheme include LCEVC: MPEG-5 Part 2 LCEVC (“Low Complexity Enhancement Video Coding”) and VC-6: SMPTE VC-6 ST-2117, the former being described in PCT / GB2020 / 050695, published as WO 2020 / 188273, (and the associated standard document) and the latter being described in PCT / GB2018 / 053552, published as WO 2019 / 111010, (and the associated standard document), all of which are incorporated by reference herein. However, the concepts illustrated herein need not be limited to these specific hierarchical coding schemes. A further example is described in WO2018 / 046940, which is incorporated by reference herein. In this example, a set of residuals are encoded relative to the residuals stored in a temporal buffer. LCEVC (Low-Complexity Enhancement Video Coding) is a standardised coding method set out in standard specification documents including the Text of ISO / IEC 23094-2 Ed 1 Low Complexity Enhancement Video Coding published in November 2021, which is incorporated by reference herein. The system describes above is suitable for generating and presenting a representation of a scene, where this scene displays media content to a user. The scene typically comprises an environment, where the user is able to move (e.g. to move their head or to turn their head) to look around the environment and / or to move around the environment. For example, the scene may be a scene of a room in a building, where the user is able to move around the room (e.g. by moving in the real-world and / or by providing an input to a user interface) in order to inspect various parts ofthe room. Typically, the scene is a XR (e.g. a VR) scene, where the user is able to move about the scene in three degrees of freedom (3DoF) or six degrees of freedom (6DoF) so as to experience the scene. As has been described with reference to Figure 1, the image generator 11 may be arranged to determine point cloud data, where each point ofthe point cloud has a 3D position and one or more attributes. More generally, the image generator (or another component) is arranged to determine a three-dimensional representation of a scene, where this three-dimensional representation is thereafter used to generate two-dimensional images that are presented to a user at the display device 17. While the points are typically points of a point cloud, more generally the disclosure extends to any point that is associated with a location and a value. Therefore, the points may, more generally, be considered to be data (or datapoints), which data is associated with a location and a value, and the ‘points’ may comprise polygons, planes (regular or irregular), Gaussian splats, etc. Referring to Figure 3, there is described a method of determining (an attribute for) a point of such a three-dimensional representation. The method comprises determining the attribute using a capture device, such as a camera or a scanner. The scene may comprise a real scene, in which attribute values are captured using a camera, or a virtual scene (e.g. a three-dimensional model of a scene), in which attribute values are captured using a virtual scanner. Where this disclosure describes ‘determining a point’ it will be understood that this generally refers to determining a point that has a location and an attribute value, where determining the point comprises determining the attribute value and / or storing a point that comprises at least an attribute value and a location value (these values may be indirect values, e.g. where the location is identified relative to another point). Once a plurality of points have been captured, these points can be stored as a three-dimensional representation (e.g. a point cloud) so as to enable the reconstruction of the three-dimensional scene based on this representation. Typically, the scene comprises a simulated scene that exists only on a computer. Such a scene may, for example, be generated using software such as the Maya software produced by Autodesk®. The attributes determined using the methods described herein may then depend on virtual objects located within the scene as well as a virtual lighting arrangement used in the scene. In a first step 11, a computer device initiates a capture process for a capture device, the capture process being initiated with an initial azimuth angle (e.g. of 0°) and an initial elevation angle (e.g. of 0°). In a second step 12, the computer device causes a point to be captured using the capture device at the current azimuth angle and current elevation angle. Capturing a point typically comprises assigning an attribute value to the point, which attribute value may, for example, be a color of the point and / or a transparency value of the point. Typically, the point has one or more color values associated with each of a left eye and a right eye of a viewer. Capturing the point may also comprise determining a normal value associated with the point, e.g. a normal of a surface on which the point lies. Typically, capturing the point further comprises determining a location of the point, e.g. by determining a distance of the point from the camera. In practice, determining the point may comprise sending a ‘ray’ from the capture device and then stepping through a computer model to determine which surface of the computer model is impacted by the ray. The color, transparency, and normal of this surface are then recorded alongside the distance of the surface from the capture device. In a third step, 13, the computer device determines whether a point has been captured for the capture device at each azimuth of a range of azimuths and in a fourth step 14, if points have not been captured at each azimuth, then the azimuth angle is incremented and the method returns to the second step 12 and another point is captured. The azimuth angle may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of azimuth angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. Once a point has been captured for each azimuth, in a fifth step 15, the computer device determines whether a point has been captured for the capture device at each elevation of a range of elevations and in a sixth step 16, if points have not been captured at each elevation, then the azimuth angle is reset to the initial value, elevation angle is incremented and the method returns to the second step 12 and another point is captured. The elevation angles may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of elevation angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. In a seventh step 17, once points have been captured for each azimuth angle and each elevation angle, the scanning process ends. This method enables a capture device to capture points at a range of elevation and azimuth angles. This point data is typically stored in a matrix. The point data may then be used to provide a representation of the scene to a user, e.g. the three-dimensional representation formed by the point data may be processed to produce two-dimensional images for each eye of a user, with these images then being shown to a user via the display device 17 to provide a virtual reality experience to the viewer. By using the captured data, a video can be provided to a viewer that enables the viewer to move their head to look around the scene (while remaining at the location of the capture device). It will be appreciated that the capture pattern (or scanning pattern) described with reference to Figure 3 is purely exemplary and that numerous capture patterns are possible. In general, the capture process for each capture device comprises capturing one or more points at one or more azimuth angles and / or one or more elevation angles. The ‘points’ captured by the capture device are typically associated with a size, such as a height, a width, or a depth. That is, the points typically relate to two-dimensional planes / pixels and / or three-dimensional voxels. In this regard, there is necessarily some space between the locations of adjacent points (since if the points had no width, then an infinite number of points would be required to capture points at each angle). The size provides points that depict a non-negligible area of the three-dimensional space so that a plurality of points can be fit together to provide a depiction of the scene to a viewer. The width and height of each point is typically dependent on the distance of that point from the capture device, where more distant points have a larger width / height. The width and height of each point is typically determined so that when each point is displayed, there is no space between adjacent points (indeed, there may be some overlap between points to ensure that no gaps appear between points). This height / width of each point can be determined at the time of capturing the points, or can be determined or defined after the capture of the points. Typically, the points comprise a size value, which is stored as a part of the point data. For example, the points may be stored with a width value and / or a height value. Typically, the minimum width and the minimum height of a point are set by the angle increment of the azimuth angle and the elevation angle respectively. The size may bethen specified in terms of this angle increment and / or in terms of this minimum width / minimum height (e.g. as being a multiple of the angle increment). In some embodiments, the size value is stored as an index, which index relates to a known list of sizes (e.g. if the size may be any of 1x1, 2x1, 1x2, 2x2, pixels this may be specified by using 3 bits and a list that relates each combination of bits to a size). The size may be stored based on an underscan value. In this regard, where an object is very near to the viewing zone it may be captured using an unnecessarily dense arrangement of points. Therefore, certain surfaces or areas of the representation may be associated with an underscan value, which underscan value defines a reduction in the number of points captured as compared to a representation without underscan. The size of the points may be defined so as to indicate this underscan value. In an exemplary embodiment, the underscan value is an integer value between 0 and 3 and the size is stored as a combination of point dimensions (e.g. a width in the range [0,2]) and a height in the range ([0,2]) and an underscan factor (e.g. an underscan factor in the range [0,3]). In some embodiments, the width and the height are dependent on the underscan factor. For example, when the underscan factor exceeds a threshold value, the possible height and width values may be limited. In a specific example, when the underscan factor is 3, the width and the height may be limited to the range [0,1]. The size may then be defined as size = underscan*9 + height*3 + width. Such a method provides efficient storage and indication of width, height, and underscan values. As shown in Figure 4a, typically, for each capture step (e.g. each azimuth angle and / or each elevation angle), a plurality of sub-points SP1, SP2, SP3, SP4, SP5 is determined. For example, where the azimuth angle increment is 0.1° then for an azimuth angle of 0°, sub-points may be determined at azimuth angles of-0.05°, -0.025°, 0, 0.025°, and 0.05° (and similar sub-points may be determined for a plurality of elevation angles). Attribute values of these sub-points may then be combined to obtain an attribute value for the point. For example, a maximum attribute value of the sub-points may be used as the value forthe point, an average attribute value of the sub-points may be used as the value for the point, and / or a weighted average of the sub-points may be used as the value forthe point. It will be appreciated that numerous other methods for combining the attribute values of the sub-points are possible. By determining the attribute of a point based on the attributes of sub-points, the accuracy of the capture process can be increased. While it would be possible to simply reduce the increment of the angle steps to provide a higher resolution scene, by considering sub-points but only storing attributes for points, a balance can be struck between accuracy and file size (since storing every sub-point would lead to a substantial increase in the amount of data that needs storing). With the example of Figure 4a, for each point of the three-dimensional representation that is captured by a capture device, this capture device may obtain attributes associated with each of the sub-points SP1, SP2, SP3, SP4, SP5, combine these attributes to obtain a point attribute, and then store a point with a distance that is an average (e.g. a weighted average) of the distances of the sub-points from the capture device, at the nominal angle of the point, with the point attribute. As shown in Figure 4b, where a plurality of sub-points SP1, SP2, SP3, SP4, SP5 are considered, these points may have different distances from the location of the capture device. In some embodiments, the attributes of the sub-points may be combined in dependence on this distance, e.g. so that sub-points nearer to the capture device have higher weightings. However, the possibility of sub-points with substantially different distances raises a potential problem. Typically, in order to determine a distance for a point, the distances forthe sub-points are averaged. But where the sub-points have substantially different distances and / or are related to different surfaces in the scene, this may result in the point having a distance that does not correspond to any actual surface in the scene. Therefore, the point may seem to hang in space (e.g. to hang between the front and rear surfaces shown in Figure 4b. Similarly, where the attribute values of the sub-points greatly differ, e.g. if the sub-points SP1 and SP2 are white in colour and the sub-points SP3 and SP4 are black in colour, then the attribute value of the point may be substantially different to the attribute value of other points in the scene. In an example, if the scene were composed of black and white objects, the point may appear as a grey point hanging in space between these objects. In some embodiments, the computer device is arranged to aggregate sub-points so as not to create any floating points. For example, the computer device may determine whether the sub-points are spatially coherent by employing a clustering algorithm (e.g. a k-means clustering algorithm). Where the sub-points are spatially coherent (e.g. where a difference in the distance of the sub-points is below a threshold value), these distances may be averaged to obtain a distance forthe point. Where the sub-points are not spatially coherent, the sub-points may be processed to ensure that the distance of any point places it upon a surface; for example, in the system of Figure 4b, sub-points SP1, SP2, and SP3 may be grouped into a first point and sub-points SP4 and SP5 may be grouped into a second point. Since each sub-point is associated with the same capture device and capture angle (all of these sub-points being associated with a capture step that has a particular azimuth angle and elevation angle), these points may be located at the same angle with respect to a capture device. Therefore, to ensure that each sub-point affects the representation considered, the first point (made up of sub-points SP1, SP2, and SP3) may have a smaller distance value than the second point (made up of sub-points SP4 and SP5) and the first point may be assigned a nonzero transparency value so that the second point can be seen through the first point. By capturing points at a plurality of azimuth angles and elevation angles, e.g. using the method described with reference to Figure 3, it is possible to provide a three-dimensional representation of the scene that can later be used to enable a viewer to view the scene from a plurality of angles. More specifically, given the three-dimensional points captured by the capture device, a computer device is able to render a two dimensional representation (e.g. a two-dimensional image) of the scene for each eye of a viewer so as to provide a representation with an impression of depth. The computer device may render a series of two-dimensional representations to enable the viewer to look around the scene, where the two-dimensional representations are rendered based on an orientation of the viewer’s head. In this way, the determined representation is useable to provide, for example, a virtual reality (VR), mixed reality (MR), augmented reality (AR), and / or extended reality (XR) experience to the viewer. To enable such a display, the display device 17 is typically a virtual reality headset, that comprises a plurality of sensors to track a head movement of the user. By tracking this head movement, the display device is able to update the images being displayed to the viewer as the viewer moves their head to look about the scene. Typically, this involves the display device sensing the sensor data to an external computer device (e.g. a computer connected to the display device via a wire). The external computer device may comprise powerful graphical processing units (GPUs) and / or computer processing units (CPUs) so that the external computer device is able to rapidly render appropriate two-dimensional images for the viewer based on the three-dimensional images and the sensor data. In some embodiments, the external computer device may comprise a server device, where the display device 17 may be connected to this server device wirelessly. This enables the two-dimensional images to be streamed from the serverto the display device so as to enable the display of high-quality images without the need for a viewer to purchase expensive computer equipment. In other words, operations that require large amounts of computing power, such as the rendering of two-dimensional images based on the three-dimensional representation, may be performed by the server, so that the display device is only required to perform relatively simple operations. This enables the experience to be provided to a wide range of viewers. In some embodiments, a first two-dimensional image is provided to the display device 17 (and / or a connected device) and this first image is ‘warped’ in order to provide an image for viewing at the display device. The warping of the image comprises processing the image based on the sensor data in order to provide an image that matches a current viewpoint of the viewer. By performing the warping at the display device or another local device, the lag between a head movement of the user and an updating of the two-dimensional representation of the scene can be reduced. One issue with the above-described method of capturing a three-dimensional representation is that it only enables a viewer to make rotational movements. That is, since the points are captured using a single capture device at a single capture location, there is no possibility of enabling translational movements of a viewer through a scene. This inability to move translationally can induce motion sickness within a viewer, can reduce a degree of immersion of the viewer, and can reduce the viewer’s enjoyment of the scene. Therefore, it is desirable to enable translational movements through the scene. To enable such movements, the three-dimensional representation of the scene may be captured using a plurality of capture devices placed at different locations (or the same capture device placed at different locations). A viewer is then able to move around the scene translationally (e.g. by moving between these locations). More generally, by capturing points for every possible surface that might be viewed by a viewer, a three-dimensional representation of a scene may be captured that allows a suitable two-dimensional representation ofthis scene to be rendered regardless of a location of a viewer (e.g. regardless of where a user is standing within a virtual room). This need to capture points for every possible surface (so as to enable movement about a scene) greatly increases the amount of data that needs to be stored to form the three-dimensional representation. Therefore, as has been described in the application WO 2016 / 061640 A1, which is hereby incorporated by reference, the three-dimensional representation may be associated with a viewing zone, or a zone of viewpoints (ZVP), where the three-dimensional representation is arranged to enable a user to move about the viewing zone so as to view the scene. Figure 5 illustrates such a viewing zone 1 and illustrates how the use of a viewing zone limits the amount of image data that needs to be stored to provide a three-dimensional representation of the scene. With the scene shown in this figure, and the viewing zone 1 shown in this figure, it is not necessary to determine attribute data for the occluded surface 2 since this occluded surface cannot be viewed from any point in the viewing zone. Therefore, by enabling the userto only move within the viewing zone (as opposed to around the whole scene) the amount of data needed to depict the scene is greatly reduced. While Figure 5 shows a two-dimensional viewing zone, it will be appreciated that in practice the viewing zone 1 is typically a three-dimensional zone or volume. The viewing zone 1 may, for example, comprise a rectangular volume, or a rectangular parallelepiped, and the viewing zone may have a height of at least 30 cm, a depth of at least 30 cm, and / or a width of at least 30 cm, where these dimensions enable a user to move their head while remaining in the viewing zone. This is merely an exemplary arrangement of the viewing zone; it will be appreciated that viewing zones of various shapes and sizes may be used (e.g. spherical viewing zones). That being said, it is preferable that the viewing zone is limited so as to cover only a part of the volume of the scene, e.g. no more than 50% of the scene no more than 25% of the scene, and / or no more than 10% of the scene. In this regard, if the viewing zone is the same size as the scene, then the three-dimensional representation will simply be a standard representation for virtual reality (that enables a user to move freely about the scene) - and so the use of the viewing zone will not provide any reduction in file size. The viewing zone 1 enables movement of a viewer around (a portion of) the scene. For example, where the scene is a room, the base representation may enable a user to walk around the room so as to view the room from different angles. In particular, the viewing zone enables a user to move through the scene with six degrees-of-freedom (6DoF) movement through the scene, where this aids in the provision of an immersive experience. In some embodiments, the viewing zone 1 may be four-dimensional, where a three-dimensional location of the viewing zone changes over time - and in such embodiments the size and location of the occluded surface 2 may also change over time. More generally, it will be appreciated that viewing zones may be formed in any size or shape, with different sizes and shapes being suitable for different scenes. The volume of the viewing zone 1 is typically selected so that a user is able to move to a degree sufficient to avoid motion sickness and to provide an immersive sensation, while still only enabling a limited amount of movement (where this leads to a smaller file size as compared to an implementation where a user is able to fully move about the scene). Typically, the viewing zone is arranged to enable a user to move their head while they are sitting or standing, but not to freely roam around a room. The viewing zone 1 may have a (e.g. real-world) volume of less than five cubic metres (5m3), less than one cubic metre (1m3), less than one-tenth of a cubic metre (0.1m3) and / or less than one-hundredth of a cubic metre (0.01m3). The viewing zone 1 may also have a minimum size, e.g. the viewing zone may have a volume of at least 1% of the volume of the scene, at least 5% of the volume of the scene, and / or at least than 10% of the volume of the scene. Similarly, the viewing zone may have a volume of at least one-thousandth of a cubic metre (0.01m3); at least one-hundredth of a cubic metre (0.01m3); and / or at least one cubic metre (1m3). The ‘size’ of the viewing zone 1 typically relates to a size in the real world, where if the viewing zone has a length of one metre this means that a user is able to move one metre in the real world while staying within the viewing zone. The size of the viewing zone in the scene may be greater than, equal to, or less than the size of the viewing zone in the real world. For example, the viewing zone may scale a real-world distance so that moving one metre in the real world moves the user less than (or more than) one metre in the scene. This enables the scene to provide different perceptions to the user (e.g. to make the user feel larger or smaller than they are in real life). Similarly, the viewing zone may scale a real-world angle so that rotating one degree in the real world rotates the user less than (or more than) one degree in the scene. Therefore, a viewing zone with a volume of one cubic metre typically connotes a viewing zone in which the user is able to move about a one cubic metre volume in the real world while remaining in the viewing zone. And this may cause the user to move about a volume that is more than, or less than, one metre in the scene. Referring to Figure 6a, in order to capture points for each surface and location that is visible from the viewing zone 1, a plurality of capture devices C1, C2, ..., C9 may be used (e.g. a plurality of virtual scanners and / or a plurality of cameras). Each capture device is typically arranged to perform a capture process, e.g. as described with reference to Figure 3, in which the capture device captures points at a plurality of azimuth angles and elevation angles. By locating the capture devices appropriately, e.g. by locating a capture device at each corner of the viewing zone, it can be ensured that most (or all) points of a scene are captured. Typically, a first capture device C1 is located at a centrepoint of the viewing zone 1. In various embodiments, one or more capture devices C2, C3, C4, C5 may be located at the centre of faces of the viewing zone; and / or one or more capture devices C6, C7, C8, C9 may be located at edges of and / or corners of the viewing zone. Figure 6a shows a two-dimensional view (e.g. a plan view) of a rectangular viewing zone. It will be appreciated that within this viewing zone each capture device may be located on a shared plane. Equally, the various capture devices may be located on different planes. Referring, for example, to Figure 6b, there is shown a three-dimensional view of a cuboid viewing zone, where there is a capture device located: at the centre of the viewing zone; at the centre of each face of the viewing zone; and at each corner of the viewing zone. With this arrangement, many locations in the scene (e.g. specific surfaces) will be captured by a plurality of capture devices so that there will be overlapping points relating to different capture devices. This is shown in Figure 7, which shows a first point P1 being captured by each of a first capture device C1, a sixth capture device C6, and a seventh capture device C7. Each capture device captures this point at a different angle and distance and may be considered to capture a different ‘version’ of the point. Typically, only a single version of the point is stored, where this version may be the highest quality version of the point and / or may be the version of the point associated with the nearest and / or least angled capture device. In this regard, the highest ‘quality’ version of the point is captured by the capture device with the smallest distance and smallest angle to the point (e.g. the smallest solid angle). In this regard, as described with reference to Figures 4a and 4b, capturing a point for a given azimuth angle and elevation angle typically comprises capturing a plurality of sub-points at varying sub-point azimuth and elevation angles spread around the point azimuth and elevation angles. Due to the different spreads of sub-points, each capture device will capture a different version of the point (that has a different attribute) even when the points are at the same location. Capture devices that are close to the point and less angled with respect to the point typically have a smaller spread of sub-points and so typically obtain a version of a point that is sharper than a version of that point captured by more distant capture devices. In some embodiments, a quality value of a version of the point is determined based on the spread of subpoints associated with this version (e.g. based on the perimeter formed by these sub-points and / or based on a surface area or volume bounded by these sub-points). The version of the point that is stored may depend on the respective quality values of possible versions of the points. Regarding the 'versions’ of the points, it will be appreciated that two ‘points’ in approximately the same location captured by each capture device may not have exactly the same location in the three-dimensional representation. More specifically, since each capture device typically projects a ‘ray’ at a given angle, the rays of differing capture devices may contact the surface at different locations for each capture device. Two points may be considered to be two 'versions' of a single point when they are within a certain proximity, e.g. a threshold proximity. For example, where the first capture device C1 captures a first point and a second point at subsequent azimuth angles, and the sixth capture device C6 captures a further point that is in between the locations of the first point and the second point, this further point may be considered to be a ‘version’ of one of the first point and the second point. This difference in the points captured by different capture devices is illustrated by Figures 8a and 8b, which show the separate captured grids that are formed by two different capture devices. As shown by these figures, each capture device will capture a slightly different ‘version’ of a point at a given location and these captured points will have different sizes. Each capture step is associated with a particular range of angles (e.g. a nominal capture angle of 1° might encompass angles from 0.9° to 1.1°), and therefore capture devices that are far from a point to be captured represent a wider region at the capture distance than capture devices closer to that point to be captured. As shown in Figure 8a, the capture device C1 would capture the points P1 and P2 in separate brackets, whereas for the capture device C2 these points are in the same bracket. Therefore, the capture device C2 might determine a single point that encompasses both points P1 and P2, whereas the capture device C1 would determine separate points for these two points. Considering then a situation in which points P1 and P2 are captured separately, and capture device C1 is used to capture point P1 while capture device C2 being used to capture point P2, it should be apparent that the 'sizes’ of these captured points, and the locations in space that are encompassed by the captured points will be based on different grids. For example, the width of the captured point P2 captured by the capture device C2 will be larger than the width of the captured point P1 captured by the capture device C1. The capture process may be determined based on the existence of these different grids, and on the different bracket widths that occur at different distances from a capture device. Figure 8a shows an exaggerated difference between grids for the sake of illustration. Figure 8b shows a more realistic embodiment in which the three-dimensional representation comprises a plurality of points associated with different capture devices, where these points lie on different grids associated with these different capture devices. In order to store the points of the three-dimensional representation, the points may be stored as a string of bits, where a first portion of the string indicates a location of the point (e.g. using x, y, z coordinates) and a second portion of the string locates an attribute of the point. In various embodiments, further portions of the string may be used to indicate, for example, a transparency of the point, a size of the point, and / or a shape of the point. A computer device that processes the three-dimensional representation after the generation of this representation is then able to determine the location and attribute of each point so as to recreate the scene. This location and attribute may then be used to render a two-dimensional representation of the scene that can be displayed to a viewer wearing the display device 17. Specifically, the locations and attributes of the points of the three-dimensional representation can be used to render a two-dimensional image for each of the left eye of the viewer and the right eye of the viewer so as to provide an immersive extended reality (XR) experience to the viewer. The present disclosure considers an efficient method of storing the locations of the points (e.g. at an encoder) and of determining the locations of the points (e.g. at a decoder). As has been described with reference to Figures 5a and 5b, the points of the three-dimensional representation are determined using a set of capture devices placed at locations about the viewing zone, where these capture devices are arranged to capture points at a series of azimuth angles and elevation angles. Typically, each of the capture devices is arranged to use the same capture process (e.g. the same series of azimuth angles and elevation angles), though it will be appreciated that different series of capture angles are possible. For example, there may be a plurality of possible series of capture angles, where different capture devices use different capture angles. In general, the present disclosure considers a method in which points are stored based on a capture device identifier and an indication of a distance of the point from the capture device associated with this capture device identifier. Typically, the point is also associated with an angular indicator, which indicates an azimuth angle and / or an elevation angle of the point relative to the identified capture device. It will be appreciated that the storage of the distance and the angle may take many forms. For example, the distance and the angle of each point may be converted into a universal coordinate system, where each capture device has a different location in this universal coordinate system. In particular, each point may be stored with reference to a centre of this universal coordinate system, which centre may be co-located with a central capture device. Where a point is determined based on a distance and an angle from a capture device of a known location in this universal coordinate system, the coordinates of the point in this universal coordinate system can be determined trivially - and the location of the point may then be stored either relative to the capture device or as a coordinate in the universal coordinate system. The capture device identifier may comprise a location of a capture device (e.g. a location in a co-ordinate system of the three-dimensional representation). Equally, the capture device identifier may comprise an index of a capture device. Similarly, the indication of the azimuth angle and the elevation angle for a point may comprise an angle with reference to a zero-angle of a co-ordinate system of the three-dimensional representation. Equally, the azimuth angle and / or the elevation angle may be indicated using an angle index. In some embodiments, the three-dimensional representation is associated with configuration information, which configuration information comprises one or more of: a set of capture device indexes; locations associated with the capture devices and / or the capture device indexes; a spacing of capture devices (e.g. so that locations of the capture devices can be determined from a location of a first capture device and the spacing); angles associated with a capture process for the capture devices; an azimuth angle increment and / or an elevation angle increment associated with the capture process; and a set of angle indexes (e.g. to match an angle index to an angle). With this configuration information, it is possible to determine a location of each capture device from an index of that capture device and / orto determine a capture angle from a known capture process. Therefore, given two numbers: a capture device index and an angle index (that is associated with a combination of a specific azimuth angle and a specific elevation angle), a location of a capture device and a direction of a point from this capture device can be determined. By also signalling a distance of the point from the signalled capture device, a precise location of the point in the three-dimensional space can be signalled efficiently. Typically, the point is associated with each of: a camera index, a distance, a first angular index (e.g. a first azimuth), and a second angle (e.g. a second elevation) This method of indicating a location of a point enables point locations to be identified using a much smaller number of bits than if each point location is identified using x, y, z coordinates. Referring to Figure 9, there is shown a method of determining a location of a point. This method is carried out by a computer device, e.g. the image generator 11 and / or the decoder 15. In a first step 21, the computer device identifies an indicator of a capture device used to capture the point. Typically, this comprises identifying a portion of a string of bits associated with a capture device index. In a second step 22, the computer device identifies an indicator of an angle of the point from the capture device. Typically, this comprises identifying an angle index, e.g. an azimuth index and / or an elevation index and / or a combined azimuth / elevation index, which index(es) identifies a step of the capture process during which the point was captured. In a third step 23, based on the identifiers, the computer device determines the location of the capture device and the angle of the point from the capture device. The capture device identifier is typically a capture device index, which is related to a capture device location based on configuration information that has been sent before, or along with, the point data. For example, the configuration information may specify: Location of first capture device is (0,0,0). Step between capture devices is (0,0,1) along the grid, then across the grid, then up the grid. - The grid is (10,10,10). With this information, a capture device with an index of 1 can be determined to be located at (0,0,0); a capture device with an index of 5 can be determined to be located at (0,0,4); a capture device with an index of 12 can be determined to be located at (0,1,0), and so on. Equally, the configuration information may specify a list of camera indexes and locations associated with these indexes, where this enables the use of a wide range of setups of capture devices. Typically, the three-dimensional representation is associated with a frame of video. The configuration information may be constant over the frames of the video so that the configuration information needs to be signalled only once for an entire video. Therefore, the configuration information may be transmitted alongside a three-dimensional representation of a first frame of the video, with this same information being used for any subsequent frames (e.g. until updated configuration information is sent). The angle identifier may similarly be related to an angle by a location and an increment that are signalled in a configuration file. For example, the configuration information may specify: An azimuth increment and an elevation increment are each 1°. There are 359 increments for each angle type. With this information: a capture angle with an index of 1 can be determined to be at an azimuth angle of 0° and an elevation angle of 0°; a capture angle with an index of 10 can be determined to be at an azimuth angle of 10° and an elevation angle of 0°; a capture angle with an index of 360 can be determined to be at an azimuth angle of 0° and an elevation angle of 1°; and a capture angle with an index of 370 can be determined to be at an azimuth angle of 9° and an elevation angle of 1°; etc. In a fourth step 24, based on the determined location of the capture device and the determined angle, a location of the point is determined. Typically, this comprises determining the location of the point based on the location of the capture device, the capture angle, and a distance of the point from the capture device (where this distance is specified in the point data for the point). Determining the location of the point typically comprises determining the location of the point relative to a centrepoint of the three-dimensional representation, this location of the point may then be converted into a desired coordinate system and / orthe point may be processed based on its location (e.g. to stitch together adjacent points). The angular identifier typically comprises a first angular identifier and a second angular identifier, where the first identifier provides the azimuthal angle of the point and the second identifier provides the elevation angle of the point. Referring to Figure 10, each angular identifier may be provided as an index of a segment of the three-dimensional representation, where, for example, an index of 0 may identify the point as being in a first angular bracket 101 and an index of 1 may identify the point as being in a second angular bracket 102. In this regard, the capture devices are arranged to perform a capture process, e.g. as described with reference to Figure 3, with a non-infinite angular resolution. Given this non-infinite resolution, each point is not a one-dimensional point located at a precise angle. Instead, each point is a point for a particular area of space, with the size of this area being dependent on the angular resolution as well as the distance of the point from the capture device. In other words, each capture angle determines a point for an angular range (with the range being dependent on the angular resolution). That is, if the capture process leads to points being captured at angles of 10°, 11°, and 12° then this can equally be considered to relate to points being captured at a first range of 9.5°-10.5°, a second range of 10.5°-11.5°, and a third range of 11.5°-12.5°. This is shown in Figure 10, which shows a series of angular brackets, with the size of these angular brackets at a given distance being dependent on the angular resolution. The angular identifier(s) typically comprise a reference to such an angular bracket. Consider, for example, a cube placed with the capture device C1 at the centre of this cube. By dividing this cube into x segments at regular azimuth angles and y segments at regular elevation angles, it is possible to identify any angular range of the representation by reference to an x segment and a y segment (and then the space bracketed by this angular range will depend on both the angular resolution (e.g. the angle between adjacent brackets) and the distance of the point from the capture device). Typically, each capture device has the same capture pattern so that the angular bracketing of each device is the same (albeit centred differently at the location of the relevant capture device). For example, in an embodiment with 1000 equal angular brackets, the angle for each bracket may be 360 / 1000. In some embodiments, different capture devices are associated with different capture patterns, where this may be signalled in configuration information relating to the three-dimensional representation. In some embodiments, each capture device is arranged to capture a point for a plurality of angular brackets, where each bracket is associated with a different angle. The angular spread of each bracket (that is, the angle between a first, e.g. left, angular boundary of the bracket and a second, e.g. right, angular boundary of the bracket) may be the same; equally, this angular spread may vary. In particular, the angular spread may vary so as to be smaller for points which are directly in front of (or behind, or to a side of) the capture device. For example, the embodiment shown in Figure 7 shows an angular bracketing system that is based on a cube. With this system, a cube is placed such that a capture device is located at the centre of the cube and the cube is then split into 1000 sections of equal size (it will be appreciated that the use of 1000 sections is exemplary and any number of sections may be used). Each of these sections is then associated with an angular index. With this arrangement, the angular spread of each section (or bracket) varies, as has been described above. Figure 10 shows a two-dimensional square, where each angular bracket of the square is referenced by an index number (between 1 and 100). In a three-dimensional implementation, an angular bracket of a cube could be indicated with two separate numbers (with a first azimuthal indicator that identifies a 'column' of the cube and a second elevational indicator that identifies a ‘row’ of the cube). Equally, a singular indicator may be provided that indicates a specific bracket of the cube. Therefore, for a cube that is divided into 1000 elevational sections and 1000 azimuthal sections, the bracket may be indicated with two separate indicators that are each between 0 and 999 or with a single indicator that is between 0 and 999999. It will be appreciated that the use of a cube to define the brackets is exemplary and that other bracketing systems are possible. For example, a spherical bracketing system may be used (where this leads to curve angular brackets). Equally, a lookup table may be provided that relates angular indexes to angles, where this enables irregularly spaced brackets to be used. Typically, determining the location of the point comprises determining the location of the point so as to be at the centre of the angular bracket identified by the angular identifier(s). Movement vectors The three-dimensional representation is typically used to render one or more two-dimensional images in order to form a video that can be presented on the display device 17. In particular, a computer device (e.g. the display device 17) may render a series of two-dimensional images for each eye of a user, where a user viewing these two series of images simultaneously is provided with the impression of a three-dimensional scene. Referring to Figure 11a, the series of images typically comprises a plurality of frames of a video, e.g. first to sixth frames F1 ... F6. By showing these frames in succession, the display device is able to provide a viewer with the impression of an unbroken video. Achieving this impression of an unbroken video requires the viewer to be shown a number of frames every second. The exact rate / density at which the frames F1 ... F6 are shown typically depends on a frame rate (or a refresh rate) of the display device 17 and / or a frame rate defined by a creator of a video. Typically, the display device is arranged to display video with a frame rate of at least 60 frames per second (60 ‘fps’), e.g. the frame rate of the display device may be 72 fps. In some embodiments, the image generator 11 is arranged to generate three-dimensional representations at a corresponding rate so that each three-dimensional representation is associated with a single frame of a video. Therefore, if the refresh rate of the display device 17 is 72 fps then the image generator may be arranged to generate 72 three-dimensional representations for every second of video (e.g. to generate three-dimensional representations at 72 fps). However, each three-dimensional representation is typically large in size so that a file that contains this rate of three-dimensional representations might be prohibitively large (leading to difficulties in storing and / or transmitting the file). Furthermore, since the scene may be viewed on a plurality of different display devices with different refresh rates, generating a set of three-dimensional representations that enables each three-dimensional representation to be associated with a single frame of a video for each of the plurality of display devices can require generating a prohibitively large number of sets of three-dimensional representations. Therefore, the present disclosure considers methods of rendering a two-dimensional video, where a number of frames of the two-dimensional video is different to (e.g. greater than or less than) a number of three-dimensional representations used to render these frames. This may be considered as the frame rate of the two-dimensional video being different to the frame rate of the three-dimensional representations. Referring to Figure 11b, the three-dimensional representations may be generated and / or rendered so as to provide a two-dimensional video that has a frame rate that is a multiple of a frame rate of the three-dimensional representations, e.g. where the video comprises six frames F1 ... F6, the image generator 11 may be arranged to generate three three-dimensional representations R1, R2, R3 where each three-dimensional representation is associated with a pair of two-dimensional images. For example, the three-dimensional representations may have a frame rate of 36 fps (i.e. 36 three-dimensional representations may be generated for each second of the scene) and these three-dimensional representations may be used to render a video with a frame rate of 72 fps. More generally, referring to Figure 11c, the present disclosure envisages methods by a plurality of three-dimensional representations of a scene may be used to render a two-dimensional video of any frame rate. This enables the three-dimensional representations to be used with various devices and for various situations. With the example of Figure 11c, the first second and third video frame F1, F2, F3 may be rendered based on the first three-dimensional representation R1, the fourth and fifth video frame F4, F5 may be rendered based on the second three-dimensional representation R2 and the sixth video frame F6 may be rendered based on the third three-dimensional representation R3. For example, the three-dimensional representations may have a frame rate of 24 fps (i.e. 24 three-dimensional representations may be generated for each second of the scene) and these three-dimensional representations may be used to render a video with a frame rate of 70 fps. More generally, it will be appreciated that any pairing of frame rates may be provided using the methods described herein. A simple solution to this problem of differing frame rates is to repeat a frame of the video. For example, the first frame F1 that is rendered based on the first three-dimensional representation R1 could also be used as the second frame F2 and the third frame F3. However, rendering a video in this way can leads to stuttering and ghosting. Regarding ghosting, the brain of a viewer expects an object to move in a consistent manner so that copying frames in the above-described manner can result in a viewer seeing objects in the second and third frames at both expected locations (the locations that would be expected from a previously movement of the objects) and also at the displayed location (the location of the object in the repeated first frame). Therefore, simply repeating frames in order to match the frame rate of the three-dimensional representations to the frame rate of a rendered two-dimensional video is typically unsatisfactory (unless there is very little movement between successive representations). Referring to Figures 12a - 12c, there is described a method of rendering a two-dimensional image from a three-dimensional representation that avoids the problems of stuttering and ghosting. This method is performed by a computer device, for example the image generator 11 or the display device 17. In a first step 101, the computer device identifies a point of a three-dimensional representation. As described above, the points typically comprise a location and at least one attribute value (e.g. a colour). Furthermore, the points may comprise a motion vector that identifies a motion of the point (or a ‘movement vector’, these terms are used interchangeably in this document). This is shown in Figure 16b, which figure shows a point that has a first location L1 and a first motion vector MV1. In a second step 102, the computer device identifies this first (initial) location L1 and identifies the first movement vector MV1. The initial location is typically a location defined in point data of the point. In a third step 103 that is shown in Figure 16c, the computer device identifies a second, predicted, location L2 of the point based on the first location L1 and the first movement vector MV1. In a fourth step 104, the computer device renders a two-dimensional image based on the predicted location L2 of the point In some embodiments, the first motion vector MV1 indicates an end location of the point, that is the motion vector indicates where the point will be at the end of a time period associated with the three-dimensional representation. The predicted location L2 can then be determined by interpolating between the first location L1 and the end location. In practice, the three-dimensional representation is typically associated with a first time period, e.g. t=10 to t=11, with the first location L1 being at t=10 and the first motion vector MV1 indicating the end location of the point at t=11. In order to find the location of the point at t=10.25, the computer device may interpolate between the first location and the end location so that the predicted location L2 is between these locations and is 25% of the way along a line drawn between these locations. The computer device may then consider a further three-dimensional representation that is associated with a second time period, e.g. t=11 to t=12 so as to form a continuous video. This method enables the rendering of a two-dimensional video with a first frame rate based on three-dimensional representations with a second frame rate (which second frame rate is typically smaller than a frame rate of the two-dimensional video). Typically, this involves rendering a plurality of frames of the two-dimensional video based on a single frame of the three-dimensional representation. In a simple example, the first frame F1 may be rendered based on the first (e.g. defined) location of each point in the first three-dimensional representation R1 - e.g. each of F1 and R1 may relate to the same time t=0. Thereafter, for example at time t=0.1, the second frame F2 may be rendered based on the three-dimensional representation. For this second frame F2, the predicted location L2 of each point may be determined based on the first locations L1 and the first movement vectors MV1 of those points. As described in Figure 12a, the method may comprise rendering a two-dimensional image based on the predicted locations L2 of points in a first three-dimensional representation. Additionally, or alternatively, the method may comprise determining a further (e.g. intermediate) three-dimensional representation based on the predicted locations of the points of the three-dimensional representation. For example, based on the movement vectors, a three-dimensional representation may be formed for each frame of a video. This enables a set of three-dimensional representations (and / or a set of two-dimensional images) to be generated for a video with any frame rate. It will be appreciated that in the description below any features described with reference to the generation of an intermediate three-dimensional representation may equally be applied to the rendering of a two-dimensional image (and vice versa). Typically, the first frame F1 is associated with an earlier time than the second frame F2 and the first movement vector MV1 define a ‘forwards’ movement of the point (e.g. defines a direction in which the point is moving). Additionally, or alternatively, the motion vector MV2 may comprise a ‘backwards’ motion vector that indicates a backwards movement (e.g. that defines a direction from which the point has come). Therefore, referring to Figure 12c, the third frame F3 may be determined based on locations and motion vectors from either of, or both of, the first three-dimensional representation R1 and the second three-dimensional representation R2. Any combination of forwards and backwards movement vectors may be used to determine the locations of points in the frames of the video. For example, a first location L1 and a first, forwards, movement vector MV1 from the first three-dimensional representation R1 may be combined with a second location and a second, backwards, movement vector from the second three-dimensional representation in order to determine the predicted location of a point in the frames F2 and F3. This determination may involve a weighted combination of locations and / or motion vectors from a plurality of frames. The backwards movement vector for each point may be the inverse of the forwards movement vector (this implementation is generally used only when the velocity of the point is constant across a time period associated with a plurality of three-dimensional representations). Equally, the backwards movement vector may be a different vector where each point may be associated with one or more of a backwards movement vector and a forwards movement vector. Therefore, the location of the point in the second frame F2, the ‘predicted’ location, may be determined as any one or more of: ^pred = + Lpred = T2 — M72 * t2 Lpred = a(Li + * tj) + b(L2 — MV2 * t2), where a + b = 1 Where: ^pred = the predicted location of a point in an intermediate three-dimensional representation and / or in a two-dimensional image (e.g. a frame of a video). = the location of the point in a first three-dimensional representation before the intermediate three-dimensional representation. L2 - the location of the point in a second three-dimensional representation after the intermediate three-dimensional representation. MVt = the movement vector of the point in a first three-dimensional representation. MV2 = the movement vector of the point in a first three-dimensional representation (it will be appreciated that the sign of the movement vector, i.e. whether the movement vector is positive or negative, is not fixed and is selected in dependence on the exact implementation). is the time between the first three-dimensional representation and the intermediate three-dimensional representation. t2 is the time between the second three-dimensional representation and the intermediate three-dimensional representation. a and b are (optional) weighting factors. Typically, the movement vector has units of velocity (e.g. m / s) so that this movement vector can be combined with a time in order to obtain a (vector) movement of a point. Typically, each three-dimensional representation and each frame is associated with a time, so that the method of Figure 16a may involve determining a time difference between a three-dimensional representation and a frame and determining the predicted location L2 in dependence on this time difference. Typically, the movement vector is a single value (e.g. 5m / s). In some embodiments, the movement vector comprises a range of values, a formula, and / or a reference to a range of values. For example, the movement vector may relate to a speed value of: 5m / s + 0.1t. Such embodiments enable the signalling of movement vectors that change over time (e.g. to indicate that an object is accelerating). In practice, the frame rate of the three-dimensional representations is typically high enough that changes in the velocity of a point in between representations are insufficient to cause problems with video quality. Therefore, using movement vectors that are a single (constant) value typically provides sufficient quality while maintaining an efficient three-dimensional representation. In some embodiments, e.g. those with many erratically moving objects, the use of variable movement vectors (e.g. that have an acceleration and / or a dependence on time) can cause a sufficient improvement in video quality to justify the increase in data that is required to signal these movement vectors. As described above, the location of each point is typically stored in dependence on a capture device used to capture that point. The movement vectors may similarly be stored with reference to a capture device, where, for example, the movement vectors may identify a number of angular brackets through which a point moves in a unit time. Typically, the movement vectors are stored in absolute values. For example, the movement vectors may be stored with x, y, z values based on a cartesian grid or as a speed, an elevation angle, and an azimuthal angle based on a grid. In this regard, typically the scene is associated with a grid that, for example, has a center at the center of the viewing zone and that is aligned with the viewing zone. Therefore, determining the predicted location L2 of the point may comprise converting the first location L1 into an absolute value (e.g. determining a coordinate value for L1 based on a location of the capture device used to capture the point and the distance / angle of the point from this capture device) and then combining this absolute location value with the determined motion of the point (this motion being determined based on the movement vector and the time difference between the first three-dimensional representation and the intermediate three-dimensional representation). The movement vectors of each point of a three-dimensional representation are typically determined during the formation of that three-dimensional representation. In this regard, the three-dimensional representation may be determined based on a computer-generated model (e.g. an animated scene), in which case the objects in this model may be associated with movements. The movement vectors may then be determined directly from the model. For example, the description above has described a method of capturing point data based on scanning rays that are sent by a capture device. These scanning rays may be used to determine each of: a location; an attribute; and a movement vector of a surface impacted by the scanning rays. In some embodiments, the scene represents a real scene where the point data may be captured using (real) cameras. In such embodiments, the motion vectors may be determined based on an analysis of the frames captured by the cameras (e.g. a movement of an object between subsequent frames may be determined and this movement may be used to determine a movement vector). In some embodiments, the determination of the movement vectors uses artificial intelligence or machine learning algorithms. In some embodiments, the movement vectors are determined or modified (e.g. refined) based on a postprocessing process in which similar points in different three-dimensional representations are processed in order to determine a motion vector for these points. For example, the locations of corresponding points in each of the first three-dimensional representation R1 and the second three-dimensional representation R2 may be compared to determine a suitable motion vector (e.g. to determine a forwards motion vector for inclusion a point of the first three-dimensional representation). Such embodiments typically require the representations to be evaluated to determine corresponding points in successive representations; this may involve determining points with similar attributes in successive representations. Viewing zone movement In some embodiments, the viewing zone 1 is arranged to move and so the viewing zone may be associated with a movement matrix or vector. For example, the viewing zone may be moving forward and / or rotating within the scene, where this causes a relative movement between each of the points of the three-dimensional representation and the viewing zone (even points that would otherwise be stationary in the three-dimensional representation).In some embodiments, the viewing zone movement matrix is signalled using one or more (or all) of the following components: a translation movement vector, a rotation element (e.g. in the form of a quaternion), and a scale change. Such a representation enables a computer device to determine appropriate relative movements for each point of the representation based on the viewing zone movement matrix. In this regard, in these embodiments, the locations of one or more points in the three-dimensional representation at a second time may be predicted based on the locations of the points at a first time and based on a movement matrix associated with the viewing zone. The prediction of these points may occur using a process similar to that described above, with the viewing zone movement matrix being inverted to determine a suitable movement vector for applying to each points. In some embodiments, the movement of the viewing zone is instead effected by adding a viewing zone movement matrix to each of the movement vectors of the points in the scene (and then storing these points with this movement vector). However, such an implementation can result in every point in the scene needing a movement vector and this is typically less efficient than associating the viewing zone with a viewing zone movement matrix that can be applied to the points at the time of rendering the scene. In some embodiments, one or more points of the three-dimensional representation is arranged to move with the viewing zone. For example, the viewing zone may be located inside a virtual cockpit, where certain points surrounding the viewing zone are also part of the cockpit (and these points move with the viewing zone). In orderto identify these points, one or more (or each) points of the three-dimensional representation may be associated with a flag (e.g. a ‘cockpit flag'), which flag defines whether a viewing zone movement vector should be applied to the points. Equally, these points may be associated with movement vectors that are defined so as to counteract the viewing zone movement matrix where this precludes the need for the cockpit flag, but requires numerous movement vectors to be added to the three-dimensional representation that would not be needed otherwise. Furthermore, due for example to rounding errors that may occur during the determination of the predicted locations of each point, embodiments that use counteracting movement vectors may result in some relative movement between points. Therefore, the cockpit flag typically provides a benefit in embodiments where there is a viewing zone movement matrix. The viewing zone movement matrix may comprise a single vector so that the viewing zone moves through the scene as a fixed shape. Equally, the viewing zone movement matrix may comprise a rotational component and / or a deforming component. De-occlusion As described above, in orderto generate a continuous video, the computer device (e.g. the image generator 11 or the display device 17) may be arranged to determine an intermediate three-dimensional representation or two-dimensional image based on movement vectors of a first three-dimensional representation. Referring to Figure 13, there is shown a scene similar to that of Figure 5. In this scene of Figure 19, a surface moves from a first position LS1 to a second position LS2, where this movement is based on a movement vector associated with the points of the surface. Both the starting position, with the surface at LS1, and the final position, with the surface at LS2 are covered by a single three-dimensional representation. A subsequent three dimensional representation may then cover future actions in the scene. For the sake of this description, we consider a situation in which a first three-dimensional representation covers a period t=0 to t=1 and a second three-dimensional representation covers a period t=1 to t=2. The period t=0 to t=1 is associated with a plurality of frames of a rendered two dimensional video, where, for example, a first frame may be rendered when the surface is in the first position LS1 and a second frame may be rendered when the surface is in the second position LS2. The first three-dimensional representation is captured such that each point has a position relating to the time t=0 and a movement vector that indicates the movement of that point over the period t=0 to t=1. As has been described with reference to Figure 5, a benefit of the present method in which points are determined in dependence on a viewing zone 1 is that occluded surfaces such as the occluded surface 2 do not need to be captured. Therefore, the first three-dimensional representation, which is associated with the scene of Figure 5 in which the surface is at the first position LS1, does not need to contain points for the occluded surface 2 that are occluded at the time t=0. This provides an increased efficiency over representations in which each and every point (included occluded points) needs to be captured and stored. However, since occluded points are typically not captured in the three-dimensional representation, a problem can occurs for the intermediate three-dimensional representation: at the time of the intermediate three-dimensional representation (e.g. at t=0.5), the surface has moved to the second position LS2 so that a portion 2A of the occluded surface 2 is now visible from the viewing zone, but the first three-dimensional representation may not contain points for this portion of the occluded surface. A potential solution to this problem is to determine all of the surfaces that will be visible at any time covered by the first three-dimensional representation and to determine points for these surfaces. In particular, the computer device may determine a time range associated with the first three-dimensional representation; determine each surface that will be visible from the viewing zone during this time range; and determine points for each of these surfaces. This may involve the computer device capturing points that are initially behind the surface, where this capturing of points can be straightforward when the scene comprises a computer-generated scene (but can be more difficult when the scene comprises a real scene and the capture devices may not have sight of these initially occluded points at the time of capturing points for the first three-dimensional representation). In some embodiments, as described below, the computer device is arranged to process the first three-dimensional representation based on one or more further three-dimensional representations in order to include points in the first three-dimensional representation that capture each surface that will be visible during the time period associated with the first three-dimensional representation. In this regard, referring to Figures 14a and 14b, typically a video is rendered based on a plurality of three-dimensional representations relating to successive time periods. By way of example, the scene of Figures 5 and 13 is shown in Figures 15a and 15b. Figure 15a shows the scene at the time t=0, where the surface is at the position LS1. Figure 15b shows the scene at the time t=1, where the surface is at the position LS2. As can be seen, the occluded portion of the scene moves with the surface, so that the points associated with the portion 2A of the occluded surface are captured in a second three-dimensional representation associated with the second time but not in a first three-dimensional representation captured at a first time. Therefore, as is described below, the present disclosure considers a method in which a first three-dimensional representation is processed based on a second three-dimensional representation to include points that will be visible during a time period associated with the first three-dimensional representation (which points are occluded in an initial arrangement of the first three-dimensional representation). Referring to Figure 16, in a first step 121, the computer device identifies a first location of a hole in a first three-dimensional representation. As used herein, a 'hole' typically connotes a portion of the scene for which no attribute data is available. Within the three-dimensional representation, the hole may present as a gap within an array of points. Due to this lack of attribute data, a computer device that is attempting to render an image that includes the hole is not able to identify attribute values (e.g. colours) that should be rendered at the location of the hole. In practice, the three-dimensional representation may comprise a number of surfaces for which no attribute data is available (e.g. typically no attribute data is captured for the occluded surface 2). As used herein ‘hole’ typically refers to a portion ofthe scene for which no attribute data is available, where this portion is (or becomes) visible to a viewer. In particular, the hole (e.g. the portion 2A ofthe occluded surface) typically relates to a surface ofthe scene that is occluded at a first time and becomes visible at an intermediate time (e.g. a second time), where each ofthe first time and the intermediate time are within a time period associated with the first three-dimensional representation (e.g. where the intermediate time is between t=0 and t=1, e.g. where the intermediate time is at t=0.5). The hole may be determined by determining an area ofthe three-dimensional representation without any points at the intermediate time. For example, the computer device may convert each point of the representation into absolute values and then determine an area ofthe three-dimensional representation for which no values are available (e.g. for which no opaque values are available). The hole may be determined by comparing the points of a second three-dimensional representation to the points of a first three-dimensional representation. In particular, the computer device may be arranged to determine any differences between a first set of points that is formed by combining the locations and movement vectors ofthe points ofthe first three-dimensional representation to a second set of points based on the locations ofthe points ofthe second three-dimensional representation. Any discrepancies between these sets of points may be used to identify a hole in the first three-dimensional representation. In some embodiments, the hole is determined for a specific capture device. In particular, one or more angular brackets may be determined for a capture device in which points are present at the first time, but not at the intermediate time. This determination may be dependent on a movement vector of a point, where the computer device may: identify a first location associated with a point at a first time (e.g. an angular bracket associated with that point); identify that the point has moved away from the first location at an intermediate time; and determine whetherthere is any data available for the location covered by the angular bracket. In certain situations, a surface, e.g. a wall, behind the point may already be captured since this point is visible from another part of the viewing zone. In certain situations, the surface exposed by the movement of the point will not have been previously captured. In some embodiments, the computer device is arranged to determine that the hole is entirely surrounded by points captured by a specific capture device (e.g. the hole relates to a subset of angular brackets for a capture device that is entirely within a larger arrangement of angular brackets for which point data is available). In such a situation, it is unlikely that another capture device will have captured relevant point data. When the hole has been identified, in a second step 122, the computer device identifies points associated with the first location in a second three-dimensional representation (e.g. an immediately preceding three-dimensional representation and / or an immediately successive three-dimensional representation). In particular, the computer device may identify that a point in a first angular bracket associated with a capture device is moving. This may comprise identifying an angle of the point from the first capture device. The computer device may then identify that, at the intermediate time, the point has moved out of this original angular bracket. The computer device may then identify a point captured for the same angular bracket in another three-dimensional representation (e.g. an immediately successive three-dimensional representation). This point is likely to represent a surface that has been uncovered by the movement of the point. In a third step 123, the computer device processes the first three-dimensional representation based on the identified point of the second three dimensional representation. In particular, the computer device may include the identified point of the second three-dimensional representation in the first three-dimensional representation. Typically, the computer device adds a point to the first three-dimensional representation that has a location and / or an attribute that is the same as a location and / or attribute of the identified point. The added point may have a different motion vector to the identified point. In some embodiments, the location of the added point is determined based on the location and the movement vector of the identified point of the second three-dimensional representation (e.g. the movement vector may be used to predict a location of the point at the first time, where this predicted location is used as the location ofthe added point). As described above, in some embodiments a hole in the first three-dimensional representation is determined by comparing points of a second three-dimensional representation to the points of the first three-dimensional representation. For example, the hole may be determined by comparing the points of the second three-dimensional representation to predicted points of the first three-dimensional representation (e.g. where the predicted points indicate predicted locations of these points at the time of the second three-dimensional representation). More generally, the hole(s) may be determined by comparing the points of the first three-dimensional representation to the points of the second three-dimensional representation. Typically, the first three-dimensional representation is associated with a first time period (e.g. a period oft=0 to t=1) and the second three-dimensional representation is associated with a second time period (e.g. a period of t=1 to t=2). Therefore, it can be assumed that many (or all) of the surfaces in the second three-dimensional representation will become visible during the period covered by the first three-dimensional representation. Therefore, the first three-dimensional representation and the second three-dimensional representation may be compared to ensure that all of the locations covered by the points ofthe second three-dimensional representation are also covered by points ofthe first three-dimensional representation. Such a comparison may be carried out to determine a hole without considering any movement vectors of either three dimensional representation. A consideration of points without considering movement vectors (or of points without movement vectors) can be particularly useful to ensure that any stationary surfaces that become de-occluded due to the movement of objects are captured in the first three-dimensional representation. In some embodiments, the points of the second three-dimensional representation are compared to predicted points of the first three-dimensional representation (the predicted points being determined by combining initial locations of points in the first three-dimensional representation with movement vectors of these points). This method ensures that the first three-dimensional representation leads smoothly into the second three-dimensional representation so that there will be no discontinuity between a final frame rendered from the first three-dimensional representation and a first frame rendered from the second three-dimensional representation. The comparing of the first and second three-dimensional representations may comprise comparing the points captured by each capture device (e.g. stepping through each angular bracket for each capture device and identifying a hole in the first three-dimensional representation where there is a point in an angular bracket of the second three-dimensional representation and no point in a corresponding angular bracket of the first three-dimensional representation). However, such a method may fail to account for the capturing of different points at the same location by different capture devices (e.g. a first capture device may capture a point at a specific location in the first three-dimensional representation and a second capture device may capture a corresponding point at this specific location in the second three-dimensional representation). Therefore, typically, the method of comparing the points of the first three-dimensional representation and the second three-dimensional representation comprises: determining a coverage (e.g. an area ora volume covered by) of one or more points of the second three-dimensional representation; and determining one or more parts ofthis coverage (parts of a scene covered by the second three-dimensional representation) that are not covered by points of the first three-dimensional representation. Determining the coverage of the points of the second three-dimensional representation may involve converting the points of the second three-dimensional representation to, e.g. quads in, a global coordinate system. This may involve converting points captured by different capture devices into a shared, global, coordinate system. Determining the one or more parts of the coverage that are not covered by points of the first three-dimensional representation may comprise determining a coverage of one or more points of the first three-dimensional representation and identifying any discrepancies between the coverages. Equally, this may comprise identifying for each part of the coverage of the second three-dimensional representation (e.g. for each angular bracket that forms a part of the coverage) whether the first three-dimensional representation contains a point that provides corresponding coverage. In an exemplary implementation, the computer device iterates through the capture devices and the angular brackets of these capture devices and identifies, for each angular bracket, whether this angular bracket is covered by a point of the second three-dimensional representation. If the angular bracket is covered by a point of the second three-dimensional representation, then the computer device determines whether the angular bracket is covered by a point of the first three-dimensional representation (which point may be captured by any capture device). This may involve generating a predicted final arrangement of the points of the first three-dimensional representation by combining the points of the first three-dimensional representation with movement vectors of the three-dimensional representation. It will be appreciated that other methods are possible to determine coverages of the representations, for example, the points of each three-dimensional representation may be converted to a global coordinate system and the computer device may then divide this global coordinate system into regions and step through each region to identify regions that are covered by a point of the second three-dimensional representation, butthat are not covered by any point of the first three-dimensional representation. These regions may be identified to be holes in the first three-dimensional representation. Any holes that have been determined in the three-dimensional representation can then be filled in using points from the second three-dimensional representation. This may comprise copying points from the second three-dimensional representation into the first three-dimensional representation. This may comprise defining a new point in the first three-dimensional representation based on an existing point in the second three-dimensional representation, where a location and / or a movement vector of the new point may be based on a movement vector of the existing point (e.g. the location of the new point may be determined assuming that the point has a constant velocity so that the location of the new point is determined using the inverse of the motion vector of the existing point and the new point is then given this motion vector of the existing point). Equally, the new point may be determined based on other points in the first three-dimensional representation (as is described further below with reference to Figures 19 - 20d). Where the hole is a part of a moving surface, it may be preferable to define the new point based on existing points in the first three-dimensional surface that surround the new point to ensure that the movement of the new point matches the movement of these surrounding points. Referring to Figure 16, the supplemented three-dimensional representation may comprise points for the initially occluded portion 2A, where these points come into view as the surface moves from the first position LS1 to the second position LS2. These points may be unused for the first frame F1 that is rendered based on the first three-dimensional representation R1, but may be used for the second frame F2 that is rendered based on this first three-dimensional representation. Referring to Figures 17a - 17c, holes may also appear if the surface expands (or rotates) due to the movement vectors. Referring to Figure 17a, the surface may expand from a first location LS1 at a first time to a second location LS3 at an intermediate time, where this expanded surface comprises the same number of points as the original surface (since the expansion is a result of points moving apart). In this regard, referring to Figures 17b and 17c, the expansion may result in the points moving apart at the intermediate time so that at the first time (shown in Figure 17b), the points cover a contiguous arrangement of angular brackets for a capture device and at the intermediate time (shown in Figure 17c), there are empty angular brackets between the points. In some embodiments, these holes may be filled in as described with reference to Figure 15. In this regard, if the surface expands during the time period covered by the first three-dimensional representation, then it is likely to be in its expanded form at the initial time of the second three-dimensional representation. Therefore, points from the second three-dimensional representation may be determined that would fill in the holes. Typically, the first three-dimensional representation needs to cover the entire time period associated with the first three-dimensional representation. That is, the points in the first three-dimensional representation need to be suitable for the first frame F1 that is associated with the first time and also the second frame F2 that is associated with the intermediate time. An issue with adding points from the second three-dimensional representation to fill in holes that form at the intermediate time, is that these points may not be allocatable during the first time. That is, an angular bracket AB1 that is empty at the intermediate time shown in Figure 17c may be filled at the first time shown in Figure 17b. Therefore, adding a point from the second three-dimensional representation that relates to this angular bracket may cause conflict with the point that already exists in this bracket in the first three-dimensional representation. In some embodiments, the point is nevertheless added to the first three-dimensional representation (where this can cause an overlap between points at the first time). In some embodiments, to avoid placing a plurality of points into a single angular bracket for a capture device, instead of adding a point from a further three-dimensional representation the computer device may be arranged to blend a plurality of points associated with a hole (e.g. to fill in the first angular bracket AB1 at the intermediate time based on the values of adjacent angular brackets). In some embodiments, the computer device is arranged to identify a plurality of points (e.g. a first point and a second point) that are adjacent a hole at the intermediate time and the computer device is arranged to determine a new point based on the plurality of points (e.g. to determine this new point at the time of generating an intermediate three-dimensional representation or at the time of rendering a two-dimensional image). For example, the new point may be located at the center of the identified points and may have a motion vector that is a combination of (e.g. an average of) the motion vectors of the identified points. Therefore, the new point may move into the hole as the identified points diverge between the first time and the intermediate time. Referring to Figure 18, there is described another method for processing the first three-dimensional representation to avoid holes, which method is suitable in situations with expanding objects. In a first step 131, the computer device identifies a first capture device associated with a first point (with this first point being associated with a hole, e.g. the first point fills the first angular bracket AB1 at the first time and then moves out of the first angular bracket at the intermediate time). In a second step 132, the computer device identifies a second capture device. In a third step 133, the computer device determines a second point associated with the second capture device, the second point being arranged to fill the hole associated with the first point at the intermediate time. For example, the second point may be co-located with the first point so that the second point appears as the first point moves out of the first angular bracket. The computer device then includes this second point in the first three-dimensional representation. The second capture device may, for example, be located adjacent the first capture device. By determining the second point based on a second capture device, this method avoids any problems that might be caused by having a plurality of points associated with a single capture device and a single angular bracket (e.g. where this may cause shimmering as these two points overlap each other). Typically, the computer device is arranged to determine a movement vector for the second point, which movement vector is based on the first point and / or based on a plurality of points associated with the hole. Therefore, the second point may be arranged to move into the location of the first angular bracket AB1 at the intermediate time (in this regard, the first angular bracket is associated with the first capture device, so the second point associated with the second capture device does not move into the first angular bracket but rather moves into a different angular bracket that is substantially co-located with the first angular bracket. The second capture device used to capture the second point may be determined based on a quality value associated with the second capture device and the second point, which quality value is a function of the distance of the surface from the second capture device and the angle of the surface from the second capture device (e.g. so that the quality value indicates a spread of sub-points associated with the second capture device capturing the second point, where a lower spread indicates a higher quality of capture). Typically, the first capture device is the device with the highest quality value for the surface associated with the points. The second capture device may be selected as the capture device with the second highest quality value for this surface. In some embodiments, the movement vectors can cause a change in the normal of a point between the first time and the intermediate time (e.g. to enable the point to rotate). In such embodiments, this rotation may cause holes to appear between points of a surface. These holes may again be filled as described above (e.g. by generating a new point that blends a surface between rotating points, which new point may be generated based on the points adjacent a hole and may have a normal that is a combination of the normals of these points adjacent the hole). Referring to Figures 20a and 20b, there is shown a further situation in which a first object 01 is moving behind a second object 02. At the initial time ofa first three-dimensional representation, as shown in Figure 26a, this second object occludes a surface of the first object. During a time period associated with the first three-dimensional representation, the first object O1 moves so that the surface is no longer occluded by the second object. This leads to a hole H1 at this surface. Since the first object O1 is moving during the time period associated with the first three-dimensional representation, it may not be possible to fill in the hole based on a second three-dimensional representation (e.g. because the first object O1 is differently located in the first and second three-dimensional representations, it is not possible to simply copy (unmodified) point data from the second three-dimensional representation into first three-dimensional representation). Therefore, referring to Figure 19, there is described a method of processing the first three-dimensional representation so as to determine a point forthe visible portion of the occluded surface 2 (this visible portion being visible for at least a part of a first time period associated with the first three-dimensional representation. In a first step 141, the computer device determines an edge ofa hole. As described previously, this may comprise identifying a location and / or an angular bracket that will be visible to a viewer but is not covered by any point of the first three-dimensional representation. For example, referring to Figure 20c, the computer device may identify points P1 and P2 that are points in the first three-dimensional representation that are at the edges of the hole H1, and the computer device may identify that there is no point data available to cover the space between these points P1 and P2. In a second step 142, the computer device determines an attribute value associated with (e.g. adjacent) the edge of the hole. In particular, for the one or more edges of the hole, the computer device may identify an attribute value of the closest point to the hole. In an example where the hole is associated with an empty angular bracket, the computer device may identify the attribute value of a point of the angular bracket adjacent this empty angular bracket. Referring again to Figure 20c, the computer device may determine attribute values (and / or movement vectors) associated with the points P1 and P2. In a third step 143, the computer device determines a new point (or new pixel value) that covers at least a portion of the hole (e.g. that fills an empty angular bracket), with an attribute value of the new point being associated with the attribute value determined in the second step 142. This new point is then included in the first three-dimensional representation so that it becomes visible as the first object 01 moves and thereby exposes the hole H1. Referring again to Figure 20c, the computer device may determine new points NP1, NP2 based on the existing points P1 and P2, where the attribute values of the new points may be determined based on (e.g. may be defined as being equal to the attribute values of the existing points and the movement vectors of the new points may be determined based on (e.g. may be defined as being equal to )the movement vectors of the existing points. Therefore, these new points move with the existing points in order to provide a consistent surface. Typically, the method of Figure 19 is repeated so as to propagate points from each edge of the hole H1 until the entirety of the hole is filled. For example, a first edge of the hole may be filled with new points based on the values of points adjacent that first edge and a second edge of the hole may be filled with new points based on the values of points adjacent that second edge. As the new points are propagated towards the centre of the hole, the computer device may determine new points based on a combination of the values of the points adjacent the first edge and the second edge of the hole. For example, the computer device may determine the new points so as to smoothly blend the hole based on the values of the points adjacent the first edge and the second edge of the hole. This process is shown in Figure 20d, where a first new point NP1 that is at a first edge of the hole H1 is determined based on an existing point P1 atthe edge of the second hole. The attribute value and movement vector of the new point NP1 is determined based on the attribute value and movement vector of the existing point P1. Similarly, a second new point NP2 is determined at the second edge of the hole based on an existing point P2 atthe second edge of the hole. This process is then repeated, with a third new point NP3 being determined based on the (attribute values and movement vector) of the first new point NP1 and a fourth new point NP4 being determined based on the (attribute values and movement vector) of the second new point NP2. The process iterates until the hole H1 is filled (e.g. until each angular bracket associated with the hole is filled with a new point). This method of propagating points in order to fill a hole provides a coherent surface in which the new points will move with the existing points to avoid any gaps forming between the points. In general, the method may comprise an iterative process in which, at each stage of the iteration, a new point at an edge of a hole is generated based on an existing point atthe edge of the hole. So in a first stage, a first new point NP1 and a second new point NP2 are generated based on an existing first point P1 and an existing second point P2; then, in a second stage, a third new point NP3 and a fourth new point NP4 are formed based on the first new point NP1 and the second new point NP2. In this regard, at the onset of the second stage, the first new point NP1 and the second new point NP2 have already been added to the first three-dimensional representation so that, at the onset of the second stage, these points may be considered to be existing points of the first three-dimensional representation. In some embodiments, the determination of the new points may comprise interpolating between (e.g. attribute values of and / or movement vectors of) a plurality of existing points. For example, where the first point P1 and the second point P2 have the same movement vectors (and the surface is moving without deforming), then each of the new points that are generated in order to fill the hole H1 may have the same movement vectors as these points P1 and P2. Where P1 and P2 have different movement vectors, the new points may be determined so as to have movement vectors that are based on the movement vectors of each of the existing points P1 and P2. In particular, the movement vector for each new point may be determined by interpolating between the movement vectors of the existing points P1 and P2 based on the relative distances of each new point from the existing points. For example, if the second existing point P2 is moving more quickly than the first existing point P1, then the new points may be determined to provide a speed gradient between these existing points. This ensures that the surface formed by the new points remains coherent during the entirety of the time period associated with the scene. Similarly, the attribute values of the new points may be interpolated based on the attribute values of the existing points. In a simple example, where the first edge of the hole is black and the second edge of the hole is white, the computer device may be arranged to determine one or more new points so as to fill the hole with a black to white colour gradient. Equally, the computer device may be arranged to determine one or more new points so as to provide a sharp edge where a surface changes from black to white. The locations and / or normals and / or movement vectors of the new point(s) may be determined based on the locations and / or normals and / or movement vectors point(s) adjacent the edge of the hole, where these attribute values may be blended during the formation of the new points. This method of propagating information from the edges of the hole towards the centre of the hole may be performed on a three-dimensional representation (e.g. to add new points to the three-dimensional representation in order to cover the hole). Equally, this method may be performed on a rendered two-dimensional image (e.g. to add new pixels to the image in order to cover a hole in the image). This method of propagating may also be used where the second object 02 is moving instead of, or as well as, the first object O1. While this method has been described in two dimensions, e.g. to describe the propagation of points towards the center of a hole from a first and second direction, it will be appreciated that in practice this method is typically performed in three dimensions so that the propagation of points may occur in additional directions. In some embodiments, the determination of the new points comprises obtaining one or more attribute values (e.g. colour values and / or normal values) from a second three-dimensional representation. In this regard, where the first object O1 is moving during the time period associated with the first three-dimensional representation, it may not be possible to simply copy points from the second three-dimensional representation into the first three-dimensional representation in order to cover a hole associated with the first object (since the points associated with the first object are in different locations in the first and second three-dimensional representations and might have different movement vectors in these representations). However, it may still be possible to obtain colour values, normals, etc. from the second three-dimensional representation in order to fill the hole. Therefore, the determination of the new points (that fill in the hole in the first three-dimensional representation) may comprise obtaining movement vectors for the new points based on existing points of the first three-dimensional representation and obtaining attribute values for the new points based on points of the second three-dimensional representation. Determining these points of the second three-dimensional representation may involve: identifying existing points of the first three-dimensional representation that are adjacent the hole; identifying corresponding points in the second three-dimensional representation; and identifying points of the second three-dimensional representation that cover the hole (e.g. by identifying points adjacent the corresponding points). For example, the computer device may identify points P1a and P2a in the second three-dimensional representation that correspond to the points P1 and P2 in the first three-dimensional representation. The computer device may then identify further points NP1a, NP2a adjacent these points P1a and P2a in the second three-dimensional representation and may then determine the attribute values for NP1 and NP2 based on the attribute values of these adjacent points NP1a and NP2a. Identifying the points P1a and P2a in the second three-dimensional representation may involve determining a final location of the points P1 and P2 ofthe first three-dimensional representation based on the initial locations and movement vectors of these points and then identifying points in the second three-dimensional representation at these final locations. In this regard, the initial locations of points of the second three-dimensional representation typically correspond to the final locations of points ofthe first three-dimensional representation (where this enables a continuous scene to be formed using a series of three-dimensional representations). Such a method of obtaining attribute values from the second three-dimensional representation ensures that attributes (e.g. colours) ofthe first object O1 are consistent throughout a scene. By obtaining the movement vectors for the new points from other points in the first three-dimensional representation, it can be ensured that the first object remains coherent throughout the time period associated with the first three-dimensional representation. Quantization of movement vectors Typically, the movement vectors are quantized prior to the storage and / or transmission of the three-dimensional representation and prior to the rendering of any two-dimensional images based on the three-dimensional representations. This quantization enables the movement vectors to be efficiently stored and to be efficiently signalled in a bitstream that contains the three-dimensional representation(s). This quantization can cause a difference in the predicted location of each point as compared to a prediction that uses the precise movement vectors (in particular, where an object is moving quickly this quantization may, for example, cause the actual predicted point to be in a different angular bracket than the ‘precise’ predicted point). Typically, the identification of a hole and / or the determination of a new point in the first three-dimensional representation to fill in the hole is performed using the quantized motion vectors. This ensures that the new point covers the hole when the two-dimensional images are rendered. Typically, the above-described methods are performed as a post-processing stage after the initial capture and storage of the three-dimensional representations. In these situations, the movement vectors are typically already quantized before the method is performed so that no further quantization is required. Formation of a bitstream The processing of a three-dimensional representation, and the rendering of two-dimensional images, that has been described above results in a three-dimensional representation and / or a two-dimensional image that may then be stored and / or transmitted by a computer device. In particular, the three-dimensional representation and / or the two-dimensional image may be encoded in a bitstream that is transmitted to another device and / or may be used to render one or more two-dimensional images that can be encoded in a bitstream that is transmitted to another device. This bitstream can then be decoded by this other device (e.g. the display device 17) in order to extract the three-dimensional representation(s) and / or the two-dimensional image(s) from the bitstream. The present disclosure envisages a bitstream that contains or is determined based on a three-dimensional representation that comprises one or more ‘new’ points, which new points are arranged to cover holes that might otherwise appear in the scene. Figure 21 shows a schematic of such a bitstream. Specifically, Figure 21 shows a bitstream that comprises a plurality of bits Bit-a to Bit-d. These bits signal one or more points of a three-dimensional representation and / or that defines one or more immersive two-dimensional images. This bitstream can be decoded by a decoding device and this may allow the three-dimensional representation and / or the two-dimensional images to be extracted from the bitstream. In some embodiments, the bitstream comprises one or more flags that indicate features of the bitstream and / or of the three-dimensional representations or two-dimensional images signalled by the bitstream. For example, the bitstream may comprise one or more flags that indicate: whether movement vectors are present in the three-dimensional representation; whether a viewing zone movement vector is present in the three-dimensional representation; whether (for each point), a point is arranged to move with the viewing zone (e.g. a cockpit flag); an accuracy of the movement vectors and / or a quantisation curve associated with the movement vectors; a time spacing of three-dimensional representations in the bitstream (e.g. so that a time period for each three-dimensional representation can be determined); a unit type of the movement vectors signalled in the bitstream; and whether new points have been included in a three-dimensional representation, e.g. due to de-occlusion occurring. These flags may be used to determine whether a de-occlusion process might have occurred or might need to occur and may affect the playback of the two-dimensional images that are rendered based on the three-dimensional representation. For example, if flags indicate that there are many fast moving objects in a scene and that processing has been performed to account for de-occlusion, then the scene may be rendered at a lower resolution than might occur otherwise to reduce the noticeability of any imperfections that might have been introduced by the de-occlusion process and the generation ofthe new points in the first three-dimensional representation. Alternatives and modifications It will be understood that the present invention has been described above purely by way of example, and modifications of detail can be made within the scope ofthe invention. The representation is typically arranged to provide an extended reality (XR) experience (e.g. a representation that is useable to render a XR video). The term extended reality (XR) covers each of virtual reality (VR), augmented reality (AR), and mixed reality (MR) and it will be appreciated that the disclosures herein are applicable to any of these technologies. The representation may be encoded into, and / or transmitted using, a bitstream, which bitstream typically comprises point data for one or more points ofthe three-dimensional representation. The point data may be compressed or encoded to form the bitstream. The bitstream may then be transmitted between devices before being decoded at a receiving device so that this receiving device can determine the point data and reform the three-dimensional representation (or form one or more two-dimensional images based on this three-dimensional representation). In particular, the encoder 13 may be arranged to encode (e.g. one or more points of) the three-dimensional representation in order to form the bitstream and the decoder 14 may be arranged to decode the bitstream to generate the one or more two-dimensional images. In some embodiments, the scene comprises a static scene; alternatively, in some embodiments the scene comprises a video and / or a moving (e.g. non-static) scene. That is, in some embodiments the scene comprises a static scene, such as a building, where a viewer is able to move through this scene, e.g. to view different rooms ofthe building, but where the scene itself does not change. In some embodiments, the scene comprises a moving scene, where elements of the scene vary in time even where the viewer remains stationary. It will be appreciated that typically the scene comprises both static and moving elements where, for example, non-static elements move in front of a static background. Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope ofthe claims.
Claims
1. A method of processing a three-dimensional representation of a scene, the method comprising: identifying a hole in the three-dimensional representation;determining a location of the hole;determining a new point, the new point being associated with the hole location; and adding the new point to the three-dimensional representation.
2. The method of claim 1, wherein the new point is determined based on a point of a further three-dimensional representation.
3. The method of any preceding claim, wherein the new point is determined based on existing points of the three-dimensional representation, preferably wherein the existing points are adjacent the hole.
4. The method of claim 3, wherein:a movement vector of the new point is determined based on existing points of the three-dimensional representation; andan attribute value of the new point is determined based on a point of a further three-dimensional representation.
5. The method of preceding claim, wherein the hole comprises a portion of the scene and / or a surface of the scene for which no attribute data is available.
6. The method of any preceding claim, wherein identifying the hole comprises identifying an area of the three-dimensional representation that is not associated with a point and / or that is not associated with an attribute value, preferably not associated with an opaque attribute value.
7. The method of any preceding claim, wherein identifying the hole comprises identifying a surface of the scene that is occluded from a viewer in a viewing zone of the three-dimensional representation ata first time and that is visible from the viewing zone of the scene at a second time, preferably wherein the first time and the second time are each within a time period associated with the three-dimensional representation.
8. The method of any preceding claim, wherein identifying the hole comprises comparing one or more points of the three-dimensional representation to one or more points of a further three-dimensional representation, preferably comprising determining the hole based on a discrepancy between the points of the three-dimensional representation and the points of the further three-dimensional representation.
9. The method of claim 8, comprising:determining one or more predicted points for the first three-dimensional representation, preferably determining the predicted points by combining locations and movement vectors of points of the three-dimensional representation; andcomparing the predicted points to the points of the further three-dimensional representation.
10. The method of claim 8 or 9, comprising determining a coverage of the points of the further three-dimensional representation, preferably comprising converting points of the further three-dimensional representation to a global coordinate system.
11. The method of any of claims 8 to 10, comprising determining a coverage of the points of further three-dimensional representation, preferably comprising converting points of the three-dimensional representation to a global coordinate system.
12. The method of claim 10 or 11, comprising comparing a coverage of the points of the further three-dimensional representation to a coverage of the points of the three-dimensional representation, preferably to a coverage of predicted points of the three-dimensional representation.
13. The method of any preceding claim, wherein the location is associated with an angle.
14. The method of any preceding claim, wherein each point of the three-dimensional representation is associated with a capture device, and the location is determined in dependence on a capture device, preferably wherein the location is associated with an angular bracket of a capture device, preferably wherein identifying the hole comprises identifying one or more angular brackets associated with the capture device for which no information is available at the second time, preferably wherein the angular brackets are surrounded by angular brackets for which information is available.
15. The method of any preceding claim, wherein identifying the hole comprises determining a range of surfaces that will be visible during a first time period associated with the three-dimensional representation.
16. The method of any preceding claim, wherein identifying the hole comprises identifying a movement vector associated with a moving point of the three-dimensional representation, the movement vector identifying a movement of the point during a first time period associated with the three-dimensional representation, preferably wherein determining the point of the furtherthree-dimensional representation comprises determining a point on a surface behind the moving point of the three-dimensional representation, preferably wherein:the point is a point that will be revealed during a time period associated with the three-dimensional representation due to the movement of the moving point; and / orthe method comprises determining a quantised movement vector, and determining the point on the surface based on the quantised movement vector.
17. The method of any preceding claim, wherein the further three-dimensional representation is an immediately preceding or an immediately successive three-dimensional representation.
18. The method of any preceding claim, comprising:identifying a location of the determined point of the furtherthree-dimensional representation; identifying a motion vector of the determined point of the furtherthree-dimensional representation; predicting a location of the determined point at the time of the three-dimensional representation based on the identified location and the identified motion vector; andadding the determined point to the three-dimensional representation based on the predicted location.
19. The method of any preceding claim, comprising:identifying a hole in the three-dimensional representation;determining a location of the hole;determining attribute values of one or more points adjacent the hole; andadding one or more points to the three-dimensional representation so as to fill the hole the attribute values of the points being defined based on the determined attribute values.
20. The method of claim 19, wherein:the points adjacent the hole are associated with a first capture device; and adding one or more points comprises adding one or more points associated with a second capture device; and / orthe one or more points are associated with a surface that is expanding; and adding the one or more points comprises adding one or more points so as to fill in gaps formed between the points as the surface expands; and / orthe one or more points are associated with a surface that is rotating; and adding the one or more points comprises adding one or more points so as to fill in gaps formed between the points as the surface rotates; and / or21. The method of any of claim 19 or 20, comprising:identifying movement vectors associated with the one or more points adjacent the hole; and determining movement vectors for the one or more points based on the identified movement vectors, preferably by averaging the identified movement vectors; and / oridentifying normals associated with the one or more points adjacent the hole; and determining normals for the one or more points based on the identified movement vectors, preferably by averaging the identified normals.
22. The method of any preceding claim, wherein the three-dimensional representation is associated with a viewing zone, the viewing zone comprising a subset of the scene and / or the viewing zone enabling a user to move through a subset of the scene, preferably wherein:the user is able to move within the viewing zone with six degrees of freedom (6DoF); and / orthe viewing zone has a volume of less than 50% of the volume of the scene, less than 20% of the volume of the scene, and / or less than 10% of the volume of the scene; and / orthe viewing zone has, or is associated with, a volume, preferably a real-world volume, of less than five cubic metres (5m3), less than one cubic metre (1 m3), less than one-tenth of a cubic metre (0.1m3) and / or less than one-hundredth of a cubic metre (0.01m3).
23. An apparatus for determining a point of a three-dimensional representation of a scene, the apparatus comprising:means for identifying a hole in the three-dimensional representation;means for determining a location of the hole;means for determining a new point, the new point being associated with the hole location; andmeans for adding the new point to the three-dimensional representation.
24. A bitstream comprising the three-dimensional representation and / or the two-dimensional image of any preceding claim.
25. A bitstream associated with a three-dimensional representation, the bitstream comprising one or more flags indicating:whether movement vectors are present in the three-dimensional representation;whether a viewing zone movement vector is present in the three-dimensional representation;whether (for each point), a point is arranged to move with the viewing zone;an accuracy of the movement vectors and / or a quantisation curve associated with the movement vectors;a time spacing of three-dimensional representations in the bitstream;a unit type of the movement vectors signalled in the bitstream; andwhether new points have been included in a three-dimensional representation due to de-occlusion occurring.
Citation Information
Patent Citations
Point cloud completion method and electronic equipment
CN118037601A
Three-dimensional imaging system point cloud hole completion system and method
CN118052741A
Augmented point cloud for a visualization system and method
US20170287216A1
Method and System for Completing Point Clouds Using Planar Segments
US20180218510A1
Method and apparatus for reconstructing a point cloud representing a 3D object
US20200258202A1