Processing a point of a three-dimensional representation

By denoising three-dimensional representations through threshold-based point processing, the method addresses the challenge of large file sizes, enhancing efficiency in processing and storage for virtual reality environments.

GB2642548APending Publication Date: 2026-01-14V NOVA INT LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
GB2024010601
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-09
Filing Date
2024-07-19
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Three-dimensional representations, such as point clouds, require substantial processing power and storage space, necessitating a reduction in file size for efficient handling and transmission.

Method used

A method of denoising three-dimensional representations by identifying and processing points based on threshold differences in location and attribute values, forming time-ordered series, and applying denoising operations to reduce redundancy.

Benefits of technology

Reduces the file size of three-dimensional representations while maintaining quality, enabling efficient processing and storage, suitable for virtual reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method and apparatus for processing three-dimensional representations of a scene comprise: identifying a first point in a first three-dimensional representation 21, the first point having a first lo
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Disclosure The present disclosure relates to methods, systems, and apparatuses for processing a point of a three-dimensional representation. In particular, the disclosure relates to methods of denoising a series of points of a three-dimensional representation, in particular to implement the denoising of a time-ordered series of points. Background to the Disclosure Three-dimensional representations of environments (e.g. point clouds) are used in many contexts, including for the generation of virtual reality videos. In such a context, a three-dimensional representation typically comprises a plurality of points that each have a location and a colour value for each of a left eye and a right eye of a user. A two-dimensional image can then be generated for each eye of a user based upon these locations and colour values to provide a virtual reality experience to this user. Typically, substantial processing power is required to form such a three-dimensional representation, and then once the representation has been formed a large amount of space is required to store the representation. Therefore, it is desirable to provide technologies which enable reductions in the file size of three-dimensional representations. Summary of the Disclosure According to an aspect of the present disclosure, there is described: a method of processing a three-dimensional representation of a scene, the method comprising: identifying a first point in a first three-dimensional representation, the first point having a first location; identifying a second point in a second three-dimensional representation, the second point having a second location; determining that a difference between the first location and the second location is below a threshold difference; and processing one or more of the points in dependence on the difference being below the threshold difference (e.g. to provide a processed three-dimensional representation including the processed points). Preferably, processing the points comprises forming a series of points, the series comprising the first point and the second point, preferably comprising forming a time ordered series of points. Preferably, processing the points comprises modifying an attribute value of one or more of the points. Preferably, processing the points comprises modifying a location of one or more of the points. Preferably, processing the points comprises performing a denoising operation on the series of points. Preferably, each point is associated with a capture device (e.g. used to capture that point). Preferably, identifying the first point comprises identifying a first point associated with a first capture device; and identifying the second point comprises identifying a second point associated with the (e.g. same) first capture device. Preferably, the method comprises identifying a third point associated with a second capture device, the third point having a third location; identifying a fourth point associated with the first capture device, the fourth point having a fourth location; determining that a difference between the third location and the fourth location is below a further threshold difference (e.g. the threshold distance); and processing one or more of the third point and the fourth point in dependence on the difference being below the further threshold difference. Preferably, the location of each point defines a distance of that point from a capture device. Preferably, the location of each point is defined based on: a capture device; a distance of the respective point from the capture device; and one or more angles of the respective point from the capture device. Preferably, the threshold difference is associated with a threshold distance difference. Preferably, the threshold difference is dependent on the distance of the first point and / or the second point from a viewing zone associated with the first three-dimensional representation and / or the second three-dimensional representation. Preferably, the method comprises determining a plurality of components of the location of the first point and a plurality of components of the location of the second point. Preferably, the components comprise a distance component and an angular component; and processing one or more of the points in dependence on a difference between each component being below a respective threshold difference (e.g. a distance difference being below a distance threshold difference and an angular difference being below an angular threshold difference). Preferably, the three-dimensional representations of a scene are associated with a plurality of (e.g. consecutive) frames in a video. Preferably, each of the three-dimensional representations of the scene comprise a plurality of points, wherein each point is associated with a location and at least one attribute. Preferably, the threshold difference depends on the attribute values of the first point and / or the second point, preferably an average value of the first point and the second point. Preferably, the threshold difference is dependent on a user input. Preferably, the threshold difference is dependent on a standard deviation of the locations of further points in the series. Preferably, determining that a difference between the first location and the second location is below a threshold difference further comprises: determining a probability that the first point and second point have the same location based on one or more of: the value of one or more attributes of the first point and the second point; and a component of the location of the first point and the second point. Preferably, the threshold difference is selected from a plurality of threshold differences, preferably selected by a user. Preferably, the method further comprises forming a plurality of series of points in parallel, each series comprising a plurality of co-located points. Preferably, the method further comprises performing a denoising method on values of attributes of points within the series. Preferably, the method further comprises predicting an attribute value of a further point of the series based at least in part on the prediction of an attribute value of the first point and the second point. Preferably, the method further comprises predicting an attribute value of a further point of the series based on the attribute values of the first point and the second point. Preferably, the method further comprises predicting the attribute value of the further point of the series based on a value of the further point. Preferably, processing the one or more points comprises modifying an attribute value of a point of a series based on a predicted attribute value forthat point, preferably replacing an initial attribute value of the point with the predicted attribute value. Preferably, the method further comprises fitting a line to the series of points, preferably wherein the line is associated with the attribute values of the points. Preferably, the method further comprises compressing a plurality of three-dimensional representations of a scene. Preferably, the compressing comprises encoding a function corresponding to the values of the attributes of the points within the series. Preferably, the compressing comprises defining the attribute values for one or more points with reference to a delta value, wherein the delta value indicates a difference between a value of an attribute at a first point in the series from a value of an attribute at a second point in the series. Preferably, the compressing is performed simultaneously to forming the series. Preferably, the plurality of three-dimensional representations of a scene are a first time ordered group and wherein the method further comprises forming a second time ordered group by reversing the time ordering of the plurality of three-dimensional representations of a scene within the first time ordered group and repeating the method of processing on the second time ordered group. Preferably, the second time ordered group is formed simultaneously to the first time ordered group. Preferably, the method further comprises altering the location of a point in the time series. Preferably, the method further comprises determining a third point in the first three-dimensional representation and a fourth point in the second three-dimensional representation, and forming a second series of points including the third point and the second point in dependence on a difference in a location of the third point and the fourth point being below a threshold difference. Preferably, the method further comprises determining a fourth point in a fifth three-dimensional representation and a fifth point in a sixth three-dimensional representation, and forming a further series of points including the fifth point and the sixth point in dependence on a difference in a location of the fifth point and the sixth point being below a threshold difference. Preferably, the method comprises determining a plurality of series of points. Preferably, each series is associated with a single capture device (e.g. each point in each series is captured with the same capture device). Preferably, the method comprises: identifying a capture device; identifying a plurality of points associated with this capture device; and determining a series based on the (e.g. locations of the) identified points. Preferably, the series and the further series are collated into a single composite series. Preferably, the method is repeated to determine a plurality of pairs of points for forming a plurality of time ordered series. Preferably, the method further comprises determining a first ordered series of points from a set of three-dimensional representations by evaluating each representation while moving through the set of representations in a first direction; determining a second ordered series of points from the set of three-dimensional representations by evaluating each representation while moving through the set of representations in a second direction, the second direction being opposite to the first direction; and processing one or more points of the first ordered series and / or the second ordered series. Preferably, the method comprises: determining a first ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the first ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a first direction; determining a second ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the second ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a second direction, the second direction being opposite to the first direction; and processing one or more points of the first ordered series and / or the second ordered series. According to another aspect of the present disclosure, there is described a method of processing a three-dimensional representation of a scene, the method comprising: determining a first ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the first ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a first direction; determining a second ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the second ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a second direction, the second direction being opposite to the first direction; and processing one or more points of the first ordered series and / or the second ordered series (e.g. to provide a processed three-dimensional representation including the processed points). Preferably, each of the first ordered series and the second ordered series is associated with the same points of the three-dimensional representations, preferably wherein the second ordered series comprises a reversed version of the first ordered series. Preferably, the first ordered series is associated with a forwards pass through the set of three-dimensional representations; and / or the second ordered series is associated with a backwards pass through the set of three-dimensional representations. Preferably, the processing one or more points comprises modifying an attribute value of one or more points of the first ordered series and the second ordered series. Preferably, the method further comprises determining a modified value of an attribute for one or more points in the first ordered series and / or the second ordered series; preferably comprising: determining a modified series of attribute values for the points in the first ordered series based on the initial values of these points; and / or determining a modified series of attribute values for the points in the first ordered series based on the initial values of these points. Preferably, the method further comprises selecting a sub-series of points from within the series of points; and determining a modified attribute value for a further point in the series of points based on the attribute values of the points in the sub-series of points. Preferably, the method further comprises sliding a window along the series of points so as to select a plurality of sub-series of points from the series of points; and determining a modified attribute value for each sub-series so as to form a modified series of attribute values. Preferably, the method further comprises sliding the window along the series of points in a first direction so as to form a forwards modified series of attribute values; and sliding the window along the series of points in a second direction, the second direction being opposite the first direction, so as to form a backwards modified series of attribute values. Preferably, the method further comprises forming a combined modified series of points based on the forwards modified series and the backwards modified series, preferably wherein forming the combined modified series comprises: determining one or more attribute values for the combined modified series by averaging corresponding attribute values of the forwards modified series and the backwards modified series; determining one or more attribute values forthe combined modified series as being one ofthe values of the forwards modified series and the backwards modified series; and determining one or more attribute values forthe combined modified series as being one ofthe original attribute values ofthe points ofthe first ordered series and / or the second ordered series. Preferably, the method further comprises weighting the attribute values ofthe sub-series so as to determine the modified attribute values, preferably wherein the weighting is biased to favour points towards an end of the sub-series, more preferably to favour chronologically later points in the sub-series. Preferably, the method further comprises weighting the attribute values for one or more sub-series used to determine the forwards modified series so as to favour chronologically later points in said sub-series; and / or weighting the attribute values for one or more sub-series used to determine the backwards modified series so as to favour chronologically earlier points in said sub-series. Preferably, the method further comprises modifying the attribute values of one or more points in the first ordered series and / or the second ordered series based on one or more of: the modified attribute series; the forwards modified series; the backwards modified series; and the combined modified series; preferably, comprising replacing the attribute values of the one or more points in the first ordered series and / or the second ordered series with values from one or more of: the modified attribute series; the forwards modified series; the backwards modified series; and the combined modified series. Preferably, the three-dimensional representation is associated with a viewing zone, the viewing zone comprising a subset of the scene and / orthe viewing zone enabling a user to move through a subset of the scene, preferably wherein the user is able to move within the viewing zone with six degrees of freedom (6DoF). Preferably, the viewing zone has a volume of less than 50% of the volume of the scene, less than 20% of the volume of the scene, and / or less than 10% of the volume of the scene; and / orthe viewing zone has, or is associated with, a volume, preferably a real-world volume, of less than five cubic metres (5m3), less than one cubic metre (1m3), less than one-tenth of a cubic metre (0.1 m3) and / or less than one-hundredth of a cubic metre (0.01m3). Preferably, each three-dimensional representation comprises a point cloud. Preferably, the method further comprises storing a processed three-dimensional representation, the processed three-dimensional representation comprising one or more processed points. Preferably, the method further comprises generating an image and / or a video based on the three-processed dimensional representation. Preferably, the method further comprises forming one or more two-dimensional representations of the scene based on the three-dimensional representation, preferably comprising forming a two-dimensional representation for each eye of a viewer. Preferably, each point is associated with one or more of: a location; an attribute; a transparency; a colour; and a size. Preferably, each point is associated with an attribute for a right eye and an attribute for a left eye. Preferably, the scene comprises one or more of: an extended reality (XR) scene; a virtual reality (VR) scene; an augmented reality (AR) scene; and a mixed reality (MR) scene. Preferably, the method further comprises forming a bitstream comprising one or more of the processed points. According to another aspect of the present disclosure, there is described a system for carrying out the method of any preceding claim, the system comprising one or more of: a processor; a communication interface; and a display. According to another aspect of the present disclosure, there is described an apparatus for processing a three-dimensional representation of a scene, the apparatus comprising: means for (e.g. a processor for) identifying a first point in a first three-dimensional representation, the first point having a first location; means for (e.g. a processor for) identifying a second point in a second three-dimensional representation, the second point having a second location; means for (e.g. a processor for) determining that a difference between the first location and the second location is below a threshold difference; and means for (e.g. a processor for) processing one or more of the points in dependence on the difference being below the threshold difference. According to another aspect of the present disclosure, there is described an apparatus for processing a three-dimensional representation of a scene, the apparatus comprising: means for (e.g. a processor for) identifying a first point in a first three-dimensional representation, the first point having a first location; means for (e.g. a processor for) identifying a second point in a second three-dimensional representation, the second point having a second location; means for (e.g. a processor for) determining that a difference between the first location and the second location is below a threshold difference; and means for (e.g. a processor for) processing one or more of the points in dependence on the difference being below the threshold difference. According to another aspect of the present disclosure, there is described an apparatus for processing a three-dimensional representation of a scene, the apparatus comprising: means for (e.g. a processor for) determining a first ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the first ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a first direction; means for (e.g. a processor for) determining a second ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the second ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a second direction, the second direction being opposite to the first direction; and means for (e.g. a processor for) processing one or more points of the first ordered series and / or the second ordered series. According to another aspect of the present disclosure, there is described a bitstream comprising one or more points processed using the aforesaid method. According to another aspect of the present disclosure, there is described an apparatus (e.g. an encoder) for forming and / or encoding the aforesaid bitstream. According to another aspect of the present disclosure, there is described an apparatus (e.g. a decoder) for receiving and / or decoding the aforesaid bitstream. Any feature in one aspect of the disclosure may be applied to other aspects of the invention, in any appropriate combination. In particular, method aspects may be applied to apparatus aspects, and vice versa. Furthermore, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be construed accordingly. Any apparatus feature as described herein may also be provided as a method feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure, such as a suitably programmed processor and associated memory. It should also be appreciated that particular combinations of the various features described and defined in any aspects of the disclosure can be implemented and / or supplied and / or used independently. The disclosure also provides a computer program and a computer program product comprising software code adapted, when executed on a data processing apparatus, to perform any of the methods described herein, including any or all of their component steps. The disclosure also provides a computer program and a computer program product comprising software code which, when executed on a data processing apparatus, comprises any of the apparatus features described herein. The disclosure also provides a computer program and a computer program product having an operating system which supports a computer program for carrying out any of the methods described herein and / or for embodying any of the apparatus features described herein. The disclosure also provides a computer readable medium having stored thereon the computer program as aforesaid. The disclosure also provides a signal carrying the computer program as aforesaid, and a method of transmitting such a signal. The disclosure extends to methods and / or apparatus substantially as herein described with reference to the accompanying drawings. The disclosure will now be described, by way of example, with reference to the accompanying drawings. Description of the Drawings Figure 1 shows a system for generating a sequence of images. Figure 2 shows a computer device on which components of the system of Figure 1 may be implemented. Figure 3 shows a method of determining a three-dimensional representation of a scene. Figure 4 shows a scene comprising a viewing zone. Figures 5a and 5b show arrangements of capture devices for determining points of the three-dimensional representation. Figure 6 shows a method of processing a point of a three-dimensional representation. Figure 7 shows a graph of a time ordered series of points. Figure 8a and 8b show a method of denoising a time ordered series of points. Figure 9 shows a graph of an example attribute for a time ordered series. Description of the Preferred Embodiments Referring to Figure 1, there is shown a system for generating a sequence of images. This system can be used to generate, and then display, a representation of an environment, which may comprise a VR environment (or an XR environment). The system comprises an image generator 11, an encoder 12, a transmitter 13, a network 14, a receiver 15, a decoder 16 and a display device 17. These components may each be implemented on separate apparatuses. Equally, various combinations of these components may be implemented on a shared apparatus; for example, the image generator 11, the encoder 12, and the transmitter 13 may all be part of a single image data generation device. Similarly, the receiver 15, the decoder 16, and the display device 17 may all be a part of a single image rendering device. Typically, the system comprises at least one encoding computer device (e.g. a server of a content provider) and at least one rendering computer device (e.g. a VR headset). Referring to Figure 2, each of the components, and in particular the image generator 11, the encoder 12, the transmitter 13, the receiver 15, the decoder 16 and the display device 17 is typically implemented on a computer device 20, where, as described above, a plurality of these components may be implemented on a shared computer device. Each computer device comprises one or more of: a processor 21 for executing instructions (e.g. so as to perform one or more of the steps of the various methods described below), a communication interface 22 for facilitating communication between computer devices (e.g. an ethernet interface, a Bluetooth® interface, or a universal serial bus (UBS) interface, a memory 23 and / or storage 24 for storing information and instructions (e.g. a random access memory (RAM), a read only memory (ROM), a hard drive disk (HDD) a solid state drive (SSD), and / or a flash memory, and a user interface 25 (e.g. a display, a mouse, and / or a keyboard) for enabling a user to interact with the computer device. These components may be coupled to one another by a bus 25 of the computer device. The computer device 20 may comprise further (or fewer) components. In particular, the computer device (e.g. the display device 17) may comprise one or more sensors, such as an accelerometer, a GPS sensor, or a light sensor. These sensors typically enable the computer device to identify an environmental condition and / or an action of wearer of the display device. Turning back to Figure 1, the image generator 11 is configured to generate a sequence of image data (e.g. a sequence of image frames) to enable the display device 17 to use this image data to display a plurality of images. The image data may comprise one or more digital objects and the image data may be generated or encoded in any format. For example, the image data may comprise point cloud data, where each point has a 3D position and one or more attributes. These attributes may, for example, include, a surface colour, a transparency value, an object size and a surface normal direction. Each attribute may have a value chosen from a continuous range or may have a value chosen from a discrete set. The image data enables the later rendering of images. This image data may enable a direct rendering (e.g. the image data may directly represent an image). Equally, the image data may require further processing in order to enable rendering. For example, the image data may comprise three-dimensional point cloud data, where rendering a two-dimensional image using this data requires processing based on a viewpoint of this two-dimensional image. The image data may comprise depth map data, where one or more pixels or objects in the image is associated with a depth that is specified by the depth map data. The depth map data may be provided as a depth map layer, separate from an image layer. In some contexts, such as MPEG Immersive Video (MIV), the image layer may instead be described as a texture layer. Similarly, in some contexts, the depth map layer may instead be described as a geometry layer. The image data may include a predicted display window location. The predicted display window location may indicate a portion of an image that is likely to be displayed by the display device 17. The predicted display window location may be based on a viewing position (such as a virtual position and / or orientation of the user in a 3D environment) of the user, where this viewing position may be obtained from the display device. The predicted display window location may be defined using one or more coordinates. For example, the predicted display window location may be defined using the coordinates of a corner or center of a predicted display window, and may be defined using a size of the predicted display window. The predicted display window location may be encoded as part of metadata included with the frame. The image data for each image (e.g. each frame) may include further information, which may be provided as a part of an image, e.g. as part of the point cloud data, or as separate layers. In particular, the image data may include audio information or haptic feedback information indicating audio or haptics which can accompany displayed visual data. An audio layer or haptic layer may accompany each image, and may be omitted for images where no accompanying audio or haptics are required. Similarly, the image data may comprise interactivity information, where the image data may contain or indicate elements with which a user can interact. The interactivity information may, for example, define a behaviour of an element, where a user is able to interact with the element based on this behaviour. The behaviour typically defines a change in an element that occurs as a result of a user interaction where this change may comprise a change in the attributes of the element or in the rendering of the element. As an example, where an image contains a target element, the target element may be arranged to disappear when a user interacts with this element, or to provide feedback indicating that the user has interacted with the target. This interactivity data may be provided as part of, or separately to, the image data. The image data may indicate, or may be combinable with, a state of the virtual environment, a position of a user, or a viewing direction of the user. Here, the position and viewing direction may be physical properties of the user in the real-world, or position and viewing direction may also be purely virtual, for example being controlled using a handheld controller. The image generator 11 may, for example, obtain information from the display device 17 that indicates the position, viewing direction, or motion of the user. Equally, the image generator may generate image data such that it can later be combined with this position, viewing direction, or motion, where the image generator may generate a full scene which is only partially viewed by a user depending on the position of that user. In some cases, the generated image may be independent of user position and viewing direction. This type of image generation typically requires significant computer resources such as a powerful GPU, and may be implemented in a cloud service, or on a local but powerful computer. For example, a cloud service (such as a Cloud Rendering Service (CRN)) may reduce the cost per-user and thereby make the image frame generation more accessible to a wider range of users. Here “rendering” refers at least to an initial stage of rendering to generate an image. Further rendering may occur at the display device 17 based on the generated image to produce a final image which is displayed. The image generator 11 may, for example, comprise a rendering engine for initially rendering a virtual environment such as a game or a virtual meeting room. The encoder 12 is configured to encode frames to be transmitted to the display device 17. The encoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. In some embodiments, the image generator 11 may transmit raw, unencoded, data through the network 14. However, such transmission typically leads to a high file size and requires a high bandwidth so that it is typically desirable to encode the data prior to the transmission. The encoder 12 may encode the image data in a lossless manner or may encode the data a lossy manner. The encoder may apply inter-frame or intra-frame compression based on a currently-encoded frame and optionally one or more previously encoded frames. The encoder may be a multi-layer encoder, such as an low complexity enhancement video codec (LCEVC) enabled encoder. Where the generated frames comprise depth map data, the encoder 12 may perform layered encoding on each instance of image data (e.g. each frame) to generate an encoded frame comprising a base depth map layer and an enhancement depth map layer. Encoding a depth map in this way may improve compression. In some applications, such as HDR video, depth maps are desirably highly detailed with a bit depth of up to twelve or fourteen bits, which is a significant increase in the data to be transmitted. As a result, providing ways to improve compression of the depth map can make more realistic depth map-based displays viable when performing rendering or transmission of rendered data in real-time. Furthermore, this type of layered encoding makes it easy to drop (and then pick back up) one or more of the layers, which provides flexibility and tools for bandwidth management. Layered encoding is also helpful as the final decoder / user device (such as a user display device) can choose whether to process these extra layers. For example, in a non-layered approach, the best the end device (i.e. the receiver, decoder or display device associated with a user that will view the images) can do is determine that it does not have enough resources for a given quality (be it resolution, frame rate, inclusion of depth map) and then signal to the controller / renderer / encoder that it does not have enough resources. The controller then will send future images at a lower quality. In that alternative scenario, the end device still unfortunately has to process the higher quality data until the lower quality data arrives, if it can process the received images at all. In some of the described embodiments, this situation is improved upon because when / if the end device determines for example that it does not have the processing capabilities to handle the highest level of quality, then it can drop and / or choose not to process certain layers. The end device may also signal to the controller that it needs a lower level of quality, but in the meantime the end device can only process the number of layers that it can handle. Therefore, the end device can react to conditions much more quickly. In some cases, depth map data may be embedded in image data. In this case, the base depth map layer may be a base image layer with embedded depth map data, and the enhancement depth map layer may be an enhancement image layer with embedded depth map data. Alternatively, when the generated images comprise a depth map layer separate from an image layer and multi-layer encoding is applied, the encoded depth map layers may be separate from the encoded image layers. This has the advantage that the encoded depth map layers can be dropped under some conditions while still retaining image layers that can be displayed (albeit with a lower level of realism). For example, the encoded depth map layers can be dropped by a transmitter or encoder when available communication resources are reduced, or can be dropped by an end device which lacks the processing resources to handle the highest level of quality. Similarly, if some images comprise an audio base layer, a haptic feedback base layer, an audio enhancement layer or a haptic feedback enhancement layer, these can be processed or dropped flexibly. Again similarly, if some images comprise an interactivity data base layer or an interactivity enhancement layer these can be processed or dropped flexibly. For example, certain interactions may only be possible where a threshold bandwidth is available, where complex interactions (e.g. those enabling a conversation with a digital object) may be disabled before less complex interactions (e.g. changing a pixel colour) are disabled. Additionally or alternatively, where the image data comprises point cloud data, the encoder may apply a point cloud data encoding technique such as described in European patent application EP21386059.6, which is incorporated herein by reference. Such a point cloud encoder may act as a base encoder for a layered encoding technique such as LCEVC or VC-6. Notably LCEVC and VC-6 techniques encode and decode a layered signal, but are agnostic about the content type of data encoded in the signal. For example, the signal can include textures, video frames, geometry or depth data, meshes, point clouds, rendering attributes or physics engine attributes. The transmitter 13 may be any known type of transmitter for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The transmitter 13 may be configured to make decisions about how to transmit the image data, and / or may provide feedback to the encoder 12 or the image generator 11. For example, the transmitter may determine available communication resources (e.g. bandwidth) for transmitting image data, and may drop one or more layers from an encoded frame, or indicate to the image generator and / or encoder that image data should be generated and encoded with fewer layers, when insufficient bandwidth is available for transmission of all generated data. As specific examples, the transmitter may be configured to drop a depth map layer, an LCEVC enhancement layer, or a VC-6 enhancement layer from a frame when insufficient communication resources are available. The network 14 provides a channel for communication between the transmitter 13 and the receiver 15, and may be any known type of network such as a WAN or LAN or a wireless Wi-Fi or Bluetooth network. The network may further be a composite of several networks of different types. Many users only have access to a network with a bandwidth of 30MBps which can lead to latency jitter when streaming. The required bandwidth and the observed latency can be reduced by means of tactics such as forward-looking rendering and last-millisecond reprojection, which are enabled by improved compression. The receiver 15 may be any known type of receiver for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The decoder 16 is configured to receive and decode an encoded frame. The decoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. The display device 17 may for example be a television screen or a VR headset. The timing of the display may be linked to a configured frame rate, such that the display device may wait before displaying the image. The display device may be configured to perform warping, that is, to obtain a final display window location, adjust a warpable image to obtain a final image corresponding to a final viewing direction of the user, and display the final image. In this regard, the image data is typically arranged to provide a warpable image for which a portion of the image that is displayed at the display device 17 is dependent on a position or orientation of a viewer. The warpable image may then be rendered before a most up to date viewing direction of the user is known. The warpable image may be transmitted to the display device, or the warpable image may be transmitted to a rendering node which is near to the display device, and the display device or rendering node may perform time warping to generate a displayed image portion based on the warpable image and the most up to date viewing direction of the user. As mentioned above, a single device may provide a plurality of the described components. For example, a first rendering node may comprise the image generator 11, encoder 12 and transmitter 13. Additional similar rendering nodes may be included in the system, and may work together to generate the sequence of frames. In one case, multiple rendering nodes may each provide separate image data to an image data assembling node; for example, each rendering node may provide a part of a sequence of frames to a frame assembling node. For example, the receiver 15, decoder 16 or display device 17 may be configured to assemble parts of image data from multiple sources to generate a sequence of images for display on the display device. Alternatively, the image data assembling node may be separate from the receiver 15, decoder 16 and display device 17. Additionally or alternatively, multiple rendering nodes may be chained. In other words, successive rendering nodes may add to a sequence of image data as it passes from rendering node to rendering node, and eventually a complete sequence of image data is then provided to the receiver 15. Furthermore, each rendering node may obtain components of a render from multiple upstream rendering nodes and / or distribute components of a render to multiple downstream rendering nodes. A chain of rendering nodes may be useful for performing different rendering tasks that require different quantities of processing resources, or different frame rates. For example, a company may provide distributed processing in the form of a centralised hub which has abundant processing resources but is distant from users, and peripheral locations which have more scarce processing resources but are closer to users. Expensive but fairly static rendering features such as background lighting or environmental impact on sound may be generated at the central hub (for example using ray tracing), while features that require fewer resources but faster responses or higher frame rates may be generated closer to the user. In other words, the more responsive a rendering feature needs to be, the lower latency it needs between the rendering node which generates the feature and the user display and, in a chain of rendering nodes, the node which generates each rendering feature can be chosen based on a required maximum latency of that feature. On the other hand, if it is expensive to generate a rendering feature, then it may be preferable to generate the feature less frequently and with a higher maximum latency. For example, a static, high-quality background feature may be generated early in the chain of rendering nodes and a dynamic, but potentially lower-quality, foreground feature may be generated later in the chain of rendering nodes, closer to the user device. Here, environmental impact on sound means, for example, a set of surfaces may be constructed where each surface has different sound reflection and absorption properties depending upon material and shape. The frame rates may be matched by creating multiple frames with features generated at the lower frame rate, and combining them with the frames with features generated at the higher frame rate. In a nonlimiting embodiment, a preliminary rendering generates volumetric object data including motion vectors at a first (lowest) frame rate, then produces 2D rendered frames plus depth information for a specific user at a second (higher) frame rate, then transmits video plus depth data to the user device, which produces final frames for display via space warping (depth-based reprojections) at a third (highest) frame rate. One or more of these steps may be performed in combination with the other described embodiments. The viewing position of the user may change as additional rendering tasks are performed at different rendering nodes in the chain. Each or any rendering node may obtain an updated viewing position before performing its respective rendering task. Additionally, the system may simultaneously generate multiple sequences of image data for different respective users or different respective display devices. For example, in the context of a VR or AR experience, each user or display device may view a different 3D environment, or may view different parts of a same 3D environment. When using a chain of rendering nodes, each node may serve multiple users or just one user. For example, a starting rendering node (e.g. at a centralised hub) may serve a large group of users. For example, the group of users may be viewing nearby parts of a same 3D environment. In this case, the starting node may render a wide zone of view (“field of view”) which is relevant for all users in the large group. The starting node may send this wide field of view to a first middle rendering node which renders additional aspects of the 3D environment. These additional aspects may for example be aspects which require less processing power to render, or may be aspects which are specific to individual users of the group. Additionally, the middle rendering node may render features in a smaller field ofviewthan the starting node - this smaller field of view may be relevant to each user rather than the group of users. The first middle rendering node may additionally only serve a smaller number of users (e.g. half of the large group of users), with the remaining users being served by a second middle rendering node which also receives the wide field of view from the starting node. The middle rendering node(s) may then send sequences of second partially or fully rendered frames to an end device for each user. The end device may perform further processes such as warping or focal distance adjustments, optionally using depth map data. Preferably, each rendering node encodes the partially or fully rendered frames before transmitting them on to a next rendering node or to the receiver 15. This means that the required communication resources can be reduced when the rendering nodes are separated by one or more networks, or more generally are implemented in a distributed system such as a cloud. However, each rendering node in a chain is encoding a different partially or fully rendered frame, with different data. Therefore, it may be advantageous for different rendering nodes to use different rendering formats and / or encoding formats. For example, the output from a first rendering node may be point cloud data which logically describes a 3D scene. This point cloud data can be encoded using the techniques of EP21386059.6. A second rendering node may then operate on the point cloud data to generate image data that is more readily displayed by a generic display device, without requiring the display device to model the 3D environment. This image data may be encoded using video coding techniques. The chaining of rendering nodes may be extended to arbitrary tree structures, where a rendering node obtains partially rendered frames from more than one preceding rendering node, and generates further partially or fully rendered frames based on the multiple obtained sequences of partially rendered frames. For example, a content rendering network (CRN) comprising numerous rendering nodes may be used to serve a volumetric event to a large number of same-time users, such as users participating in a shared virtual environment. Rendering the same event for each user is far more expensive in terms of computation time and power consumption than rendering the volumetric effect once and performing the rendering equivalent of multicasting the volumetric effect for multiple users. For example, each user may have a second rendering node (such as a VR headset), and the network may comprise a central first rendering node. The first rendering node may render the volumetric event, and distribute partially rendered frames depicting the volumetric event to the different second rendering nodes. The second rendering node for each user may then integrate the partially rendered frames depicting the volumetric event into a view of the virtual environment which is currently being shown to each user, based on parameters such as the user’s virtual position. The receiver 15, decoder 16 and display device 17 may be consolidated into a single device, or may be separated into two or more devices. For example, some VR headset systems comprise a base unit and a headset unit which communicate with each other. The receiver 15 and decoder 16 may be incorporated into such a base unit. In some embodiments, the network 14 may be omitted. For example, a home display system may comprise a base unit configured as an image source, and a portable display unit comprising the display device 17. In the event that the decoder 16 orthe display device 17 does not or cannot handle one or more layers, the receiver 15 or another transmitter associated with the decoder or display device may send a corresponding layer drop indication back through the network 14. The layer drop indication may be received by each rendering node. A rendering node which generates partially or fully rendered frames for that specific decoder or display device may cease generating the dropped layer. On the other hand, a rendering node which generates partially or fully rendered frames for multiple end devices may disregard a layer drop indication received from one end device (as the dropped layer is still needed for other devices). Alternatively, rendering nodes which serve multiple end devices may record received layer drop indications, and may cease generating the dropped layer only when all end devices served by the rendering node indicate that the layer is to be dropped. In preferred examples, the encoders or decoders are part of a tier-based hierarchical coding scheme or format. Hierarchical coding enables frames to be communicated with higher resolution and / or higher frame rate than is possible in single-tier coding schemes. In hierarchical coding, one or more enhancement layers is communicated with base data, where the enhancement layers can be used to up-sample the base data at the decoder, for example providing up-sampling in a spatial or temporal dimension. When combined with equivalent down-sampling of the original frames and generation of the enhancement layer at an encoder, hierarchical coding can overall provide lossless compression of data, with higher resolution and / or higher frame rate for a given transmission bit rate. Examples of a tier-based hierarchical coding scheme include LCEVC: MPEG-5 Part 2 LCEVC (“Low Complexity Enhancement Video Coding”) and VC-6: SMPTE VC-6 ST-2117, the former being described in PCT / GB2020 / 050695, published as WO 2020 / 188273, (and the associated standard document) and the latter being described in PCT / GB2018 / 053552, published as WO 2019 / 111010, (and the associated standard document), all of which are incorporated by reference herein. However, the concepts illustrated herein need not be limited to these specific hierarchical coding schemes. A further example is described in WO2018 / 046940, which is incorporated by reference herein. In this example, a set of residuals are encoded relative to the residuals stored in a temporal buffer. LCEVC (Low-Complexity Enhancement Video Coding) is a standardised coding method set out in standard specification documents including the Text of ISO / IEC 23094-2 Ed 1 Low Complexity Enhancement Video Coding published in November 2021, which is incorporated by reference herein. The system describes above is suitable for generating and presenting a representation of a scene, where this scene displays media content to a user. The scene typically comprises an environment, where the user is able to move (e.g. to move their head or to turn their head) to look around the environment and / or to move around the environment. For example, the scene may be a scene of a room in a building, where the user is able to move around the room (e.g. by moving in the real-world and / or by providing an input to a user interface) in order to inspect various parts ofthe room. Typically, the scene is a XR (e.g. a VR) scene, where the user is able to move about the scene in three degrees of freedom (3DoF) or six degrees of freedom (6DoF) so as to experience the scene. As has been described with reference to Figure 1, the image generator 11 may be arranged to determine point cloud data, where each point ofthe point cloud has a 3D position and one or more attributes. More generally, the image generator (or another component) is arranged to determine a three-dimensional representation of a scene, where this three-dimensional representation is thereafter used to generate two-dimensional images that are presented to a user at the display device 17. Referring to Figure 3, there is described a method of determining (an attribute for) a point of such a three-dimensional representation. The method comprises determining the attribute using a capture device, such as a camera ora scanner. The scene may comprise a real scene, in which attribute values are captured using a camera, or a virtual scene (e.g. a three-dimensional model of a scene), in which attribute values are captured using a virtual scanner. Where this disclosure describes 'determining a point’ it will be understood that this generally refers to determining a point that has a location and an attribute value, where determining the point comprises determining the attribute value and / or storing a point that comprises at least an attribute value and a location value (these values may be indirect values, e.g. where the location is identified relative to another point). Once a plurality of points have been captured, these points can be stored as a three-dimensional representation (e.g. a point cloud) so as to enable the reconstruction ofthe three-dimensional scene based on this representation. Typically, the three-dimensional representation is associated with a plurality of images or frames; for example, the three-dimensional representation may be associated with a video. In such embodiments, determining a point may involve determining a point of a four-dimensional representation, which has three spatial dimensions and a time dimension. The location and / or attributes of such a point may be dependent on time (e.g. the point may have an initial location as well as a speed and / or acceleration so that the location of the point can be determined at a plurality of different times based on these variables). In such embodiments, the method may be considered to comprise determining a plurality of three-dimensional representations, where each frame of a video is dependent on a different three-dimensional representation. In such a situation, there is typically a dependency between the plurality of three-dimensional representations; for example, the points of a second three-dimensional representation may be determined / encoded based on a preceding first three-dimensional representation. Therefore, this situation could be considered to involve either of: determining a four-dimensional representation; and determining a plurality of three-dimensional representations. Typically, the scene comprises a simulated scene that exists only on a computer. Such a scene may, for example, be generated using software such as the Maya software produced by Autodesk®. The attributes determined using the methods described herein may then depend on virtual objects located within the scene as well as a virtual lighting arrangement used in the scene. In a first step 11, a computer device initiates a capture process for a capture device, the capture process being initiated with an initial azimuth angle (e.g. of 0°) and an initial elevation angle (e.g. of 0°). In a second step 12, the computer device causes a point to be captured using the capture device at the current azimuth angle and current elevation angle. Capturing a point typically comprises assigning an attribute value to the point, which attribute value may, for example, be a color of the point and / or a transparency value of the point. Typically, the point has one or more color values associated with each of a left eye and a right eye of a viewer. Capturing the point may also comprise determining a normal value associated with the point, e.g. a normal of a surface on which the point lies. Typically, capturing the point further comprises determining a location of the point, e.g. by determining a distance of the point from the capture device (e.g. camera or scanner). In practice, determining the point may comprise sending a ‘ray’ from the capture device and then stepping through a computer model to determine which surface of the computer model is impacted by the ray. The color, transparency, and normal of this surface are then recorded alongside the distance of the surface from the capture device. In a third step, 13, the computer device determines whether a point has been captured for the capture device at each azimuth of a range of azimuths and in a fourth step 14, if points have not been captured at each azimuth, then the azimuth angle is incremented and the method returns to the second step 12 and another point is captured. The azimuth angle may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of azimuth angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. Once a point has been captured for each azimuth, in a fifth step 15, the computer device determines whether a point has been captured for the capture device at each elevation of a range of elevations and in a sixth step 16, if points have not been captured at each elevation, then the azimuth angle is reset to the initial value, elevation angle is incremented and the method returns to the second step 12 and another point is captured. The elevation angles may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of elevation angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. In a seventh step 17, once points have been captured for each azimuth angle and each elevation angle, the scanning process ends. This method enables a capture device to capture points at a range of elevation and azimuth angles. This point data is typically stored in a matrix. The point data may then be used to provide a representation of the scene to a user, e.g. the three-dimensional representation formed by the point data may be processed to produce two-dimensional images for each eye of a user, with these images then being shown to a user via the display device 17 to provide a virtual reality experience to the viewer. By using the captured data, a video can be provided to a viewer that enables the viewer to move their head to look around the scene (while remaining at the location of the capture device). It will be appreciated that the capture pattern (or scanning pattern) described with reference to Figure 3 is purely exemplary and that numerous capture patterns are possible. In general, the capture process for each capture device comprises capturing one or more points at one or more azimuth angles and / or one or more elevation angles. The 'points’ captured by the capture device are typically associated with a size, such as a height, a width, or a depth. That is, the points typically relate to two-dimensional pixels and / or three-dimensional voxels (or planes). In this regard, there is necessarily some space between the locations of adjacent points (since if the points had no width, then an infinite number of points would be required to capture points at each angle). The size provides points that depict a non-negligible area of the three-dimensional space so that a plurality of points can be fit together to provide a depiction of the scene to a viewer. The width and height of each point is typically dependent on the distance of that point from the capture device, where more distant points have a larger width / height. The width and height of each point is typically determined so that when each point is displayed, there is no space between adjacent points (indeed, there may be some overlap between points to ensure that no gaps appear between points). This height / width of each point can be determined at the time of capturing the points, or can be determined or defined after the capture of the points. Typically, the points comprise a size value, which is stored as a part of the point data. For example, the points may be stored with a width value and / or a height value. Typically, the minimum width and the minimum height of a point are set by the angle increment of the azimuth angle and the elevation angle respectively. The size may be then specified in terms of this angle increment and / or in terms of this minimum width / minimum height (e.g. as being a multiple of the angle increment). In some embodiments, the size value is stored as an index, which index relates to a known list of sizes (e.g. if the size may be any of 1x1, 2x1, 1x2, 2x2, pixels this may be specified by using 3 bits and a list that relates each combination of bits to a size). By capturing points at a plurality of azimuth angles and elevation angles, e.g. using the method described with reference to Figure 3, it is possible to provide a three-dimensional representation of the scene that can later be used to enable a viewer to view the scene from a plurality of angles. More specifically, given the three-dimensional points captured by the capture device, a computer device is able to render a two-dimensional representation (e.g. a two-dimensional image) of the scene for each eye of a viewer so as to provide a representation with an impression of depth. The computer device may render a series of two-dimensional representations to enable the viewer to look around the scene, where the two-dimensional representations are rendered based on an orientation of the viewer’s head. In this way, the determined representation is useable to provide, for example, a virtual reality (VR), mixed reality (MR), augmented reality (AR), and / or extended reality (XR) experience to the viewer. To enable such a display, the display device 17 is typically a virtual reality headset, that comprises a plurality of sensors to track a head movement of the user. By tracking this head movement, the display device is able to update the images being displayed to the viewer as the viewer moves their head to look about the scene. Typically, this involves the display device sensing the sensor data to an external computer device (e.g. a computer connected to the display device via a wire). The external computer device may comprise powerful graphical processing units (GPUs) and / or computer processing units (CPUs) so that the external computer device is able to rapidly render appropriate two-dimensional images for the viewer based on the three-dimensional images and the sensor data. It will be appreciated that the use of a combination of a headset and an external device is exemplary. More generally, the processing of data and the rendering of images may be performed by various computer devices; for example, a standalone virtual reality headset may be provided, which headset is capable of processing data and rendering images without any connection to an external computer device. In some embodiments, the external computer device may comprise a server device, where the display device 17 may be connected to this server device wirelessly. This enables the two-dimensional images to be streamed from the server to the display device so as to enable the display of high-quality images without the need for a viewer to purchase expensive computer equipment. In other words, operations that require large amounts of computing power, such as the rendering of two-dimensional images based on the three-dimensional representation, may be performed by the server, so that the display device is only required to perform relatively simple operations. This enables the experience to be provided to a wide range of viewers. In some embodiments, a first two-dimensional image is provided to the display device 17 (and / or a connected device) and this first image is 'warped' in order to provide an image for viewing at the display device. The warping of the image comprises processing the image based on the sensor data in order to provide an image that matches a current viewpoint of the viewer. By performing the warping at the display device or another local device, the lag between a head movement of the user and an updating of the two-dimensional representation of the scene can be reduced. One issue with the above-described method of capturing a three-dimensional representation is that it only enables a viewer to make rotational movements. That is, since the points are captured using a single capture device at a single capture location, there is no possibility of enabling translational movements of a viewer through a scene. This inability to move translationally can induce motion sickness within a viewer, can reduce a degree of immersion of the viewer, and can reduce the viewer’s enjoyment of the scene. However, enabling a viewer to move translationally through a three-dimensional representation that has been captured using a single capture device would lead to holes in the scene wherever a viewer moves away from the capture location of this single capture device (since this movement will cause parts of the scene that were not captured by the capture device to come into view of the viewer). Therefore, it is desirable to enable translational movements through the scene while avoiding the display of these holes. To enable such movements, the three-dimensional representation of the scene may be captured using a plurality of capture devices placed at different locations (orthe same capture device placed at different locations). A viewer is then able to move around the scene translationally (e.g. by moving between these locations). More generally, by capturing points for every possible surface that might be viewed by a viewer, a three-dimensional representation of a scene may be captured that allows a suitable two-dimensional representation ofthis scene to be rendered regardless of a location of a viewer (e.g. regardless of where a user is standing within a virtual room). This need to capture points for every possible surface (so as to enable movement about a scene) greatly increases the amount of data that needs to be stored to form the three-dimensional representation. Therefore, as has been described in the application WO 2016 / 061640 A1, which is hereby incorporated by reference, the three-dimensional representation may be associated with a viewing zone, a zone of view (ZOV), or a zone of viewpoints (ZVP), where the three-dimensional representation is arranged to enable a user to move about the viewing zone so as to view the scene. Figure 4 illustrates such a viewing zone 1 and illustrates how the use of a viewing zone limits the amount of image data that needs to be stored to provide a three-dimensional representation of the scene. With the scene shown in this figure, and the viewing zone 1 shown in this figure, it is not necessary to determine attribute data for the occluded surface 2 since this occluded surface cannot be viewed from any point in the viewing zone. Therefore, by enabling the userto only move within the viewing zone (as opposed to around the whole scene) the amount of data needed to depict the scene is greatly reduced while still enabling the user to move to some extent and thereby to avoid the motion sickness that can be induced by augmented reality scenes with only three degrees of freedom. While Figure 4 shows a two-dimensional viewing zone, it will be appreciated that in practice the viewing zone 1 is typically a three-dimensional zone or volume. The viewing zone 1 may, for example, comprise a rectangular volume, ora rectangular parallelepiped, and the viewing zone may have a height of at least 30 cm, a depth of at least 30 cm, and / or a width of at least 30 cm, where these dimensions enable a userto move their head while remaining in the viewing zone. This is merely an exemplary arrangement of the viewing zone; it will be appreciated that viewing zones of various shapes and sizes may be used (e.g. spherical viewing zones). That being said, it is preferable that the viewing zone is limited so as to cover only a part of the volume of the scene, e.g. no more than 50% of the scene, no more than 25% of the scene, and / or no more than 10%ofthe scene. In this regard, if the viewing zone is the same size as the scene, then the three-dimensional representation will simply be a standard representation for virtual reality (that enables a user to move freely about the scene) - and so the use of the viewing zone will not provide any reduction in file size. The viewing zone 1 enables movement of a viewer around (a portion of) the scene. For example, where the scene is a room, the base representation may enable a user to walk around the room so as to view the room from different angles. In particular, the viewing zone enables a user to move through the scene with six degrees-of-freedom (6DoF) movement through the scene, where this aids in the provision of an immersive experience. In some embodiments, the viewing zone 1 may be four-dimensional, where a three-dimensional location of the viewing zone changes over time - and in such embodiments the size and location of the occluded surface 2 may also change over time. More generally, it will be appreciated that viewing zones may be formed in any size or shape, with different sizes and shapes being suitable for different scenes. The volume of the viewing zone 1 is typically selected so that a user is able to move to a degree sufficient to avoid motion sickness and to provide an immersive sensation, while still only enabling a limited amount of movement (where this leads to a smaller file size as compared to an implementation where a user is able to fully move about the scene). Typically, the viewing zone is arranged to enable a user to move their head while they are sitting or standing, but not to freely roam around a room. The viewing zone 1 may have a (e.g. real-world) volume of less than five cubic metres (5m3), less than one cubic metre (1m3), less than one-tenth of a cubic metre (0.1m3) and / or less than one-hundredth of a cubic metre (0.01m3). The viewing zone 1 may also have a minimum size, e.g. the viewing zone may have a volume of at least 1% of the volume of the scene, at least 5% of the volume of the scene, and / or at least than 10% of the volume of the scene. Similarly, the viewing zone may have a volume of at least one-thousandth of a cubic metre (0.01m3); at least one-hundredth of a cubic metre (0.01m3); and / or at least one cubic metre (1m3). The ‘size’ of the viewing zone 1 typically relates to a size in the real world, where if the viewing zone has a length of one metre this means that a user is able to move one metre in the real world while staying within the viewing zone. The size of the viewing zone in the scene may be greater than, equal to, or less than the size of the viewing zone in the real world. For example, the viewing zone may scale a real-world distance so that moving one metre in the real world moves the user less than (or more than) one metre in the scene. This enables the scene to provide different perceptions to the user (e.g. to make the user feel larger or smaller than they are in real life). Similarly, the viewing zone may scale a real-world angle so that rotating one degree in the real world rotates the user less than (or more than) one degree in the scene. Therefore, a viewing zone with a volume of one cubic metre typically connotes a viewing zone in which the user is able to move about a one cubic metre volume in the real world while remaining in the viewing zone. And this may cause the user to move about a volume that is more than, or less than, one metre in the scene. Referring to Figure 5a, in order to capture points for each surface and location that is visible from the viewing zone 1, a plurality of capture devices C1, C2, ..., C9 may be used (e.g. a plurality of virtual scanners and / or a plurality of cameras). Each capture device is typically arranged to perform a capture process, e.g. as described with reference to Figure 3, in which the capture device captures points at a plurality of azimuth angles and elevation angles. By locating the capture devices appropriately, e.g. by locating a capture device at each corner of the viewing zone, it can be ensured that all necessary points are captured. Typically, a first capture device C1 is located at a centrepoint of the viewing zone 1. In various embodiments, one or more capture devices C2, C3, C4, C5 may be located at the centre of faces of the viewing zone; and / or one or more capture devices C6, C7, C8, C9 may be located at edges of and / or corners of the viewing zone. Figure 5a shows a two-dimensional view (e.g. a plan view) of a rectangular viewing zone. It will be appreciated that within this viewing zone each capture device may be located on a shared plane. Equally, the various capture devices may be located on different planes. Referring, for example, to Figure 5b, there is shown a three-dimensional view of a cuboid viewing zone, where there is a capture device located: at the centre of the viewing zone (C1); at the centre of each face of the viewing zone (C2_A, C2_B, C2_C, C2_D, C2_E, C2_F); and at each corner of the viewing zone (C3_A, C3_B, C3_C, C3_D, C3_E, C3_F, C3_G, C3_H. With this arrangement, many locations in the scene (e.g. specific surfaces) will be captured by a plurality of capture devices so that there will be overlapping points relating to different capture devices. Typically, only a single version of the point is stored, where this version may be a highest quality version of the point and / or may be the version of the point associated with the nearest and / or least angled capture device. In order to store the points of the three-dimensional representation, the points may be stored as a string of bits, where a first portion of the string indicates a location of the point (e.g. using x, y, z coordinates) and a second portion of the string locates an attribute of the point. In various embodiments, further portions of the string may be used to indicate, for example, a transparency of the point, a size of the point, and / or a shape of the point. A computer device that processes the three-dimensional representation after the generation of this representation is then able to determine the location and attribute of each point so as to recreate the scene. This location and attribute may then be used to render a two-dimensional representation of the scene that can be displayed to a viewer wearing the display device 17. Specifically, the locations and attributes of the points of the three-dimensional representation can be used to render a two-dimensional image for each of the left eye of the viewer and the right eye of the viewer so as to provide an immersive extended reality (XR) experience to the viewer. As has been described with reference to Figures 5a and 5b, the points of the three-dimensional representation are determined using a set of capture devices placed at locations about the viewing zone, where these capture devices are arranged to capture points at a series of azimuth angles and elevation angles. Typically, each of the capture devices is arranged to use the same capture process (e.g. the same series of azimuth angles and elevation angles), though it will be appreciated that different series of capture angles are possible. For example, there may be a plurality of possible series of capture angles, where different capture devices use different capture angles. In some embodiments, the points are stored based on a capture device identifier and an indication of a distance of the point from the capture device associated with this capture device identifier. Typically, the point is also associated with an angular indicator, which indicates an azimuth angle and / or an elevation angle of the point relative to the identified capture device. It will be appreciated that the storage of the distance and the angle may take many forms. For example, the distance and the angle of each point may be converted into a universal coordinate system, where each capture device has a different location in this universal coordinate system. In particular, each point may be stored with reference to a centre of this universal coordinate system, which centre may be co-located with a central capture device. Where a point is determined based on a distance and an angle from a capture device of a known location in this universal coordinate system, the coordinates of the point in this universal coordinate system can be determined trivially - and the location of the point may then be stored either relative to the capture device or as a coordinate in the universal coordinate system. The capture device identifier may comprise a location of a capture device (e.g. a location in a co-ordinate system of the three-dimensional representation). Equally, the capture device identifier may comprise an index of a capture device. Similarly, the indication of the azimuth angle and the elevation angle for a point may comprise an angle with reference to a zero-angle of a co-ordinate system of the three-dimensional representation. Equally, the azimuth angle and / or the elevation angle may be indicated using an angle index. In some embodiments, the three-dimensional representation is associated with configuration information, which configuration information comprises one or more of: a set of capture device indexes; locations associated with the capture devices and / or the capture device indexes; a spacing of capture devices (e.g. so that locations of the capture devices can be determined from a location of a first capture device and the spacing); angles associated with a capture process for the capture devices; an azimuth angle increment and / or an elevation angle increment associated with the capture process; and a set of angle indexes (e.g. to match an angle index to an angle). With this configuration information, it is possible to determine a location of each capture device from an index of that capture device and / orto determine a capture angle from a known capture process. Therefore, given two numbers: a capture device index and an angle index (that is associated with a combination of a specific azimuth angle and a specific elevation angle), a location of a capture device and a direction of a point from this capture device can be determined. By also signalling a distance of the point from the signalled capture device, a precise location of the point in the three-dimensional space can be signalled efficiently. Typically, the point is associated with each of: a capture device identifier, a distance, an first angular index (e.g. a first azimuth), and a second angle (e.g. a second elevation) This method of indicating a location of a point enables point locations to be identified using a much smaller number of bits than if each point location is identified using x, y, z coordinates. The capture device identifier is typically a capture device index, which is related to a capture device location based on configuration information that has been sent before, or along with, the point data. For example, the configuration information may specify: Location of first capture device is (0,0,0). Step between capture devices is (0,0,1) along the grid, then across the grid, then up the grid. - The grid is (10,10,10). With this information, a capture device with an index of 1 can be determined to be located at (0,0,0); a capture device with an index of 5 can be determined to be located at (0,0,4); a capture device with an index of 12 can be determined to be located at (0,1,0), and so on. Equally, the configuration information may specify a list of capture device indexes and locations associated with these indexes, where this enables the use of a wide range of setups of capture devices. Point series An aspect of the present disclosure relates to the processing of a three-dimensional representation that comprises one or more points, with each point comprising a location and at least one attribute. As described above, the location of each point may be defined in relation to a capture device used to capture that point. The attribute of the point may indicate a colour of the point (e.g. a colour of the point for a left eye of a viewer and a colour of the point for a right eye of a viewer). Equally, or additionally, the attribute may indicate a transparency of the point and / or a normal of the point. The value of this attribute may vary due to a genuine change in the feature that the point represents (for example, a plurality of three-dimensional representations of a scene may represent a room in which a light is turned on and thus a ‘brightness value’ attribute of all points within the room may increase). Equally, the value may change due to ‘noise’, where noise is typically used herein to describe any variance in the attribute that is not caused by a genuine change in the feature. In various embodiments, sources of noise include: rounding errors, computational storage errors, and artifacts introduced by the lossy compression of the representation. Reducing noise-based variations in an attribute of a point improves the clarity and accuracy of the representation as it brings the value of this point closer to the value of the surface that is represented by the point. Furthermore, reducing noise-based variations can decrease the storage size required to store the representation (since a de-noised point with a fixed attribute can be represented by a steady attribute value). Therefore, the present disclosure considers a method of de-noising a (e.g. plurality of) three-dimensional representation(s) of a scene. More generally, the present disclosure considers a method of processing a three-dimensional representation of a scene based on a location of a point within the scene, and in particular based on a determination that the location of this point is similar to a location of a point in a further (e.g. subsequent) representation of the scene. This processing may comprise de-noising, but equally may comprise other operations. Referring to Figure 6, there is described a method of processing a point within a three-dimensional representation of a scene. This method is carried out by a computer device, e.g. the image generator 11 and / or the decoder 15 and / or the display device 17. In a first step 21, the computer device identifies a (first) location of a first point in a first three-dimensional representation of the scene. Identifying the location of the first point may comprise identifying coordinates of the point (e.g. X, Y, and Z coordinates). Equally, identifying the location may comprise identifying a capture device associated with the first point and identifying a distance and / or angle of the point from the capture device. In a second step 22, the computer device identifies a (second) location of a second point in a second three-dimensional representation of the scene, the second point having a second location. Typically, the first three-dimensional representation and the second three-dimensional representation are related representations. For example, each representation may be associated with the same scene but a different time; e.g. the representations may each relate to a frame of a video scene. In particular, the first representation and the second representation may relate to consecutive frames of the video. In a third step 23, the computer device determines a difference between the first location and the second location and determines that the difference is below a threshold difference. In a fourth step 24, in response to the difference being below the threshold difference, the computer device processes one or more of the points. Typically, the fourth step 24 involves the computer device forming a series of points that comprises the first point and the second point. More specifically, the method may comprise forming a time ordered series of points, the time ordered series comprising the first point and the second point. Typically, this process is repeated in order to iterate through a plurality of (e.g. successive) three-dimensional representations of a scene to determine the series (e.g. two representations, five representations, and / or ten representations). The processing of the points may comprise allocating the first point and the second point to a series of points (where the encoding of a video associated with the representations may then be dependent on this series of points). In some embodiments, processing the points comprises determining an attribute value for one or more of the points based on the series of points. For example, an attribute value for each point in a series of points may be determined based on the attribute values of each points (e.g. the attribute value for each point in the series may be set to the average value of the initial attribute values of each point). Such an implementation may be used to remove minor variations between the attribute values of the points so as to simplify the encoding of the series and reduce noise. In some embodiments, the processing comprises fitting a line to the attribute values of the series of points in order to define a (new) attribute value for one or more of the points. The attribute values for the points can then be signalled by referencing this fitted line (e.g. which may be based on an equation) instead of needing to store each attribute value separately in detail. The series is formed of points which have been determined to have locations with differences below a threshold difference. This means all points in the series have been determined to have the same - or substantially the same - location; these points can then be considered to be ‘co-located’ and can be considered to relate to a single location (and / or surface) in a scene, which location has been sampled at different points in time. Often, it can be expected that such a location of a scene will have an attribute value that is constant, or that changes in a predictable manner. For example, if the point is a point on a static, unchanging, surface, then the attribute value may be expected to be a constant value. In such a situation, any variations in the attribute value may be deemed to occur due to (undesirable) noise. Alternatively, if the attribute of the point is changing between two values (e.g. if the surface is changing colour), then this change may be expected to occur via a continuous or steady process. At least with these situations in mind, the present disclosure considers a method of processing one or more points of this series, where this typically enables the points to be stored in an efficient manner. For example, where an attribute of a point is steady over time, a plurality of points for successive representations may be stored by storing a single point of an initial representation and then indicating that this point persists through the plurality of representations. It will be appreciated that while the method presented in Figure 6 mentions only a first point and a second point, the method is typically repeated so that the series includes a third point and a fourth point and so on. Typically, the method comprises analysing a plurality of representations in order to determine one or more series of points, with each series relating to a point that is in the same location in a plurality of (e.g. successive) representations. Typically, this results in a time ordered series of points being formed comprising one point from each of a plurality of successive three-dimensional representations of a scene (e.g. where each representation is associated with a frame of a video). In various embodiments, the series may comprise at least 3 points, at least 5 points, at least 10 points, and / or at least 20 points. Processing the points typically comprises altering an attribute value and / or a location of the one or more points. For example, processing the points may comprise replacing the attribute values of one or more points of the series based on the attribute values of the other points in the series. In a (simple) practical example, if six points in a series have a first attribute value and a seventh point has a second attribute value that is slightly different to the first attribute value, the seventh point may be modified so as to have the first attribute value. Processing the points may equally comprise identifying one or more anomalous points and outputting an identifier of an anomalous point. Furthermore, processing the points may comprise identifying a (e.g. large) change in an attribute value of the point, where this can be used to determine that a substantial change has occurred in the scene. In some embodiments, processing the points comprises removing one or more of the points from a representation and / or adding one or more new points to a representation based on the series. In some embodiments, processing the points comprises fitting a line and / or curve through the (attribute of the) series of points, where one or more of the attribute values of the points may then be replaced based on this fitted line. This may be considered to involve replacing initial (or‘raw’) attribute values of the points with modified (or 'denoised’) attribute values. In some embodiments, the series comprises points from only a subset of a plurality of three-dimensional representations of the scene. In this regard, the ability to form a series from points in each of the plurality of three-dimensional representations of a scene is highly dependent on the nature of the scene. In particular, scenes which primarily comprise stationary objects (relative to the viewing zone) are likely to contain a consistent set of points throughout each of a plurality of three-dimensional representations. In contrast, scenes comprising primarily moving objects are likely to contain a substantially changing series of points (e.g. as points move behind surfaces and are thus occluded from the viewing zone). In some embodiments, the first and second three-dimensional representations of a scene referred to in the method of Figure 6 are chronologically successive three-dimensional representations from within the plurality of representations. In such embodiments, the processing can be considered to be occurring in a ‘forwards’ direction. The plurality of three-dimensional representations of a scene may also be processed in a reverse order (so that the first representation is associated with a time that is later than the second representation) - in such embodiments, the processing may be considered to occur in a ‘backwards direction. The series of points described herein can be formed from moving through the representations in either of a forward direction or a backward direction. The similarity of a time ordered series of points produced from processing in the forward or backward direction may vary depending on the implementation of the threshold difference, as is described below. In some embodiments, with the method of Figure 6, a single first point is selected from the first three-dimensional representation, a single (corresponding) second point is selected from the second three-dimensional representation), and the method continues on to form a single (typically time ordered) series. However, typically, multiple instances of the method of Figure 6 are performed in parallel so that the method may comprise determining a plurality of ‘first’ points at different locations within the first three-dimensional representation and then a plurality of corresponding second points in the second three-dimensional representation. This may allow a plurality of series to be formed simultaneously in parallel. These separate series may then be processed in parallel (e.g. by different computer devices). This may advantageously reduce the processing time needed to form and process multiple series. As described above, a location of a point may be stored in multiple ways. The implementation (and nature) of the threshold difference may vary depending on the implementation of the storage of the location. In some embodiments, the location of a point is defined based on a capture device that is used to capture the point, a distance of the point from this capture device, and at least one angular identifier (e.g. an index or a value of an angle) that indicates an angle between the capture device and the point. The point may comprise a single angular index that identifies both an elevation angle and an azimuthal angle of a point from the capture device. Equally, each angle may be indicated by a second identified. Preferably, the at least one angular identifier is stored in binary as an integer. The distance is typically stored as a floatingpoint number. In this embodiment, the computer device may compare the different pieces of a location of a point in different ways. For the binary integers representing the at least one angle index the threshold difference may be implemented as requiring a strict identity of the two binary integers being compared (e.g. if the angle index of the first point is 5 then the angle index of the second device must be exactly 5 in order to be determined to be below the threshold difference). For the floating-point number representing the distance from a capture device, a value within a specified tolerance level above or below a first point may be determined as being below the threshold difference. In some embodiments, the location of a point is stored using coordinate values (e.g. using cartesian (X, Y and Z) coordinates). The coordinate values may be captured directly by a capture device or, alternatively, may be converted to X,Y,Z coordinates using the distance from a capture device, angle index and trigonometric functions. In this embodiment, all three pieces of the location of a point are typically stored as floating-point numbers. Thus, depending on the machine accuracy of the computer device, the threshold difference for all three values may be implemented as a tolerance level as described above. In a related embodiment, if the at least one angular identifier is also implemented as a non-integer floating point number then a similar tolerance level implementation may be used. It will be appreciated that the locations of points may typically be converted between these two forms. For example, where a location is defined with reference to a capture device, this location could be converted to a (absolute) coordinate location based on a coordinate location of the capture device. The tolerance level (e.g. the threshold distance that is used to determine whether two points are substantially co-located) may be implemented in multiple ways, including but not limited to; a flat range of values above and below the first point (e.g. to determine whether the second point is +1- 0.05mm from the first point), a percentage range of values above and below the first point (e.g. to determine whether the second point is + / - 5% from the first point), and similar flat or percentage ranges above or below a mean -or weighted mean - value of distance for a plurality of previously determined points. In some embodiments, different aspects of the location of the points are evaluated differently. In particular, the computer device may consider a first distance threshold (relating to a difference in the distance of the points from a capture device and / or the viewing zone) and one or more second angular thresholds (relating to a difference in the distance of the points from a capture device and / or the viewing zone). The computer device may determine that differences between the locations of the first point and the second point satisfy each of a plurality of such distances. In some embodiments, the threshold difference is dependent on a function applied to the previous points in the series (e.g. the tolerance level may be set as + / - the value of the standard deviation of some portion of the points previously formed into the time series). In some embodiments, the threshold difference is implemented in dependence on the value of one or more attributes of points in the series or in dependence on the locations of points in the series. For example, the threshold difference may scale depending on the value of the distance from a capture device or the viewing zone, where higher separations between points may be acceptable further from the viewing zone. In some embodiments, the threshold difference scales with the average value of an attribute across all points within the three-dimensional representation. For example, the method may estimate an average value of a brightness attribute across all points within the three-dimensional representation to determine if the representation depicts a scene captured in low light levels. Upon determining the scene is captured in low light levels, the tolerance level may be increased for all points to accommodate the expected lower accuracy data for determining the location of a point under those conditions. In some embodiments, the method involves determining a probability that the first point and the second point are at the same location (‘co-located’). This probability may be based on multiple factors, including but not limited to: the attribute values of the points (e.g. where points with the same attribute may be assessed as being likely to be co-located; the locations of the points; and an expected amount of movement in the scene. This probability may be used to determine whether to allocate the points into a series, where, for points below a threshold probability, a user may be able to confirm whether the points are collocated. Typically, points that exceed a threshold probability of co-Iocation are allocated to a series, where this threshold probability may depend on a user input or on a feature of the scene. The probability may be determined by a neural network or similar probability estimating methods. Equally, the probability may be determined using a deterministic algorithm. The probability may be implemented directly as a tolerance level, wherein points are only formed into a time ordered series if the likelihood of them being the same location is sufficiently high. Optionally or alternatively, the prediction of the likelihood (or a concatenation of all likelihood predictions within the time ordered series) may be provided to the user such that the user may selectively exclude points with low or insufficient likelihoods. In some embodiments, one or more of the points is associated with (e.g. comprises) a motion vector, which motion vector indicates a direction of that point’s movement in the scene. Determining a difference between the first location (of the first point) and the second location (of the second point) may comprise evaluating a motion vector of the first or second point. For example, a motion vector associated with the first point may be evaluated to identify whether the first point at the first location will move to the second location. In a practical example, a predicted location of a first point in a first three-dimensional representation may be determined based on a location of the first point and a motion vector of the first point and this location may then be compared to a location of a second point in a second three-dimensional representation. The first point and the second point may each be included in a series based on the predicted location of the first point being within a threshold distance of the second point. The predicted location of the first point may also be compared to a predicted location of the second point. In other words, in embodiments where points are associated with a motion vector, determining a difference between the first location and the second location may further comprise determining a difference in a predicted location based at least in part on a calculation of a first motion vector at the first location and a second location. In some embodiments, the motion vector is provided as a cartesian vector with X, Y and Z dimension in a coordinate system with an origin in the three-dimensional representation. More generally, the motion vector may be provided in any form of vector representation such as, but not limited to, spherical coordinates, cylindrical coordinates and other coordinate systems wherein the origin may be centred on a capture device or a centre of a viewing zone. In some embodiments, the computer device is arranged to determine a difference between a first motion vector associated with the first point and a second motion vector associated with the second point, where the processing of the points (e.g. the formation of the time series) may depend on a difference of these motion vectors being within a threshold. This may allow denoising to be performed on attribute values of points which represent objects or features which are changing location within the three-dimensional representation, but which are effectively representing the same object or feature and hence may have similar or related attribute values. In some embodiments, the points comprise a normal, where the normal extends in a direction perpendicular to a surface represented by that point. Determining that the first point and the second point are similar (and so may be included in the series) may comprise comparing a normal of a first point and a second point. The first point and the second point may be included in a series based on the normal of the first point and the normal of the second point being within a threshold difference of each other. In some embodiments, the individual pieces of a point that define the location of this point may be stored to different levels of precision (e.g. the distance of the point from the capture device may be stored with a greater number of bits than the angular identifier(s)). At least for this reason, an input may be provided at the computer device to allow a user to select the settings for the threshold difference. Typically, a setting is provided to control the tolerance level and / or threshold value used to determine whether points are colocated. In some embodiments, the computer device trials a plurality of different threshold differences and either presents a plurality of possible threshold differences to a user or selects a 'best' threshold difference automatically. Typically, a ‘best’ threshold difference is selected based on the settings that provide a maximum length of the time ordered series and / or that provide a minimum size of the processed representation. In some embodiments, the points comprising the plurality of three-dimensional representations of a scene are captured by a plurality of capture devices, as discussed with reference to Figures 5a and 5b. The plurality of capture devices are located at different positions within the viewing zone to capture views of the scene from these different positions and thereby to enable the rendering of an image from any point within the viewing zone. Different points at different locations in the scene (and the representation) may be captured by different capture devices. Typically, the method of Figure 6 comprises identifying a first point and a second point that are captured by the same capture device. Furthermore, the method of Figure 6 may comprise identifying, for each capture device, points that have been captured by that capture device in different representations and determining whether these points are co-located. In some embodiments, the first point and the second point may be captured by different capture devices, where determining a co-Iocation of the points may then comprise converting the locations of the points into a shared coordinate system and then determining a co-Iocation of the points. However, typically the method comprises sorting the points of the representation into groups based on the capture device used to capture each point. The method as described in Figure 6 may then be performed (sequentially or in parallel) on each group of points with the same capture device identifier (e.g. different groups may be evaluated by different computer devices). Turning now to Figure 7, there is shown a graph of an example attribute for a time ordered series. More specifically, the graph of Figure 7 shows a series of points connected by a point-to-point line. Each point is arranged along the x-axis according to a time value associated with the representation containing that point. As discussed above, in addition to a location, a point typically also has at least one attribute (e.g. a color, a transparency or a surface normal direction), the value of which is determined by a capture device used to capture the point. The skilled person will appreciate that in embodiments where points are associated with multiple attributes, then a graph similar to the one depicted in Figure 7 may be produced for each attribute and then processed separately via similar methods. The data that defines the value of the attribute at each time value, as captured by the capture device, will be henceforth referred to as the ‘raw data’ and is labelled as such on the legend of the graph of Figure 7. The time value of a point is associated with the specific three-dimensional representation of a scene that contains the point (e.g. a first three-dimensional representation may be associated with a time t=0; a second three-dimensional representation may be associated with a time t=0.25, etc.; typically, the method comprises determining 24 three-dimensional representations for each second of a video). In this embodiment the time value is a value in seconds, but the skilled person will appreciate that the time value only functions to order the points chronologically and that this may be implemented in multiple ways. Typically, if the plurality of three-dimensional representations of a scene are provided as a video (i.e. a time ordered series of frames) then the time value of the point will correspond to a frame number and the actual ‘clock’ time will merely be determined by the frame rate of said video. In Figure 7, the raw data (i.e. the value captured by the capture device) is shown to vary in signal strength between each time value interval. Certain changes in the signal strength may represent a genuine change in the signal strength at the location; equally, certain changes in the signal may be a result of noise. An aspect of the present disclosure relates to a method of processing capture raw data e.g. to denoise this raw data. Such denoising is described below with reference to Figures 8a and 8b. Denoisinq Within the context of forming a (e.g. time ordered) series of points, the present disclosure considers a method of denoising attribute values of points within the series of points to produce a weighted average series. In particular, the method of denoising involves selecting a moving window in which a section of the time ordered series is considered and denoising performed before the window is 'moved along’ the time ordered series to consider a slightly different section of points. Referring to Figure 8a, there is shown a method of denoising a time ordered series of points. This method is carried out by a computer device, e.g. the image generator 11 and / or the decoder 15. In a first step 31, the computer device defines a window of a certain size. In a second step 32, the computer device selects a sub-series of points from within the series, the subseries having a length that corresponds to the size of the window. The sub-series begins at a chronologically first point in the time ordered series and continues until a number of points has been selected from the series that is equal to the size of the window. In a third step 33, the computer device predicts a further point of the series based on the selected subseries. For example, based on the attribute values of points in first, second, third, and fourth (successive) representations, the computer device may predict the attribute value of a point in a fifth representation. In some embodiments, the computer device predicts the attribute value of the point in the fifth representation based at least in part on the raw (original) value of said point in the fifth representation (in some embodiments, this original point in the fifth representation may be associated with a highest weighting). Furthermore, in some embodiments, the computer device predicts the attribute value of the point in the fifth representation based on the attribute values of points in a sixth (and, e.g. seventh etc.) representation, where the first to fourth representations are associated with values of the point prior to the fifth representation and the sixth representation is associated with a value of the point after the fifth representation. The number of representations considered before and / or after the fifth representation may be defined by a user and / or may be determined based on a feature of the scene, the points, or the representations. In a fourth step 34, the computer device generates a weighted average series by collating the predictions of a next point for every point in the time ordered series. Referring to Figure 8b, there is shown a method for predicting points of the representation, e.g. as in the third step 33 of Figure 8a. In a first step 41, the computer device applies a weighted average operation to the sub-series of points to determine a prediction of a next point. Typically, this operation comprises performing a weighted averaging of the points so as to favour chronologically later points in the sub-series. While the weighted average series of points has been described in terms of a non-recursive (Finite Impulse Response) filtering operation, in some embodiments, a next point is predicted by a recursive (Infinite Impulse Response) operation, such that future predictions may depend on previous predictions and hence further increase the accuracy of said predictions. In some embodiments the weighted average operation comprises a combination of both recursive and non-recursive filtering operations. More generally, numerous types of operations (including, and other than, filtering operations and averaging operations) may be used to predict the next point. In a second step 42, the computer device moves the window forward a step along the time ordered series, such that a chronologically first point in the sub-series is removed and that a chronologically next point in the time ordered series of points is appended into the sub-series. It will be appreciated that while the second step 42 has described moving 'forwards’ through the points, the second step may equally involve moving ‘backwards’ through the points. In a third step 43, the computer device iterates over the series (e.g. repeating the first step 41 and the second step 42) until a prediction of a next point has been determined for every point in the time ordered series. The result of this denoising method is to produce a new set of attribute values, henceforth referred to as the 'denoised data’. An example of this denoised data is shown in the graph of Figure 7 as a line of significantly reduced variation. The denoised data has many advantages over the raw data when used in the processing of three-dimensional representations of a scene. Processing a point of the representation may comprise replacing a raw (original) value of the attribute of that point with a denoised attribute value of the point. It will be appreciated that this may be done in-situ, such that a raw attribute value of a point is immediately amended or adjusted based on a prediction of its attribute value at the third step 33 of Figure 8a. In some embodiments, the adjustment is performed by the computing device. Optionally or alternatively, the computing device may store predictions of the attribute values (or the resulting adjustments to the raw value of the attribute) and apply all the adjustments at or after the fourth step 34 of Figure 8a. Performing the method of Figure 8a and 8b on the raw data causes a reduction in the average difference in attribute values between two temporally adjacent points and this may help to reduce shimmering. Shimmering is the visual effect a viewer of the plurality of three-dimensional representations may experience when the points comprising said three-dimensional representations appear to rapidly fluctuate in brightness or colour over time. Denoising may make these fluctuations smaller and hence less noticeable to a human eye. Furthermore, the lower average difference in attribute values allows the data to be encoded by a series of ‘deltas’ - a value representing the relative difference in the attribute from the previous point. The smaller value can more easily be compressed leading to an overall reduction in file size for the plurality of three-dimensional representations. The size of the window determines the quantity of points that will be considered simultaneously by the computing device. Typically, a user of the computer device is able to select a window size. Increasing the size of the window has the effect of improving the accuracy of the prediction of a next point since by providing more points as inputs to the weighted average operation, trends in the values of the attribute are more detectable by statistical methods. However, considering more points simultaneously also increases the amount of data wrangling required by the computer device. Therefore, the optimal window size may depend on the processing power and / or time available for processing the representations. In some embodiments, this window size is determined based on an available hardware or time, where the computer device may be arranged to determine the window size so as to provide a processed representation in less than an input amount of time. Referring to Figure 9, there is shown a graph of an example attribute for a time ordered series, similar to Figure 7, but comprising data which includes a change in value of an attribute of a magnitude multiple times larger than the typical variation due to noise. When the change in value of the attribute is of a consistent magnitude within the ‘moving window’ the method produces a better prediction of a next point. However, when a change in value occurs which is multiple magnitudes larger than at other points in the moving window the method may incorrectly overpredict the size of the change and hence cause an ‘overshoot’ in its prediction of a next point. An example of this can be seen in Figure 9 at around 4 seconds along the x-axis. To reduce the overshoot, the determination of the denoised attributes may be based on each of a forward pass and a backward pass over the series of points. In this regard, in the method of Figure 8a and 8b, the sub-series typically begins at a chronologically first point and the method performs a 'forward pass' (e.g. stepping through the data in order of increasing time value). Additionally, or alternatively, the sub-series may begin at a chronologically last point and perform a ‘backward pass’. Because the overshoot only effects the prediction of a point after the large attribute value change by starting at the last point and iterating through in order of decreasing time value causes the overshoot to occur at a different time value. The method of Figure 8a and 8b may further comprise combining - e.g. performing an operation on, or selecting a single one of-the denoised values obtained from each of a forward pass and a backward pass in order to remove errors relating to overshoots and hence produce an accurately denoised data set without artefacts. Typically, the forward pass and the backward pass each produce a prediction for one or more (or each) point in the original raw data. Some points may only be predicted by one of the forward pass and the backward pass (e.g. a first n points in the series may only be determined by the backwards pass and a last n points in the series may only be determined by a forwards pass). Furthermore, the method of Figure 8a and 8b may comprise generating a combined prediction of the value of an attribute of a point. The combined prediction may be a value determined by a mathematical operation or function, implemented by the computing device, based on any / all of: the prediction obtained from the forward pass, the prediction obtained from the backwards pass, and the raw 'original' value of the attribute of a point. For example, the operation may include calculating an average value for one or more denoised points where the average value is an arithmetic average of the predictions obtained from the forwards pass and the backwards pass. However, it will be appreciated that the combined prediction may be any form of mathematical operation which results in a composite value of its inputs, such as, but not limited to, a weighted average. In some embodiments, the combined prediction is formed using a weighted averaging of the form: xc = k * xa + (1 - k) * xb, where xc is a point value of the combined prediction, xa and xb are values of the forward prediction and the backwards prediction, and k is a weighting value between 0 and 1. This method of forming the combined prediction ensures a smooth interpolation of points. The method may then provide three predictions for the value of an attribute at a point in time in addition to the original raw data (the 'forwards’ prediction, the ‘backwards’ prediction, and the combined prediction). Typically, combining the forward and backward pass to produce a final denoised dataset comprises selecting a value for the attribute for each point in time, wherein the value is selected from one of; the raw data, the forward pass prediction of a next point, the backward pass prediction of a next point, and the combined prediction. The selection of an attribute value for each point may depend on one or more of: a difference between the raw value and each denoised value; a user input; and the attribute values of surrounding points. The choice of which value to select may be determined by rules which prioritise the selection of one value over another under specific circumstances. For example, if the range covered by the difference of the forward pass and backward pass prediction of a next point includes the value of the raw data then the average value may be selected, because it is likely that the average value represents a more accurate value then the raw data value. If the raw data value is lower than both the forward and backward pass predictions of a next point, then the lower or the two predictions may be selected instead. In some embodiments, constraints may also be directly applied to limit the acceptable range of values for the combined prediction. For example, requiring the combined prediction to be within a certain number of standard deviations of the original raw data. The standard deviation of the raw data may be calculated from the entire time ordered series of points or may be calculated from the points within the current window being considered by the computing device. If the combined prediction does not meet any applied constraints, the computing device may instead not provide the combined prediction. The exact set of rules used to determine which value to select may vary depending on the user’s preference so that the set of rules may be presented to the user and be amendable by the user. In some embodiments related time series may be denoised together. As described above, points may have a plurality of attribute values, in some cases some of the plurality of attribute values may be related, where these attribute values share some degree of dependence in their value such that a change in a first attribute value may be expected to cause or indicate a change in a second attribute value. For example, a point may comprise separate attribute values for each of a left eye and a right eye of a user and it may be expected that a change in the value of the left eye value would be associated with a change in the value of the right eye value, (e.g. these colour values may increase simultaneously in intensity if a brightness level of the scene increases). In some embodiments, the methods of Figures 8a and 8b may be performed simultaneously on a plurality of time ordered series comprising related attribute values, wherein the prediction of a further point in a first series of points in the third step of Figure 8a may be further based on a sub-series selected from a related second series of points. Other examples of attribute values which may be related include, but are not limited to; a brightness value and a transparency value or a motion vector value and a normal vector value. In some embodiments, if a predicted point is more than a defined (e.g. predetermined) number of standard deviations from an original point, then the original point is selected for use in the final dataset. While the above description has primarily described the use of a forward and backwards pass as part of a denoising process, more generally the method may comprise determining one or more of a forward pass and a backward pass in order to process a point (e.g. for the purpose of denoising or for another purpose). In some embodiments, the method of Figure 6 may further comprise a fifth step, wherein the processed points are recombined into a processed plurality of three-dimensional representations. Additionally, the method may further comprise presenting the processed plurality of three-dimensional representations to the user. It will be understood that presented may mean any form of rendering, displaying or playing to the user wherein the user is able to view the processed plurality of three-dimensional representations (for example, via a computer device display or extended reality experience device display). Typically, the processed plurality of three-dimensional representations is presented to the user without additional processing. However, it will be appreciated that in some embodiments it may be preferable to add artificial noise (e.g. film grain) before presenting the processed plurality of three-dimensional representations to a user (e.g. in order to represent a specific aesthetic effect). In some embodiments, the method includes generating one or more two-dimensional images from the three-dimensional representations without adding any noise to the attribute values of the points of the three-dimensional representations. Alternatives and modifications It will be understood that the present invention has been described above purely by way of example, and modifications of detail can be made within the scope of the invention. The representation is typically arranged to provide an extended reality (XR) experience (e.g. a representation that is useable to render a XR video). The term extended reality (XR) covers each of virtual reality (VR), augmented reality (AR), and mixed reality (MR) and it will be appreciated that the disclosures herein are applicable to any of these technologies. The representation may be encoded into, and / or transmitted using, a bitstream, which bitstream typically comprises point data for one or more points of the three-dimensional representation. The point data may be compressed or encoded to form the bitstream. The bitstream may then be transmitted between devices before being decoded at a receiving device so that this receiving device can determine the point data and reform the three-dimensional representation (or form one or more two-dimensional images based on this three-dimensional representation). In particular, the encoder 13 may be arranged to encode (e.g. one or more points of) the three-dimensional representation in order to form the bitstream and the decoder 14 may be arranged to decode the bitstream to generate the one or more two-dimensional images. While the detailed description has considered the processing of three-dimensional representations, it will be appreciated that many of the methods described herein are applicable to more general concepts. For example, the processing steps may be used to process representations or images of any dimensionality (e.g. 2D images, 3D images etc.). While the detailed description has considered a comparison of (e.g. locations of) points, it will be appreciated that the computer device may be arranged to compare a (e.g. indirect) feature of a plurality of 5 points. For example, the computer device may be arranged to compare a transformation of the points. In this regard, a processing stage may involve transforming the three-dimensional representations, e.g. to compress the representations. The computer device may then be arranged to evaluate the transformed representations to identify a similarity between points of these transformed representations and to form a series based on similar points in the transformed representations. 10 Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.

Claims

1. A method of processing a three-dimensional representation of a scene, the method comprising: identifying a first point in a first three-dimensional representation, the first point having a first location;identifying a second point in a second three-dimensional representation, the second point having a second location;determining that a difference between the first location and the second location is below a threshold difference; andprocessing one or more of the points in dependence on the difference being below the threshold difference.

2. The method of claim 1, wherein processing the points comprises modifying an attribute value and / or a location of one or more of the points.

3. The method of claim 1 or 2, wherein processing the points comprises forming a time-ordered series of points, the series comprising the first point and the second point.

4. The method of claim 3, wherein processing the points comprises performing a denoising operation on the series of points, preferably, wherein:processing the points comprises performing a denoising method on values of attributes of points within the series; and / orwherein performing a denoising operation comprises altering the location of a point5. The method of any preceding claim, wherein each point is associated with a capture device, and wherein:identifying the first point comprises identifying a first point associated with a first capture device; andidentifying the second point comprises identifying a second point associated with the first capture device.

6. The method of claim 5, comprising:identifying a third point associated with a second capture device, the third point having a third location; andidentifying a fourth point associated with the first capture device, the fourth point having a fourth location;determining that a difference between the third location and the fourth location is below a further threshold difference; andprocessing one or more of the third point and the fourth point in dependence on the difference being below the further threshold difference.

7. The method of any preceding claim, comprising forming a plurality of series of points in parallel, each series comprising a plurality of co-located points, preferably wherein:each series is associated with a single capture device;a subset of the plurality of series are collated into a single composite series.

8. The method of any preceding claim, wherein the threshold difference is dependent on the distance of the first point and / or the second point from a viewing zone associated with the first three-dimensional representation and / or the second three-dimensional representation, preferably wherein the viewing zone defines a subset of the scene through which a user is able to move.

9. The method of any preceding claim, comprising:determining a plurality of components of the location of the first point and a plurality of components of the location of the second point, preferably wherein the components comprise at least one distance component and at least one angular component defined in relation to a capture device; anddetermining that a difference between each component of the locations of the first point and the second point are below a respective threshold difference.

10. The method of any preceding claim, wherein the threshold difference depends on one or more of:an attribute values of the first point and / or the second point;an average of the attribute values value of the first point and the second point; anda weighted average of the attribute values value of the first point and the second point.

11. The method of any preceding claim, wherein determining that a difference between the first location and the second location is below a threshold difference further comprises:determining a probability that the first point and second point relate to the same surface based on one or more of:the value of one or more attributes of the first point and the second point;a normal of the first point and the second point; anda component of the locations of the first point and the second point.

12. The method of any preceding claim, comprising:predicting an attribute value of a further point of the series based on attribute values of the first point and the second point.

13. The method of any preceding claim, wherein processing the one or more points comprises:modifying an attribute value of a point of a series based on a predicted attribute value forthat point;preferably, comprising replacing an initial attribute value of the point with the predicted attribute value.

14. The method of any preceding claim, further comprising;defining attribute values for one or more points of the first three-dimensional representation or the second three-dimensional representation with reference to a delta value, wherein the delta value indicates a difference between a value of an attribute at a first point in a series of points and a value of an attribute at a second point in the series.

15. The method of any preceding claim, wherein the method further comprises:determining a third point in the first three-dimensional representation and a fourth point in the second three-dimensional representation, andforming a second series of points including the third point and the fourth point in dependence on a difference in a location of the third point and the fourth point being below a threshold difference.

16. The method of any preceding claim, comprising:determining a first ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the first ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a first direction;determining a second ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the second ordered series has a similar location in the scene, andwherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a second direction, the second direction being opposite to the first direction; andprocessing one or more points of the first ordered series and / or the second ordered series.

17. A method of processing a plurality of three-dimensional representations of a scene, the method comprising:determining a first ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the first ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a first direction;determining a second ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the second ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a second direction, the second direction being opposite to the first direction andprocessing one or more points of the first ordered series and / or the second ordered series.

18. The method of claim 16 or 17, comprising determining a modified value of an attribute for one or more points in the first ordered series and / or the second ordered series;preferably comprising:determining a modified series of attribute values for the points in the first ordered series based on the initial values of these points; and / ordetermining a modified series of attribute values for the points in the second ordered series based on the initial values of these points.

19. The method of any of claims 15 to 18, comprising:selecting a sub-series of points from within each of the series of points; and determining a modified attribute value for a further point in each of the series of points based on the attribute values of the points in the sub-series of points, preferably, wherein selecting a sub-series of points comprises:sliding a window along the series of points so as to select a plurality of sub-series of points from the series of points; anddetermining a modified attribute value for each sub-series so as to form a modified series of attribute values, preferably determining a modified attribute value further comprises weighting the attribute values of the sub-series, more preferably wherein the weighting is biased to favour points towards an end of the sub-series.

20. The method of claim 19, comprising:sliding the window along the series of points in a first direction so as to form a forwards modified series of attribute values; andsliding the window along the series of points in a second direction, the second direction being opposite the first direction, so as to form a backwards modified series of attribute values.

21. The method of claim 20, comprising forming a combined modified series of points based on the forwards modified series and the backwards modified series, preferably wherein forming the combined modified series comprises:determining one or more attribute values for the combined modified series by averaging corresponding attribute values of the forwards modified series and the backwards modified series;determining one or more attribute values for the combined modified series as being one of the values of the forwards modified series and the backwards modified series; anddetermining one or more attribute values for the combined modified series as being one of the original attribute values of the points of the first ordered series and / or the second ordered series.

22. An apparatus for processing a three-dimensional representation of a scene, the apparatus comprising: means for identifying a first point in a first three-dimensional representation, the first point having a first location;means for identifying a second point in a second three-dimensional representation, the second point having a second location;means for determining that a difference between the first location and the second location is below a threshold difference; andmeans for processing one or more of the points in dependence on the difference being below the threshold difference.

23. An apparatus for processing a three-dimensional representation of a scene, the apparatus comprising: means for determining a first ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the first ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a first direction;means for determining a second ordered series of points from a set of three-dimensional representations of a scene, wherein each point in the second ordered series has a similar location in the scene, and wherein determining the first ordered series comprises evaluating each representation while moving through the set of representations in a second direction, the second direction being opposite to the first direction; andmeans for processing one or more points of the first ordered series and / or the second ordered series.

24. The method of any preceding claim, comprising forming a bitstream based on the first three-dimensional representation and / or the second three-dimensional representation.

25. A bitstream formed using the method of claim 24.36Application No: GB2410601.5 Examiner: Stuart PurdyClaims searched: 1-16, 18-22, 24 and 25Date of search: 30 January 2025Patents Act 1977: Search Report under Section 17Documents considered to be relevant:Category Relevant to claims Identity of document and passage or figure of particular relevance X 1-9,15, 22, 24 and 25 CN 112132971 B (HEFEI DILUSENSE TECH CO LTD) See whole document and note in particular figure 1 and the description associated with steps 110 to 130; X 1,2, 4-12, 22, and 24, and 25; US 2018 / 0122129 Al (PETERSON et al.) See whole document and note in particular paragraphs [0032], [0033], [0043] to [0046]; v A 1, 2, 5, 8, 10, 12, 13,22, 24, and 25 US 2019 / 0087979 Al (MAMMOU et al.) See whole document and note in particular paragraphs [0003], [0004], [0280] and [0550];Categories:X Document indicating lack of novelty or inventive A Document indicating technological background and / or state step of the art. Y Document indicating lack of inventive step if P Document published on or after the declared priority' date but combined with one or more other documents of same category'. before the filing date of this invention. & Member of the same patent family E Patent document published on or after, but with priority date earlier than, the filing date of this application.Field of Search:International Classification:Subclass Subgroup Valid From G06T 0017 / 00 01 / 01 / 2006 G06T 0005 / 50 01 / 01 / 2006 G06T 0005 / 70 01 / 01 / 2024 H04N 0013 / 204 01 / 01 / 2018 H04N 0019 / 503 01 / 01 / 2014

Citation Information

Patent Citations

  • Three-dimensional human body modeling method, device, electronic device and storage medium

    CN112132971B

  • Generation, transmission and rendering of virtual reality multimedia

    US20180122129A1

  • Point cloud compression

    US20190087979A1