Determining a location of a point

By utilizing capture device identifiers and angular information, the method addresses the high processing and storage demands of three-dimensional representations, facilitating efficient rendering and transmission of virtual reality scenes.

GB2640277APending Publication Date: 2025-10-15V NOVA INT LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
GB2024005094
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Existing three-dimensional representations require substantial processing power and large file sizes, necessitating significant storage and bandwidth for transmission, particularly in generating virtual reality videos.

Method used

A method and apparatus for determining the location of a point in a three-dimensional representation by identifying capture device identifiers, distances, and angular identifiers, allowing for efficient determination of point size and coverage based on capture device configuration and angular resolution.

Benefits of technology

Reduces processing requirements and file sizes by leveraging capture device information and angular identifiers to accurately determine point locations, enabling efficient rendering and transmission of three-dimensional scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method of determining a location of a point in a 3D representation of a scene. This is carried out by a computer device, e.g. an image generator 11 and / or the decoder 15 to identifying point data associated with the point. The computer device: identifies 21 an indicator of a capture device used to capture the point, typically, this comprises identifying a portion of a string of bits associated with a capture device index; Identifies 22 an indicator of an angle of the point from the capture device, typically, identifying an angle index, e.g. an azimuth index and / or an elevation index and / or a combined azimuth / elevation index, which index(es) identifies a step of the capture process during which the point was captured; Determines the location of the capture device and the angle of the point from the capture device, typically, a capture device index, which relates to a capture device location based on configuration information that has been sent before, or along with, the point data. Determining a location of the point, typically, based on the location of the capture device, the capture angle, and a distance of the point from the capture device (where this distance is specified in the point data for the point).
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Disclosure The present disclosure relates to methods, systems, and apparatuses for determining a location of a point, in particular determining a location of a point in a point cloud. Background to the Disclosure Three-dimensional representations of environments are used in many contexts, including for the generation of virtual reality videos, in which depth information for a plurality of points of the representation is used to generate different images for a left eye and a right eye of a user. Typically, substantial processing power is required to determine such a three-dimensional representation, and the file size of files associated with these representations is typically large so that substantial amounts of storage are needed to keep the files and substantial amounts of bandwidth are required to transfer the files. Summary of the Disclosure According to an aspect of the present disclosure, there is described a method of determining a location of a point in a three-dimensional representation of a scene, the method comprising: identifying point data associated with the point; determining, from the point data: a capture device identifier associated with a capture device used to capture the point; a distance of the point from the capture device; and one or more angular identifiers that indicate an angle of the point from the capture device; and determining the location of the point based on the distance, the capture device identifier, and the one or more angle identifiers. Preferably, the capture device identifier comprises a capture device index. Preferably, the method comprises determining a location of the capture device based on the capture device identifier and configuration information associated with the representation. Preferably, determining the location of the capture device comprises comparing the capture device Index to a table of capture devices located in the configuration information. Preferably, the method comprises determining a size and / or coverage of the point. Preferably, the size and / or coverage of the point is dependent on an angular resolution of the three-dimensional representation. Preferably, determining the size comprises determining coordinates for a plurality of corners of the point, preferably four corners. Preferably, the method comprises determining the size and / or coverage and / or the coordinates of the point in dependence on the angle of the point from the capture device and the distance of the point from the capture device. Preferably, the method comprises determining the size and / or coverage and / or the coordinates of the point by extending a plane between two angular brackets associated with the capture device, the plane preferably being extended in a direction perpendicular to the angle of the point from the capture device. Preferably, the method comprises determining the size and / or coverage and / or the coordinates of the point in dependence on a normal of the point from the capture device, preferably determining the size and / or coverage and / orthe coordinates of the point by extending a plane between two angular brackets associated with the capture device, the plane preferably being extended in a direction perpendicular to the normal of the point. Preferably, the method comprises determining the coordinates of the point by: determining an angular bracket associated with the point and the angle of the point from the capture device, the point lying within this angular bracket, and the angular bracket being formed of a plurality of angular boundaries; determining a location of a centre of the point based on the angle and the distance; determining a normal of the point from the point data; and extending a plane from the location of the centre of the point to each of the angular boundaries, the plane being extended in a direction perpendicularto the normal of the point from the capture device; and determining the corners of the point as the intersections of the plane and the angular boundaries. According to an aspect of the present disclosure, there is described a method of determining the corner coordinates of a point in a three-dimensional representation of a scene, the method comprising: determining a capture device used to capture the point; determining an angle of the point from the capture device; determining a distance of the point from the capture device; determining an angular bracket associated with the capture device and the angle, the point lying within this angular bracket, and the angular bracket being formed of a plurality of angular boundaries; determining a location of a centre of the point based on the angle and the distance; determining a normal of the point; and extending a plane from the location of the centre to each of the angular boundaries, the plane being extended in a direction perpendicular to the normal of the point from the capture device; and determining the corners of the point as the intersections of the plane and the angular boundaries. Preferably, the three-dimensional representation is associated with a frame of a video, the video comprising a plurality of frames and each frame being associated with a separate three-dimensional representation of the scene, wherein the method comprises: receiving, at a first time, the configuration information; receiving, at a second time, point data for a plurality of three-dimensional representations; and determining the locations of a plurality of points in different three-dimensional representations using the configuration information and the point data. Preferably, the one or more angle identifiers comprise an azimuth angle identifier and / or an elevation angle identifier. Preferably, the one or more angle identifiers comprise an azimuth angle and / or an elevation angle. Preferably, the one or more angle identifiers comprise a reference to an angular bracket of the three-dimensional representation, the angular bracket covering a volume of a space of the three-dimensional representation. Preferably, the size of the angular bracket is dependent on an angular resolution of the three-dimensional representation. Preferably, the angular bracket is associated with one or more of: a first azimuthal angular boundary and a second azimuthal angular boundary; and a first elevational angular boundary and a second elevational angular boundary. Preferably, the capture device and / or the three-dimensional representation is associated with a plurality of angular brackets. Preferably, the angle of each angular bracket is defined by the configuration information. Preferably, the configuration information identifies an angular resolution, an azimuthal resolution, and / or an elevational resolution. Preferably, the configuration information defines a relationship between the one or more angle identifiers and a corresponding angle, preferably, between the identifiers and each of an azimuth angle and an elevation angle. Preferably, the point is associated with a quadrilateral in the three-dimensional representation. Preferably, the method comprises rendering an image based on the location of the point and / or the location of the corners of the point. Preferably, the method comprises rendering an image based on point data for a plurality of points of the three-dimensional representation. Preferably, the image is associated with an extended reality (XR) and / or virtual reality (VR) experience. Preferably, the point data comprises a size, the size defining a number of angular brackets covered by the point, preferably wherein the size comprises a size index that defines the size with reference to an index table. According to an aspect of the present disclosure, there is described a method of determining point data for a point in a three-dimensional representation of a scene, the method comprising: identifying a capture device; determining a distance of a point from the capture device, preferably wherein the point is a point on a surface of an object; determining an angle of the point from the capture device; and defining point data for the point, the point data including: a capture device identifier indicating the capture device used to capture the point; distance data indicating the distance of the point from the capture device; and one or more angle identifiers indicating the angle of the point from the capture device. Preferably, the method comprises determining an attribute value for the point, the attribute value being included in the point data. Preferably, the three-dimensional representation is associated with a viewing zone, the viewing zone comprising a subset of the scene and / or the viewing zone enabling a userto move through a subset of the scene. Preferably, the viewer is able to move within the viewing zone with six degrees of freedom (6DoF). Preferably, the viewing zone has a volume of less than 50% of the volume of the scene, less than 20% of the volume of the scene, and / or less than 10% of the volume of the scene. Preferably, the viewing zone has, or is associated with, a volume, preferably a real-world volume, of less than five cubic metres (5m3), less than one cubic metre (1 m3), less than one-tenth of a cubic metre (0.1 m3) and / or less than one-hundredth of a cubic metre (0.01 m3). Preferably, the three-dimensional representation comprises a point cloud. Preferably, the method comprises forming one or more two-dimensional representations of the scene based on the three-dimensional representation, preferably comprising forming a two-dimensional representation for each eye of a viewer. Preferably, the point is associated with one or more of: a location; an attribute; a transparency; a colour; and a size. Preferably, the point is associated with an attribute for a right eye and an attribute for a left eye. Preferably, the scene comprises one or more of: an extended reality (XR) scene; a virtual reality (VR) scene; an augmented reality (AR) scene; and a mixed reality (MR) scene. According to an aspect of the present disclosure, there is described an apparatus for performing the aforesaid method. According to an aspect of the present disclosure, there is described an apparatus for determining a location of a point in a three-dimensional representation of a scene, the apparatus comprising: means for (e.g. a processor for) identifying point data associated with the point; means for (e.g. a processor for) determining, from the point data: a capture device identifier associated with a capture device used to capture the point; a distance of the point from the capture device; and one or more angular identifiers that indicate an angle of the point from the capture device; and means for (e.g. a processor for) determining the location of the point based on the distance, the capture device identifier, and the one or more angle identifiers. According to an aspect of the present disclosure, there is described an apparatus for determining the corner coordinates of a point in a three-dimensional representation of a scene, the apparatus comprising: means for (e.g. a processorfor) determining a capture device used to capture the point; means for (e.g. a processor for) determining an angle of the point from the capture device; means for (e.g. a processor for) determining a distance of the point from the capture device; means for (e.g. a processor for) determining an angular bracket associated with the capture device and the angle, the point lying within this angular bracket, and the angular bracket being formed of a plurality of angular boundaries; means for (e.g. a processor for) determining a location of a centre of the point based on the angle and the distance; means for (e.g. a processor for) determining a normal of the point; and means for (e.g. a processor for) extending a plane from the location of the centre to each of the angular boundaries, the plane being extended in a direction perpendicular to the normal of the point from the capture device; and means for (e.g. a processor for) determining the corners of the point as the intersections of the plane and the angular boundaries. According to an aspect of the present disclosure, there is described an apparatus for determining point data for a point in a three-dimensional representation of a scene, the apparatus comprising: means for (e.g. a processor for) identifying a capture device; means for (e.g. a processor for) determining a distance of a point from the capture device; means for (e.g. a processor for) determining an angle of the point from the capture device; and means for (e.g. a processor for) defining point data for the point, the point data including: a capture device identifier indicating the capture device used to capture the point; distance data indicating the distance of the point from the capture device; and one or more angle identifiers indicating the angle of the point from the capture device. Preferably, the apparatus comprises means for (e.g. a processor for) determining an attribute value for the point, the attribute value being included in the point data. According to an aspect of the present disclosure, there is described an apparatus for carrying out the aforesaid method, the apparatus comprising one or more of: a processor; a communication interface; and a display. According to an aspect of the present disclosure, there is described a system for carrying out the aforesaid method, the system comprising one or more of: a processor; a communication interface; and a display. Any feature in one aspect of the disclosure may be applied to other aspects of the invention, in any appropriate combination. In particular, method aspects may be applied to apparatus aspects, and vice versa. Furthermore, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be construed accordingly. Any apparatus feature as described herein may also be provided as a method feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure, such as a suitably programmed processor and associated memory. It should also be appreciated that particular combinations of the various features described and defined in any aspects of the disclosure can be implemented and / or supplied and / or used independently. The disclosure also provides a computer program and a computer program product comprising software code adapted, when executed on a data processing apparatus, to perform any of the methods described herein, including any or all oftheir component steps. The disclosure also provides a computer program and a computer program product comprising software code which, when executed on a data processing apparatus, comprises any of the apparatus features described herein. The disclosure also provides a computer program and a computer program product having an operating system which supports a computer program for carrying out any of the methods described herein and / or for embodying any of the apparatus features described herein. The disclosure also provides a computer readable medium having stored thereon the computer program as aforesaid. The disclosure also provides a signal carrying the computer program as aforesaid, and a method of transmitting such a signal. The disclosure extends to methods and / or apparatus substantially as herein described with reference to the accompanying drawings. The disclosure will now be described, by way of example, with reference to the accompanying drawings. Description of the Drawings Figure 1 shows a system for generating a sequence of images. Figure 2 shows a computer device on which components of the system of Figure 1 may be implemented. Figure 3 shows a method of determining a three-dimensional representation of a scene. Figure 4 shows a scene comprising a viewing zone. Figures 5a and 5b show arrangements of capture devices for determining points of the three-dimensional representation. Figure 6 describes a method of determining a location of a point of the three-dimensional representation. Figure 7 shows a method of determining an angle of a point from a capture device used to capture the point. Figure 8 describes a method of determining a size of a point of the three-dimensional representation. Figure 9 illustrates embodiments of the method of Figure 8. Figures 10 and 11 indicate shapes that may be covered by a point. Description of the Preferred Embodiments Referring to Figure 1, there is shown a system for generating a sequence of images. This system can be used to generate, and then display, a representation of an environment, which may comprise a VR environment (or an XR environment). The system comprises an image generator 11, an encoder 12, a transmitter 13, a network 14, a receiver 15, a decoder 16 and a display device 17. These components may each be implemented on separate apparatuses. Equally, various combinations of these components may be implemented on a shared apparatus; for example, the image generator 11, the encoder 12, and the transmitter 13 may all be part of a single image data generation device. Similarly, the receiver 15, the decoder 16, and the display device 17 may all be a part of a single image rendering device. Typically, the system comprises at least one encoding computer device (e.g. a server of a content provider) and at least one rendering computer device (e.g. a VR headset). Referring to Figure 2, each of the components, and in particular the image generator 11, the encoder 12, the transmitter 13, the receiver 15, the decoder 16 and the display device 17 is typically implemented on a computer device 20, where, as described above, a plurality of these components may be implemented on a shared computer device. Each computer device comprises one or more of: a processor 21 for executing instructions (e.g. so as to perform one or more of the steps of the various methods described below), a communication interface 22 for facilitating communication between computer devices (e.g. an ethemet interface, a Bluetooth® interface, or a universal serial bus (UBS) interface, a memory 23 and / or storage 24 for storing information and instructions (e.g. a random access memory (RAM), a read only memory (ROM), a hard drive disk (HDD) a solid state drive (SSD), and / or a flash memory, and a user interface 25 (e.g. a display, a mouse, and / or a keyboard) for enabling a user to interact with the computer device. These components may be coupled to one another by a bus 25 of the computer device. The computer device 20 may comprise further (or fewer) components. In particular, the computer device (e.g. the display device 17) may comprise one or more sensors, such as an accelerometer, a GPS sensor, or a light sensor. These sensors typically enable the computer device to identify an environmental condition and / or an action of wearer of the display device. Turning back to Figure 1, the image generator 11 is configured to generate a sequence of image data (e.g. a sequence of image frames) to enable the display device 17 to use this image data to display a plurality of images. The image data may comprise one or more digital objects and the image data may be generated or encoded in any format. For example, the image data may comprise point cloud data, where each point has a 3D position and one or more attributes. These attributes may, for example, include, a surface colour, a transparency value, an object size and a surface normal direction. Each attribute may have a value chosen from a continuous range or may have a value chosen from a discrete set. The image data enables the later rendering of images. This image data may enable a direct rendering (e.g. the image data may directly represent an image). Equally, the image data may require further processing in order to enable rendering. For example, the image data may comprise three-dimensional point cloud data, where rendering a two-dimensional image using this data requires processing based on a viewpoint of this two-dimensional image. The image data may comprise depth map data, where one or more pixels or objects in the image is associated with a depth that is specified by the depth map data. The depth map data may be provided as a depth map layer, separate from an image layer. In some contexts, such as MPEG Immersive Video (MIV), the image layer may instead be described as a texture layer. Similarly, in some contexts, the depth map layer may instead be described as a geometry layer. The image data may include a predicted display window location. The predicted display window location may indicate a potion of an image that is likely to be displayed by the display device 17. The predicted display window location may be based on a viewing position (such as a virtual position and / or orientation of the user in a 3D environment) of the user, where this viewing position may be obtained from the display device. The predicted display window location may be defined using one or more coordinates. For example, the predicted display window location may be defined using the coordinates of a corner or center of a predicted display window, and may be defined using a size of the predicted display window. The predicted display window location may be encoded as part of metadata included with the frame. The image data for each image (e.g. each frame) may include further information, which may be provided as a part of an image, e.g. as part of the point cloud data, or as separate layers. In particular, the image data may include audio information or haptic feedback information indicating audio or haptics which can accompany displayed visual data. An audio layer or haptic layer may accompany each image, and may be omitted for images where no accompanying audio or haptics are required. Similarly, the image data may comprise interactivity information, where the image data may contain or indicate elements with which a user can interact. The interactivity information may, for example, define a behaviour of an element, where a user is able to interact with the element based on this behaviour. The behaviour typically defines a change in an element that occurs as a result of a user interaction where this change may comprise a change in the attributes of the element or in the rendering of the element. As an example, where an image contains a target element, the target element may be arranged to disappear when a user interacts with this element, or to provide feedback indicating that the user has interacted with the target. This interactivity data may be provided as part of, or separately to, the image data. The image data may indicate, or may be combinable with, a state of the virtual environment, a position of a user, or a viewing direction of the user. Here, the position and viewing direction may be physical properties of the user in the real-world, or position and viewing direction may also be purely virtual, for example being controlled using a handheld controller. The image generator 11 may, for example, obtain information from the display device 17 that indicates the position, viewing direction, or motion of the user. Equally, the image generator may generate image data such that it can later be combined with this position, viewing direction, or motion, where the image generator may generate a full scene which is only partially viewed by a user depending on the position of that user. In some cases, the generated image may be independent of user position and viewing direction. This type of image generation typically requires significant computer resources such as a powerful GPU, and may be implemented in a cloud service, or on a local but powerful computer. For example, a cloud service (such as a Cloud Rendering Service (CRN)) may reduce the cost per-user and thereby make the image frame generation more accessible to a wider range of users. Here “rendering” refers at least to an initial stage of rendering to generate an image. Further rendering may occur at the display device 17 based on the generated image to produce a final image which is displayed. The image generator 11 may, for example, comprise a rendering engine for initially rendering a virtual environment such as a game or a virtual meeting room. The encoder 12 is configured to encode frames to be transmitted to the display device 17. The encoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. In some embodiments, the image generator 11 may transmit raw, unencoded, data through the network 14. However, such transmission typically leads to a high file size and requires a high bandwidth so that it is typically desirable to encode the data prior to the transmission. The encoder 12 may encode the image data in a lossless manner or may encode the data a lossy manner. The encoder may apply inter-frame or intra-frame compression based on a currently-encoded frame and optionally one or more previously encoded frames. The encoder may be a multi-layer encoder, such as an low complexity enhancement video codec (LCEVC) enabled encoder. Where the generated frames comprise depth map data, the encoder 12 may perform layered encoding on each instance of image data (e.g. each frame) to generate an encoded frame comprising a base depth map layer and an enhancement depth map layer. Encoding a depth map in this way may improve compression. In some applications, such as HDR video, depth maps are desirably highly detailed with a bit depth of up to twelve or fourteen bits, which is a significant increase in the data to be transmitted. As a result, providing ways to improve compression of the depth map can make more realistic depth map-based displays viable when performing rendering or transmission of rendered data in real-time. Furthermore, this type of layered encoding makes it easy to drop (and then pick back up) one or more of the layers, which provides flexibility and tools for bandwidth management. Layered encoding is also helpful as the final decoder / user device (such as a user display device) can choose whether to process these extra layers. For example, in a non-layered approach, the best the end device (i.e. the receiver, decoder or display device associated with a user that will view the images) can do is determine that it does not have enough resources for a given quality (be it resolution, frame rate, inclusion of depth map) and then signal to the controller / renderer / encoder that it does not have enough resources. The controller then will send future images at a lower quality. In that alternative scenario, the end device still unfortunately has to process the higher quality data until the lower quality data arrives, if it can process the received images at all. In some of the described embodiments, this situation is improved upon because when / if the end device determines for example that it does not have the processing capabilities to handle the highest level of quality, then it can drop and / or choose not to process certain layers. The end device may also signal to the controller that it needs a lower level of quality, but in the meantime the end device can only process the number of layers that it can handle. Therefore, the end device can react to conditions much more quickly. In some cases, depth map data may be embedded in image data. In this case, the base depth map layer may be a base image layer with embedded depth map data, and the enhancement depth map layer may be an enhancement image layer with embedded depth map data. Alternatively, when the generated images comprise a depth map layer separate from an image layer and multi-layer encoding is applied, the encoded depth map layers may be separate from the encoded image layers. This has the advantage that the encoded depth map layers can be dropped under some conditions while still retaining image layers that can be displayed (albeit with a lower level of realism). For example, the encoded depth map layers can be dropped by a transmitter or encoder when available communication resources are reduced, or can be dropped by an end device which lacks the processing resources to handle the highest level of quality. Similarly, if some images comprise an audio base layer, a haptic feedback base layer, an audio enhancement layer or a haptic feedback enhancement layer, these can be processed or dropped flexibly. Again similarly, if some images comprise an interactivity data base layer or an interactivity enhancement layer these can be processed or dropped flexibly. For example, certain interactions may only be possible where a threshold bandwidth is available, where complex interactions (e.g. those enabling a conversation with a digital object) may be disabled before less complex interactions (e.g. changing a pixel colour) are disabled. Additionally or alternatively, where the image data comprises point cloud data, the encoder may apply a point cloud data encoding technique such as described in European patent application EP21386059.6, which is incorporated herein by reference. Such a point cloud encoder may act as a base encoder for a layered encoding technique such as LCEVC or VC-6. Notably LCEVC and VC-6 techniques encode and decode a layered signal, but are agnostic about the content type of data encoded in the signal. For example, the signal can include textures, video frames, geometry or depth data, meshes, point clouds, rendering attributes or physics engine attributes. The transmitter 13 may be any known type of transmitter for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The transmitter 13 may be configured to make decisions about how to transmit the image data, and / or may provide feedback to the encoder 12 or the image generator 11. For example, the transmitter may determine available communication resources (e.g. bandwidth) for transmitting image data, and may drop one or more layers from an encoded frame, or indicate to the image generator and / or encoder that image data should be generated and encoded with fewer layers, when insufficient bandwidth is available for transmission of all generated data. As specific examples, the transmitter may be configured to drop a depth map layer, an LCEVC enhancement layer, or a VC-6 enhancement layer from a frame when insufficient communication resources are available. The network 14 provides a channel for communication between the transmitter 13 and the receiver 15, and may be any known type of network such as a WAN or LAN or a wireless Wi-Fi or Bluetooth network. The network may further be a composite of several networks of different types. Many users only have access to a network with a bandwidth of 30MBps which can lead to latency jitter when streaming. The required bandwidth and the observed latency can be reduced by means of tactics such as forward-looking rendering and last-millisecond reprojection, which are enabled by improved compression. The receiver 15 may be any known type of receiver for wired or wireless communications, including an Ethernet transmitter or a Bluetooth transmitter. The decoder 16 is configured to receive and decode an encoded frame. The decoder may be implemented using executable software or may be implemented on specific hardware such as an ASIC. The display device 17 may for example be a television screen or a VR headset. The timing of the display may be linked to a configured frame rate, such that the display device may wait before displaying the image. The display device may be configured to perform warping, that is, to obtain a final display window location, adjust a warpable image to obtain a final image corresponding to a final viewing direction of the user, and display the final image. In this regard, the image data is typically arranged to provide a warpable image for which a portion of the image that is displayed at the display device 17 is dependent on a position or orientation of a viewer. The warpable image may then be rendered before a most upto date viewing direction of the user is known. The warpable image may be transmitted to the display device, or the warpable image may be transmitted to a rendering node which is near to the display device, and the display device or rendering node may perform time warping to generate a displayed image portion based on the warpable image and the most up to date viewing direction of the user. As mentioned above, a single device may provide a plurality of the described components. For example, a first rendering node may comprise the image generator 11, encoder 12 and transmitter 13. Additional similar rendering nodes may be included in the system, and may work together to generate the sequence of frames. In one case, multiple rendering nodes may each provide separate image data to an image data assembling node; for example, each rendering node may provide a part of a sequence of frames to a frame assembling node. For example, the receiver 15, decoder 16 or display device 17 may be configured to assemble parts of image data from multiple sources to generate a sequence of images for display on the display device. Alternatively, the image data assembling node may be separate from the receiver 15, decoder 16 and display device 17. Additionally or alternatively, multiple rendering nodes may be chained. In other words, successive rendering nodes may add to a sequence of image data as it passes from rendering node to rendering node, and eventually a complete sequence of image data is then provided to the receiver 15. Furthermore, each rendering node may obtain components of a render from multiple upstream rendering nodes and / or distribute components of a render to multiple downstream rendering nodes. A chain of rendering nodes may be useful for performing different rendering tasks that require different quantities of processing resources, or different frame rates. For example, a company may provide distributed processing in the form of a centralised hub which has abundant processing resources but is distant from users, and peripheral locations which have more scarce processing resources but are closer to users. Expensive but fairly static rendering features such as background lighting or environmental impact on sound may be generated at the central hub (for example using ray tracing), while features that require fewer resources but faster responses or higher frame rates may be generated closer to the user. In other words, the more responsive a rendering feature needs to be, the lower latency it needs between the rendering node which generates the feature and the user display and, in a chain of rendering nodes, the node which generates each rendering feature can be chosen based on a required maximum latency of that feature. On the other hand, if it is expensive to generate a rendering feature, then it may be preferable to generate the feature less frequency and with a higher maximum latency. For example, a static, high-quality background feature may be generated early in the chain of rendering nodes and a dynamic, but potentially lower-quality, foreground feature may be generated later in the chain of rendering nodes, closer to the user device. Here, environmental impact on sound means, for example, a set of surfaces may be constructed where each surface has different sound reflection and absorption properties depending upon material and shape. The frame rates may be matched by creating multiple frames with features generated at the lower frame rate, and combining them with the frames with features generated at the higher frame rate. In a nonlimiting embodiment, a preliminary rendering generates volumetric object data including motion vectors at a first (lowest) frame rate, then produces 2D rendered frames plus depth information for a specific user at a second (higher) frame rate, then transmits video plus depth data to the user device, which produces final frames for display via space warping (depth-based reprojections) at a third (highest) frame rate. One or more of these steps may be performed in combination with the other described embodiments. The viewing position of the user may change as additional rendering tasks are performed at different rendering nodes in the chain. Each or any rendering node may obtain an updated viewing position before performing its respective rendering task. Additionally, the system may simultaneously generate multiple sequences of image data for different respective users or different respective display devices. For example, in the context of a VR or AR experience, each user or display device may view a different 3D environment, or may view different parts of a same 3D environment. When using a chain of rendering nodes, each node may serve multiple users or just one user. For example, a starting rendering node (e.g. at a centralised hub) may serve a large group of users. For example, the group of users may be viewing nearby parts of a same 3D environment. In this case, the starting node may render a wide zone of view (“field of view”) which is relevant for all users in the large group. The starting node may send this wide field of view to a first middle rendering node which renders additional aspects of the 3D environment. These additional aspects may for example be aspects which require less processing power to render, or may be aspects which are specific to individual users of the group. Additionally, the middle rendering node may render features in a smaller field of view than the starting node - this smaller field of view may be relevant to each user rather than the group of users. The first middle rendering node may additionally only serve a smaller number of users (e.g. half of the large group of users), with the remaining users being served by a second middle rendering node which also receives the wide field of view from the starting node. The middle rendering node(s) may then send sequences of second partially or fully rendered frames to an end device for each user. The end device may perform further processes such as warping or focal distance adjustments, optionally using depth map data. Preferably, each rendering node encodes the partially or fully rendered frames before transmitting them on to a next rendering node or to the receiver 15. This means that the required communication resources can be reduced when the rendering nodes are separated by one or more networks, or more generally are implemented in a distributed system such as a cloud. However, each rendering node in a chain is encoding a different partially or fully rendered frame, with different data. Therefore, it may be advantageous for different rendering nodes to use different rendering formats and / or encoding formats. For example, the output from a first rendering node may be point cloud data which logically describes a 3D scene. This point cloud data can be encoded using the techniques of EP21386059.6. A second rendering node may then operate on the point cloud data to generate image data that is more readily displayed by a generic display device, without requiring the display device to model the 3D environment. This image data may be encoded using video coding techniques. The chaining of rendering nodes may be extended to arbitrary tree structures, where a rendering node obtains partially rendered frames from more than one preceding rendering node, and generates further partially or fully rendered frames based on the multiple obtained sequences of partially rendered frames. For example, a content rendering network (CRN) comprising numerous rendering nodes may be used to serve a volumetric event to a large number of same-time users, such as users participating in a shared virtual environment. Rendering the same event for each user is far more expensive in terms of computation time and power consumption than rendering the volumetric effect once and performing the rendering equivalent of multicasting the volumetric effect for multiple users. For example, each user may have a second rendering node (such as a VR headset), and the network may comprise a central first rendering node. The first rendering node may render the volumetric event, and distribute partially rendered frames depicting the volumetric event to the different second rendering nodes. The second rendering node for each user may then integrate the partially rendered frames depicting the volumetric event into a view of the virtual environment which is currently being shown to each user, based on parameters such as the user’s virtual position. The receiver 15, decoder 16 and display device 17 may be consolidated into a single device, or may be separated into two or more devices. For example, some VR headset systems comprise a base unit and a headset unit which communicate with each other. The receiver 15 and decoder 16 may be incorporated into such a base unit. In some embodiments, the network 14 may be omitted. For example, a home display system may comprise a base unit configured as an image source, and a portable display unit comprising the display device 17. In the event that the decoder 16 or the display device 17 does not or cannot handle one or more layers, the receiver 15 or another transmitter associated with the decoder or display device may send a corresponding layer drop indication back through the network 14. The layer drop indication may be received by each rendering node. A rendering node which generates partially or fully rendered frames for that specific decoder or display device may cease generating the dropped layer. On the other hand, a rendering node which generates partially or fully rendered frames for multiple end devices may disregard a layer drop indication received from one end device (as the dropped layer is still needed for other devices). Alternatively, rendering nodes which serve multiple end devices may record received layer drop indications, and may cease generating the dropped layer only when all end devices served by the rendering node indicate that the layer is to be dropped. In preferred examples, the encoders or decoders are part of a tier-based hierarchical coding scheme or format. Hierarchical coding enables frames to be communicated with higher resolution and / or higher frame rate than is possible in single-tier coding schemes. In hierarchical coding, one or more enhancement layers is communicated with base data, where the enhancement layers can be used to up-sample the base data at the decoder, for example providing up-sampling in a spatial or temporal dimension. When combined with equivalent down-sampling of the original frames and generation of the enhancement layer at an encoder, hierarchical coding can overall provide lossless compression of data, with higher resolution and / or higher frame rate for a given transmission bit rate. Examples of a tier-based hierarchical coding scheme include LCEVC: MPEG-5 Part 2 LCEVC (“Low Complexity Enhancement Video Coding”) and VC-6: SMPTE VC-6 ST-2117, the former being described in PCT / GB2020 / 050695, published as WO 2020 / 188273, (and the associated standard document) and the latter being described in PCT / GB2018 / 053552, published as WO 2019 / 111010, (and the associated standard document), all of which are incorporated by reference herein. However, the concepts illustrated herein need not be limited to these specific hierarchical coding schemes. A further example is described in WO2018 / 046940, which is incorporated by reference herein. In this example, a set of residuals are encoded relative to the residuals stored in a temporal buffer. LCEVC (Low-Complexity Enhancement Video Coding) is a standardised coding method set out in standard specification documents including the Text of ISO / IEC 23094-2 Ed 1 Low Complexity Enhancement Video Coding published in November 2021, which is incorporated by reference herein. The system describes above is suitable for generating and presenting a representation of a scene, where this scene displays media content to a user. The scene typically comprises an environment, where the user is able to move (e.g. to move their head or to turn their head) to look, around the environment and / or to move around the environment. For example, the scene may be a scene of a room in a building, where the user is able to move around the room (e.g. by moving in the real-world and / or by providing an input to a user interface) in order to inspect various parts of the room. Typically, the scene is a XR (e.g. a VR) scene, where the user is able to move about the scene in three degrees of freedom (3DoF) or six degrees of freedom (6DoF) so as to experience the scene. As has been described with reference to Figure 1, the image generator 11 may be arranged to determine point cloud data, where each point of the point cloud has a 3D position and one or more attributes. More generally, the image generator (or another component) is arranged to determine a three-dimensional representation of a scene, where this three-dimensional representation is thereafter used to generate two-dimensional images that are presented to a user at the display device 17. While the points are typically points of a point cloud, more generally the disclosure extends to any point that is associated with a location and a value. Therefore, the points may, more generally, be considered to be data (or datapoints), which data is associated with a location and a value, and the ‘points’ may comprise polygons, planes (regular or irregular), Gaussian splats, etc. Referring to Figure 3, there is described a method of determining (an attribute for) a point of such a three-dimensional representation. The method comprises determining the attribute using a capture device, such as a camera or a scanner. The scene may comprise a real scene, in which attribute values are captured using a camera, or a virtual scene (e.g. a three-dimensional model of a scene), in which attribute values are captured using a virtual scanner. Where this disclosure describes ‘determining a point’ it will be understood that this generally refers to determining a point that has a location and an attribute value, where determining the point comprises determining the attribute value and / or storing a point that comprises at least an attribute value and a location value (these values may be indirect values, e.g. where the location is identified relative to another point). Once a plurality of points have been captured, these points can be stored as a three-dimensional representation (e.g. a point cloud) so as to enable the reconstruction of the three-dimensional scene based on this representation. Typically, the three-dimensional representation is associated with a plurality of images or frames; for example, the three-dimensional representation may be associated with a video. In such embodiments, determining a point may involve determining a point of a four-dimensional representation, which has three spatial dimensions and a time dimension. The location and / or attributes of such a point may be dependent on time (e.g. the point may have an initial location as well as a speed and / or acceleration so that the location of the point can be determined at a plurality of different times based on these variables). In such embodiments, the method may be considered to comprise determining a plurality of three-dimensional representations, where each frame of a video is dependent on a different three-dimensional representation. In such a situation, there is typically a dependency between the plurality of three-dimensional representations; for example, the points of a second three-dimensional representation may be determined / encoded based on a preceding first three-dimensional representation. Therefore, this situation could be considered to involve either of: determining a four-dimensional representation; and determining a plurality of three-dimensional representations. Typically, the scene comprises a simulated scene that exists only on a computer. Such a scene may, for example, be generated using software such as the Maya software produced by Autodesk®. The attributes determined using the methods described herein may then depend on virtual objects located within the scene as well as a virtual lighting arrangement used in the scene. In a first step 11, a computer device initiates a capture process for a capture device, the capture process being initiated with an initial azimuth angle (e.g. of 0°) and an initial elevation angle (e.g. of 0°). In a second step 12, the computer device causes a point to be captured using the capture device at the current azimuth angle and current elevation angle. Capturing a point typically comprises assigning an attribute value to the point, which attribute value may, for example, be a color of the point and / or a transparency value of the point. Typically, the point has one or more color values associated with each of a left eye and a right eye of a viewer. Capturing the point may also comprise determining a normal value associated with the point, e.g. a normal of a surface on which the point lies. Typically, capturing the point further comprises determining a location of the point, e.g. by determining a distance of the point from the capture device (e.g. camera or scanner). In practice, determining the point may comprise sending a ‘ray’ from the capture device and then stepping through a computer model to determine which surface of the computer model is impacted by the ray. The colour, transparency, and normal of this surface are then recorded alongside the distance of the surface from the capture device. In a third step, 13, the computer device determines whether a point has been captured for the capture device at each azimuth of a range of azimuths and in a fourth step 14, if points have not been captured at each azimuth, then the azimuth angle is incremented and the method returns to the second step 12 and another point is captured. The azimuth angle may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of azimuth angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. Once a point has been captured for each azimuth, in a fifth step 15, the computer device determines whether a point has been captured for the capture device at each elevation of a range of elevations and in a sixth step 16, if points have not been captured at each elevation, then the azimuth angle is reset to the initial value, elevation angle is incremented and the method returns to the second step 12 and another point is captured. The elevation angles may, for example, be incremented by between 0.01° and 1° and / or by between 0.025° and 0.1°. Typically, the range of elevation angles is selected to be 360° (i.e. so that the capture device captures points surrounding the entirety of the capture device), but it will be appreciated that other ranges are possible. In a seventh step 17, once points have been captured for each azimuth angle and each elevation angle, the scanning process ends. This method enables a capture device to capture points at a range of elevation and azimuth angles. This point data is typically stored in a matrix. The point data may then be used to provide a representation of the scene to a user, e.g. the three-dimensional representation formed by the point data may be processed to produce two-dimensional images for each eye of a user, with these images then being shown to a user via the display device 17 to provide a virtual reality experience to the viewer. By using the captured data, a video can be provided to a viewer that enables the viewer to move their head to look around the scene (while remaining at the location of the capture device). It will be appreciated that the capture pattern (or scanning pattern) described with reference to Figure 3 is purely exemplary and that numerous capture patterns are possible. In general, the capture process for each capture device comprises capturing one or more points at one or more azimuth angles and / or one or more elevation angles. The ‘points’ captured by the capture device are typically associated with a size, such as a height, a width, or a depth. That is, the points typically relate to two-dimensional pixels and / or three-dimensional voxels (or planes). In this regard, there is necessarily some space between the locations of adjacent points (since if the points had no width, then an infinite numberof points would be required to capture points at each angle). The size provides points that depict a non-negligible area of the three-dimensional space so that a plurality of points can be fit together to provide a depiction of the scene to a viewer. The width and height of each point is typically dependent on the distance of that point from the capture device, where more distant points have a larger width / height. The width and height of each point is typically determined so that when each point is displayed, there is no space between adjacent points (indeed, there may be some overlap between points to ensure that no gaps appear between points). This height / width of each point can be determined at the time of capturing the points, or can be determined or defined after the capture of the points. Typically, the points comprise a size value, which is stored as a part of the point data. For example, the points may be stored with a width value and / or a height value. Typically, the minimum width and the minimum height of a point are set by the angle increment of the azimuth angle and the elevation angle respectively. The size may be then specified in terms of this angle increment and / or in terms of this minimum width / minimum height (e.g. as being a multiple of the angle increment). In some embodiments, the size value is stored as an index, which index relates to a known list of sizes (e.g. if the size may be any of 1x1, 2x1, 1x2, 2x2, pixels this may be specified by using 3 bits and a list that relates each combination of bits to a size). The size may be stored based on an underscan value. In this regard, where an object is very near to the viewing zone it may be captured using an unnecessarily dense arrangement of points. Therefore, certain surfaces or areas of the representation may be associated with an underscan value, which underscan value defines a reduction in the number of points captured as compared to a representation without underscan. The size of the points may be defined so as to indicate this underscan value. In an exemplary embodiment, the underscan value is an integer value between 0 and 3 and the size is stored as a combination of point dimensions (e.g. a width in the range [0,2]) and a height in the range ([0,2]) and an underscan factor (e.g. an underscan factor in the range [0,3]). In some embodiments, the width and the height are dependent on the underscan factor. For example, when the underscan factor exceeds a threshold value, the possible height and width values may be limited. In a specific example, when the underscan factor is 3, the width and the height may be limited to the range [0,1], The size may then be defined as size = underscan*9 + height*3 + width. Such a method provides efficient storage and indication of width, height, and underscan values. By capturing points at a plurality of azimuth angles and elevation angles, e.g. using the method described with reference to Figure 3, it is possible to provide a three-dimensional representation of the scene that can later be used to enable a viewer to view the scene from a plurality of angles. More specifically, given the three-dimensional points captured by the capture device, a computer device is able to render a two-dimensional representation (e.g. a two-dimensional image) of the scene for each eye of a viewer so as to provide a representation with an impression of depth. The computer device may render a series of two-dimensional representations to enable the viewer to look around the scene, where the two-dimensional representations are rendered based on an orientation of the viewer’s head. In this way, the determined representation is useable to provide, for example, a virtual reality (VR), mixed reality (MR), augmented reality (AR), and / or extended reality (XR) experience to the viewer.To enable such a display, the display device 17 is typically a virtual reality headset, that comprises a plurality of sensors to track a head movement of the user. By tracking this head movement, the display device is able to update the images being displayed to the viewer as the viewer moves their head to look about the scene. Typically, this involves the display device sensing the sensor data to an external computer device (e.g. a computer connected to the display device via a wire). The external computer device may comprise powerful graphical processing units (GPUs) and / or computer processing units (CPUs) so that the external computer device is able to rapidly render appropriate two-dimensional images for the viewer based on the three-dimensional images and the sensor data. It will be appreciated that the use of a combination of a headset and an external device is exemplary. More generally, the processing of data and the rendering of images may be performed by various computer devices; for example, a standalone virtual reality headset may be provided, which headset is capable of processing data and rendering images without any connection to an external computer device. In some embodiments, the external computer device may comprise a server device, where the display device 17 may be connected to this server device wirelessly. This enables the two-dimensional images to be streamed from the server to the display device so as to enable the display of high-quality images without the need for a viewer to purchase expensive computer equipment. In other words, operations that require large amounts of computing power, such as the rendering of two-dimensional images based on the three-dimensional representation, may be performed by the server, so that the display device is only required to perform relatively simple operations. This enables the experience to be provided to a wide range of viewers. In some embodiments, a first two-dimensional image is provided to the display device 17 (and / or a connected device) and this first image is ‘warped’ in order to provide an image for viewing at the display device. The warping of the image comprises processing the image based on the sensor data in order to provide an image that matches a current viewpoint of the viewer. By performing the warping at the display device or another local device, the lag between a head movement of the user and an updating of the two-dimensional representation of the scene can be reduced. One issue with the above-described method of capturing a three-dimensional representation is that it only enables a viewer to make rotational movements. That is, since the points are captured using a single capture device at a single capture location, there is no possibility of enabling translational movements of a viewer through a scene. This inability to move translationally can induce motion sickness within a viewer, can reduce a degree of immersion of the viewer, and can reduce the viewer’s enjoyment of the scene. However, enabling a viewer to move translationally through a three-dimensional representation that has been captured using a single capture device would lead to holes in the scene wherever a viewer moves away from the capture location of this single capture device (since this movement will cause parts of the scene that were not captured by the capture device to come into view of the viewer). Therefore, it is desirable to enable translational movements through the scene while avoiding the display of these holes. To enable such movements, the three-dimensional representation of the scene may be captured using a plurality of capture devices placed at different locations (orthe same capture device placed at different locations). A viewer is then able to move around the scene translationally (e.g. by moving between these locations). More generally, by capturing points for every possible surface that might be viewed by a viewer, a three-dimensional representation of a scene may be captured that allows a suitable two-dimensional representation of this scene to be rendered regardless of a location of a viewer (e.g. regardless of where a user is standing within a virtual room). This need to capture points for every possible surface (so as to enable movement about a scene) greatly increases the amount of data that needs to be stored to form the three-dimensional representation. Therefore, as has been described in the application WO 2016 / 061640 A1, which is hereby incorporated by reference, the three-dimensional representation may be associated with a viewing zone, a zone of view (ZOV), or a zone of viewpoints (ZVP), where the three-dimensional representation is arranged to enable a user to move about the viewing zone so as to view the scene. Figure 4 illustrates such a viewing zone 1 and illustrates how the use of a viewing zone limits the amount of image data that needs to be stored to provide a three-dimensional representation of the scene. With the scene shown in this figure, and the viewing zone 1 shown in this figure, it is not necessary to determine attribute data for the occluded surface 2 since this occluded surface cannot be viewed from any point in the viewing zone. Therefore, by enabling the user to only move within the viewing zone (as opposed to around the whole scene) the amount of data needed to depict the scene is greatly reduced while still enabling the user to move to some extent and thereby to avoid the motion sickness that can be induced by augmented reality scenes with only three degrees of freedom. While Figure 4 shows a two-dimensional viewing zone, it will be appreciated that in practice the viewing zone 1 is typically a three-dimensional zone or volume. The viewing zone 1 may, for example, comprise a rectangular volume, or a rectangular parallelepiped, and the viewing zone may have a height of at least 30 cm, a depth of at least 30 cm, and / or a width of at least 30 cm, where these dimensions enable a userto move their head while remaining in the viewing zone. This is merely an exemplary arrangement of the viewing zone; it will be appreciated that viewing zones of various shapes and sizes may be used (e.g. spherical viewing zones). That being said, it is preferable that the viewing zone is limited so as to cover only a part of the volume of the scene, e.g. no more than 50% of the scene, no more than 25% of the scene, and / or no more than 10% of the scene. In this regard, if the viewing zone is the same size as the scene, then the three-dimensional representation will simply be a standard representation for virtual reality (that enables a user to move freely about the scene) - and so the use of the viewing zone will not provide any reduction in file size. The viewing zone 1 enables movement of a viewer around (a portion of) the scene. For example, where the scene is a room, the base representation may enable a userto walk around the room so as to view the room from different angles. In particular, the viewing zone enables a userto move through the scene with six degrees-of-freedom (6DoF) movement through the scene, where this aids in the provision of an immersive experience. In some embodiments, the viewing zone 1 may be four-dimensional, where a three-dimensional location of the viewing zone changes over time - and in such embodiments the size and location of the occluded surface 2 may also change over time. More generally, it will be appreciated that viewing zones may be formed in any size or shape, with different sizes and shapes being suitable for different scenes. The volume of the viewing zone 1 is typically selected so that a user is able to move to a degree sufficient to avoid motion sickness and to provide an immersive sensation, while still only enabling a limited amount of movement (where this leads to a smaller file size as compared to an implementation where a user is able to fully move about the scene). Typically, the viewing zone is arranged to enable a userto move their head while they are sitting or standing, but not to freely roam around a room. The viewing zone 1 may have a (e.g. real-world) volume of less than five cubic metres (5m3), less than one cubic metre (1 m3), less than one-tenth of a cubic metre (0.1 m3) and / or less than one-hundredth of a cubic metre (0.01m3). The viewing zone 1 may also have a minimum size, e.g. the viewing zone may have a volume of at least 1% of the volume of the scene, at least 5% of the volume of the scene, and / or at least than 10% of the volume of the scene. Similarly, the viewing zone may have a volume of at least one-thousandth of a cubic metre (0.01 m3); at least one-hundredth of a cubic metre (0.01 m3); and / or at least one cubic metre (1 m3). The ‘size’ of the viewing zone 1 typically relates to a size in the real world, where if the viewing zone has a length of one metre this means that a user is able to move one metre in the real world while staying within the viewing zone. The size of the viewing zone in the scene may be greater than, equal to, or less than the size of the viewing zone in the real world. For example, the viewing zone may scale a real-world distance so that moving one metre in the real world moves the user less than (or more than) one metre in the scene. This enables the scene to provide different perceptions to the user (e.g. to make the user feel larger or smaller than they are in real life). Similarly, the viewing zone may scale a real-world angle so that rotating one degree in the real world rotates the user less than (or more than) one degree in the scene. Therefore, a viewing zone with a volume of one cubic metre typically connotes a viewing zone in which the user is able to move about a one cubic metre volume in the real world while remaining in the viewing zone. And this may cause the user to move about a volume that is more than, or less than, one metre in the scene. Referring to Figure 5a, in order to capture points for each surface and location that is visible from the viewing zone 1, a plurality of capture devices C1, C2 C9 may be used (e.g. a plurality of virtual scanners and / or a plurality of cameras). Each capture device is typically arranged to perform a capture process, e.g. as described with reference to Figure 3, in which the capture device captures points at a plurality of azimuth angles and elevation angles. By locating the capture devices appropriately, e.g. by locating a capture device at each corner of the viewing zone, it can be ensured that all necessary points are captured. Typically, a first capture device C1 is located at a centrepoint of the viewing zone 1. In various embodiments, one or more capture devices C2, C3, C4, C5 may be located at the centre of faces of the viewing zone; and / or one or more capture devices C6, C7, C8, C9 may be located at edges of and / or corners of the viewing zone. Figure 5a shows a two-dimensional view (e.g. a plan view) of a rectangular viewing zone. It will be appreciated that within this viewing zone each capture device may be located on a shared plane. Equally, the various capture devices may be located on different planes. Referring, for example, to Figure 5b, there is shown a three-dimensional view of a cuboid viewing zone, where there is a capture device located: at the centre of the viewing zone (C1); at the centre of each face of the viewing zone (C2_A, C2_B, C2_C, C2_D, C2_E, C2_F); and at each corner of the viewing zone (C3_A, C3_B, C3_C, C3_D, C3_E, C3_F, C3_G, C3_H. With this arrangement, many locations in the scene (e.g. specific surfaces) will be captured by a plurality of capture devices so that there will be overlapping points relating to different capture devices. Typically, only a single version of the point is stored, where this version may be a highest quality version of the point and / or may be the version of the point associated with the nearest and / or least angled capture device. Camera scanning order In order to store the points of the three-dimensional representation, the points may be stored as a string of bits, where a first portion of the string indicates a location of the point (e.g. using x, y, z coordinates) and a second portion of the string locates an attribute of the point. In various embodiments, further portions of the string may be used to indicate, for example, a transparency of the point, a size of the point, and / or a shape of the point. A computer device that processes the three-dimensional representation after the generation of this representation is then able to determine the location and attribute of each point so as to recreate the scene. This location and attribute may then be used to render a two-dimensional representation of the scene that can be displayed to a viewer wearing the display device 17. Specifically, the locations and attributes of the points of the three-dimensional representation can be used to render a two-dimensional image for each of the left eye of the viewer and the right eye of the viewer so as to provide an immersive extended reality (XR) experience to the viewer. The present disclosure considers an efficient method of storing the locations of the points (e.g. at an encoder) and of determining the locations of the points (e.g. at a decoder). As has been described with reference to Figures 5a and 5b, the points of the three-dimensional representation are determined using a set of capture devices placed at locations about the viewing zone, where these capture devices are arranged to capture points at a series of azimuth angles and elevation angles. Typically, each of the capture devices is arranged to use the same capture process (e.g. the same series of azimuth angles and elevation angles), though it will be appreciated that different series of capture angles are possible. For example, there may be a plurality of possible series of capture angles, where different capture devices use different capture angles. In general, the present disclosure considers a method in which points are stored based on a capture device identifier and an indication of a distance of the point from the capture device associated with this capture device identifier. Typically, the point is also associated with an angular indicator, which indicates an azimuth angle and / or an elevation angle of the point relative to the identified capture device. By storing the point using a capture device index, an azimuth, an elevation, and a distance it is possible to efficiently provide information about howto properly depict this point to a device that has an awareness of the capture process (e.g. the increments used to capture points at the various elevation and azimuth angles). The angular indicator may be signalled in a number of forms. For example, the angular indicator may comprise an angle (e.g. in degrees), or the angular indicator may comprise an index that enables an angle to be determined based on this index and a known capture pattern (e.g. where the capture pattern involves capturing points at increments of 2° with a starting angle of 0°, then an index of 5 may be used to signal a capture angle of 8°). More generally, it will be appreciated that the storage of the distance and the angle may take many forms. For example, the distance and the angle of each point may be converted into a universal coordinate system, where each capture device has a different location in this universal coordinate system. In particular, each point may be stored with reference to a centre of this universal coordinate system, which centre may be co-located with a central capture device. Where a point is determined based on a distance and an angle from a capture device of a known location in this universal coordinate system, the coordinates of the point in this universal coordinate system can be determined trivially - and the location of the point may then be stored either relative to the capture device or as a coordinate in the universal coordinate system. The capture device identifier may comprise a location of a capture device (e.g. a location in a co-ordinate system of the three-dimensional representation). Equally, the capture device identifier may comprise an index of a capture device. Similarly, the indication of the azimuth angle and the elevation angle for a point may comprise an angle with reference to a zero-angle of a co-ordinate system of the three-dimensional representation. Equally, the azimuth angle and / or the elevation angle may be indicated using an angle index. In some embodiments, the three-dimensional representation is associated with configuration information, which configuration information comprises one or more of: a set of capture device indexes; locations associated with the capture devices and / or the capture device indexes; a spacing of capture devices (e.g. so that locations of the capture devices can be determined from a location of a first capture device and the spacing); angles associated with a capture process for the capture devices; an azimuth angle increment and / or an elevation angle increment associated with the capture process; and a set of angle indexes (e.g. to match an angle index to an angle). With this configuration information, it is possible to determine a location of each capture device from an index of that capture device and / or to determine a capture angle from a known capture process. Therefore, given two numbers: a capture device index and an angle index (that is associated with a combination of a specific azimuth angle and a specific elevation angle), a location of a capture device and a direction of a point from this capture device can be determined. By also signalling a distance of the point from the signalled capture device, a precise location of the point in the three-dimensional space can be signalled efficiently. Typically, the point is associated with each of: a camera index, a distance, an first angular index (e.g. a first azimuth), and a second angle (e.g. a second elevation) This method of indicating a location of a point enables point locations to be identified using a much smaller number of bits than if each point location is identified using x, y, z coordinates. Referring to Figure 6, there is shown a method of determining a location of a point. This method is carried out by a computer device, e.g. the image generator 11 and / or the decoder 15. In a first step 21, the computer device identifies an indicator of a capture device used to capture the point. Typically, this comprises identifying a portion of a string of bits associated with a capture device index. In a second step 22, the computer device identifies an indicator of an angle of the point from the capture device. Typically, this comprises identifying an angle index, e.g. an azimuth index and / or an elevation index and / or a combined azimuth / elevation index, which index(es) identifies a step of the capture process during which the point was captured. In a third step 23, based on the identifiers, the computer device determines the location of the capture device and the angle of the point from the capture device. The capture device identifier is typically a capture device index, which is related to a capture device location based on configuration information that has been sent before, or along with, the point data. For example, the configuration information may specify: Location of first capture device is (0,0,0). Step between capture devices is (0,0,1) along the grid, then across the grid, then up the grid. - The grid is (10,10,10). With this information, a capture device with an index of 1 can be determined to be located at (0,0,0); a capture device with an index of 5 can be determined to be located at (0,0,4); a capture device with an index of 12 can be determined to be located at (0,1,0), and so on. Equally, the configuration information may specify a list of camera indexes and locations associated with these indexes, where this enables the use of a wide range of setups of capture devices. Typically, the three-dimensional representation is associated with a frame of video. The configuration information may be constant over the frames of the video so that the configuration information needs to be signalled only once for an entire video. Therefore, the configuration information may be transmitted alongside a three-dimensional representation of a first frame of the video, with this same information being used for any subsequent frames (e.g. until updated configuration information is sent). The angle identifier may similarly be related to an angle by a location and an increment that are signalled in a configuration file. For example, the configuration information may specify: An azimuth increment and an elevation increment are each 1 °. There are 359 increments for each angle type. With this information: a capture angle with an index of 1 can be determined to be at an azimuth angle of 0° and an elevation angle of 0°; a capture angle with an index of 10 can be determined to be at an azimuth angle of 10° and an elevation angle of 0°; a capture angle with an index of 360 can be determined to be at an azimuth angle of 0° and an elevation angle of 1°; and a capture angle with an index of 370 can be determined to be at an azimuth angle of 9° and an elevation angle of 1°; etc. In a fourth step 24, based on the determined location of the capture device and the determined angle, a location of the point is determined. Typically, this comprises determining the location of the point based on the location of the capture device, the capture angle, and a distance of the point from the capture device (where this distance is specified in the point data for the point). Determining the location of the point typically comprises determining the location of the point relative to a centrepoint of the three-dimensional representation, This location of the point may then be converted into a desired coordinate system and / orthe point may be processed based on its location (e.g. to stitch together adjacent points). The angular identifier typically comprises a first angular identifier and a second angular identifier, where the first identifier provides the azimuthal angle of the point and the second identifier provides the elevation angle of the point. Referring to Figure 7, each angular identifier may be provided as an index of a segment of the three-dimensional representation, where, for example, an index of 0 may identify the point as being in a first angular bracket 101 and an index of 1 may identify the point as being in a second angular bracket 102. In this regard, the capture devices are arranged to perform a capture process, e.g. as described with reference to Figure 3, with a non-infinite angular resolution. Given this non-infinite resolution, each point is not a one-dimensional point located at a precise angle. Instead, each point is a point for a particular area of space, with the size of this area being dependent on the angular resolution as well as the distance of the point from the capture device. In other words, each capture angle determines a point for an angular range (with the range being dependent on the angular resolution). That is, if the capture process leads to points being captured at angles of 10°, 11°, and 12° then this can equally be considered to relate to points being captured at a first range of 9.5°-10.5°, a second range of 10.5°-11.5°, and a third range of 11.5°-12.5°. This is shown in figure 7, which shows a series of angular brackets, with the size of these angular brackets at a given distance being dependent on the angular resolution. The angular identifiers) typically comprise a reference to such an angular bracket. Consider, for example, a cube placed with the capture device C1 at the centre of this cube. By dividing this cube into x segments at regular azimuth angles and / segments at regular elevation angles, it is possible to identify any angular range of the representation by reference to an x segment and a y segment (and then the space bracketed by this angular range will depend on both the angular resolution (e.g. the angle between adjacent brackets) and the distance of the point from the capture device). Typically, each capture device has the same capture pattern so that the angular bracketing of each device is the same (albeit centred differently at the location of the relevant capture device). For example, in an embodiment with 1000 equal angular brackets, the angle for each bracket may be 360 / 1000. In some embodiments, different capture devices are associated with different capture patterns, where this may be signalled in configuration information relating to the three-dimensional representation. In some embodiments, each capture device is arranged to capture a point for a plurality of angular brackets, where each bracket is associated with a different angle. The angular spread of each bracket (that is, the angle between a first, e.g. left, angular boundary of the bracket and a second, e.g. right, angular boundary of the bracket) may be the same; equally, this angular spread may vary. In particular, the angular spread may vary so as to be smaller for points which are directly in front of (or behind, or to a side of) the capture device. For example, the embodiment shown in Figure 7 shows an angular bracketing system that is based on a cube. With this system, a cube is placed such that a capture device is located at the centre of the cube and the cube is then split into 1000 sections of equal size (it will be appreciated that the use of 1000 sections is exemplary and any number of sections may be used). Each of these sections is then associated with an angular index. With this arrangement, the angular spread of each section (or bracket) varies, as has been described above. Figure 7 shows a two-dimensional square, where each angular bracket of the square is referenced by an index number (between 1 and 100). In a three-dimensional implementation, an angular bracket of a cube could be indicated with two separate numbers (with a first azimuthal indicator that identifies a ‘column’ of the cube and a second elevational indicatorthat identifies a ‘row’ of the cube). Equally, a singular indicator may be provided that indicates a specific bracket of the cube. Therefore, for a cube that is divided into 1000 elevational sections and 1000 azimuthal sections, the bracket may be indicated with two separate indicators that are each between 0 and 999 or with a single indicatorthat is between 0 and 999999. It will be appreciated that the use of a cube to define the brackets is exemplary and that other bracketing systems are possible. For example, a spherical bracketing system may be used (where this leads to curve angular brackets). Equally, a lookup table may be provided that relates angular indexes to angles, where this enables irregularly spaced brackets to be used. Typically, determining the location of the point comprises determining the location of the point so as to be at the centre of the angular bracket identified by the angular identifier(s). Referring then to Figure 8, there is described a method of determining the size of the point based on the angular identifiers. In this regard, a size of the point is typically determined based on the distance of the point to the capture device as well as the angular identifiers, where the point may be determined to extend across an entire width of an angular bracket. In a first step 31, a computer device determines a location of a point, e.g. using the method of Figure 7. For example, given a capture device location of (a,h), an angle of 0, and a distance of D, the point can be determined as being located at the coordinate (a + Dsin(0), b + Dcos(0)). It will be appreciated that this is a two-dimensional example and typically the point is located at a three-dimensional coordinate (where this coordinate may depend on the same distance and a further angle 0). For example, where the capture device is located at a position (a,b,c), the point may be determined as being located at the coordinate (a + Dcos(0)sin(0), b + Dcos(^cos(9), c + Dsin(0)) As described above, the ‘angle’ provided is typically an indication of an angular bracket, where the point can be determined as being located at the centre of the angular bracket. The determination of the angular bracket may depend on the bracketing system used (e.g. with a spherical bracketing system, the angular bracket may be determined by the equation: centre of angular bracket = index* angular step, with the cubic bracketing system described with reference to Figure 7, each index may be associated with a known angle). So that adjacent points fit together, in a second step 32, the computer device determines an angle of a first angular boundary associated with the point and in a third step 33 the computer device determines an angle of a second angular boundary associated with the point. Then, in a fourth step 34, the computer device determines a size (e.g. a coverage) of the point based on the first angular bracket and the second angular bracket. This typically comprises determining a size (e.g. a coverage) of the point such that the point extends between the first angular boundary and the second angular boundary. The fourth step 34 may, additionally and / or alternatively, comprise determining coordinates of one or more corners of the point, where this enables a quadrilateral area covered by a point to be determined based on the configuration information and the point data. Where a ‘size’ may be suitable for signalling a regular shape (with a consistent width and height) covered by a point, such as a square or a rectangle, this size may be less suitable for signalling a point that covers a shape with a varying width or height, such as a trapezium. The fourth step 34 may comprise aligning (or overlapping) the corners of a first point in a first angular bracket with the corners of a second point in a second angular bracket (e.g. to align a first point in the first bracket with a second point in a second bracket so as to form a continuous surface). Referring to Figures 9a and 9b, two different methods of determining the size and / or corner coordinates of the point are described. Referring first to Figure 9a, we consider an example of a point that lies in the second angular bracket 101 for a capture device C1. This angular bracket covers an angle p, which angle depends on an angular resolution of the three-dimensional representation. As described above, the method of determining the point typically involves determining a size orcoverage of this point, where the size may be dependent on the first angular bracket and the second angular bracket. In particular, the point may be sized so as to extend between these brackets. In the embodiment of Figure 9a, the coordinates of the corners of the point are determined based on the angle between the capture device and the point. In particular, the coordinates of the point may be determined so that the plane is perpendicular to the line from the capture device to the centre of the point. Where the point is extended so as to face the capture device, the point will extend equal amounts in both directions between the nominal location of the point and the first and second angular boundaries, which can simplify the signalling / determination of the point. The nominal location of the point is the location of the point based on the determined distance and angle of the point from the capture device. With the example of Figure 9a, the centrepoint of the point can be determined as (Dsni(0),Deos(0) and the coordinates of the corners of the point can then be determined as (Dsin(0) - x,Dcos(0) + y and (Dsin(0) + x,Dcos(0) - y, where the specific values of x and y depend on the values of: D (the distance of the (nominal) point from the capture device); 9 (the angle of the (nominal) point from the capture device); and p (the angle of the angular bracket, which depends on the angular resolution of the representation). However, such an implementation can result in points extending in a different direction to an actual surface on which these points lie. In this regard, each point is typically associated with a normal vector, which normal vector is normal to a surface on which the point lies. In order to obtain an accurate depiction of the three-dimensional representation, the coordinates of the corners of the point may be determined based on this normal vector. More specifically, the coordinates of the point may be determined so that the point extends between the first angular boundary and the second angular boundary perpendicularly to the normal vector of the point. As shown in Figure 9b, this leads to coordinates that are not equidistant from the point. As can be seen from Figure 9b, with such an implementation the coordinates of the corners depend on the values of D, 0, p, and a, where a is the angle between the normal ofthe point and the angle of the (nominal) point from the capture device. More generally, the coordinates ofthe corners ofthe point may depend on the normal ofthe point. With such an implementation, the corners are not located equidistantly from the nominal point, but instead the corners can be determined as (£>sin(0) - x,,Dcos(0) + yP and (Dsin(0) + x,:, Dcos(0) — yP). Such a method of determining the corners of a point can be used to lead to a representation in which the adjacent points are substantially aligned. In this regard, determining the corner locations based solely on D, 0, and p, can lead to incongruities where adjacent points are captured by different capture devices. By also determining the corner locations based on a these incongruities can be minimised. It will be appreciated that the examples of Figures 9a and 9b show a two-dimensional example. Typically, the point is located at a three-dimensional location in space. In this case, the locations ofthe corners will be dependent on each of: a distance of the point from the capture device; an azimuthal angle of the point from the capture device; an elevational angle of the point from the capture device; an azimuthal angle of the bracket; an azimuthal angle of the bracket; and (optionally) the normal of the point. As shown in Figure 10, using the method of Figure 9b, in which the corner coordinates are determined based on: a capture device identifier, a distance of the point from the capture device, one or more (e.g. two) angular identifiers, an angular resolution of the representation, and a normal of the point, it is possible to determine a two-dimensional quadrilateral that is associated with the point. And it should be apparent that (particularly for large / dense three-dimensional representations for which the coordinates are precisely located) this method of signalling a two-dimensional quadrilateral is substantially more efficient than a method of signalling each corner coordinate separately. For example, the capture device location may be signalled with a high degree of accuracy (e.g. using 16 bits or more) in the configuration file, with the point data then requiring a capture indexthat may be indicated using 8 bits (for an embodiment with 255 capture device locations). The distance may then be signalled using, for example, a 15 bit string and the angular identifiers may be signalled using a 12 or 13 bit string to provide an angular resolution of 360 / 4096 (0.08789° per bracket) or 360 / 8192 (0.04394° per bracket). The normal may be encoded as two 8 bit vectors using octahedron normal encoding. In some embodiments, the three-dimensional representation is signalled in the form of a cubemap, where the angular identifies may then have different lengths (e.g. the cubemap may have 3 faces horizontally and 2 faces vertically, so that the azimuthal angular identifier may be of a greater length than the elevational angular identifier). As described above, the size of the point may be determined based on the angular resolution of the capture process, where the point is determined to extend between two angular brackets, with the location of these angular brackets depending on the angular resolution. With the examples given, the point is extended from a centrepoint of a bracket until it intersects with a first angular boundary and a second angular boundary of this bracket (or, in three-dimensional implementations, with each of a first to fourth angular boundary). These angular boundaries may be spaced symmetrically about the centrepoint (so that the distance between the first angular boundary and the nominal point location is the same as the distance between the nominal point location and the second angular boundary), but this is not necessary. More generally, the size of the point is determined so that the point extends from the nominal location of the point (e.g. the point has a width and / or a height). Typically, the point is determined to have four corner coordinates that are located on the angular boundaries of a bracket, but different numbers of corner coordinates or different methods of projection may be used. For example, the point coordinates may be determined so that they extend beyond these angular boundaries in order to cause adjacent points to have an overlap. In some embodiments, the points are associated with a size, which size indicates a number of brackets covered by the points. For example, if a point has a size of 2, this may indicate that the point extends over two angular brackets. The corner coordinates of such a point may still be determined based on the factors mentioned above (albeit instead of being located on the second angular boundary of the second angular bracket 102, the second corner would be located on a second angular boundary of an adjacent bracket). The configuration indication may indicate a location of extension of points. That is, if a point has a size of 2, the configuration information may indicate that this point extends in a positive azimuthal direction (so from the second angular bracket 120 into a third angular bracket as opposed to from the second angular bracket 120 into the first angular bracket 101). It will be appreciated that various methods of signalling a direction of extension are possible. In various embodiments, the size of the point comprises a width (e.g. an azimuthal size) and / or a height (e.g. an elevation size). The size may be signalled via a size index that is associated with a list of sizes. For example, the maximum possible width and height may be 3 angular brackets; such an implementation provides 9 options for the size of a point: 1x1,2x1, 3x1, 1x2, 2x2, 3x2,1x3, 2x3, 3x3. This size option can then be signalled using 4 bits to signal an index of the size from the list. It will be appreciated that other options for indicating the size are possible. The size may be stored based on an underscan value. In this regard, where an object is very near to the viewing zone it may be captured using an unnecessarily dense arrangement of points. Therefore, certain surfaces or areas of the representation may be associated with an underscan value, which underscan value defines a reduction in the number of points captured as compared to a representation without underscan. The size of the points may be defined so as to indicate this underscan value. In an exemplary embodiment, the underscan value is an integer value between 0 and 3 and the size is stored as a combination of point dimensions (e.g. a width in the range [0,2]) and a height in the range ([0,2]) and an underscan factor (e.g. an underscan factor in the range [0,3]). In some embodiments, the width and the height are dependent on the underscan factor. For example, when the underscan factor exceeds a threshold value, the possible height and width values may be limited. In a specific example, when the underscan factor is 3, the width and the height may be limited to the range [0,1]. The size may then be defined as size = underscan*9 + height*3 + width. Such a method provides efficient storage and indication of width, height, and underscan values. Returning to Figure 10, based on the factors and the methods described above, the computer device is able to determine coordinates of four corners of a quadrilateral associated with the point. Noting that the point is also associated with an attribute and / or a transparency, the computer device is then able to render a depiction of the point that occupies an area on a two-dimensional plane of a three-dimensional space, the area being a quadrilateral. This rendering may involve rendering a plurality of triangles that cover the two-dimensional quad. By rendering a plurality of points (e.g. quadrilaterals and / or triangles) on a plurality of two-dimensional planes located at different distances from the viewing zone, the computer device is able to provide a depiction of the three-dimensional representation. This depiction may be generated in the form of two two-dimensional images that are displayed to (respective eyes of) a viewer at the display device 17 to provide an extended reality experience to this viewer. Referring to Figure 11, there is shown a three-dimensional representation of a point that lies within an angular bracket, the angular bracket having a first azimuthal angular boundary (between the angular lines B1 and B3), a second azimuthal angular boundary (between the angular lines B2 and B3), a first elevational angular boundary (between the angular lines B1 and B2), and a second elevational angular boundary (between the angular lines B3 and B4). As shown in Figure 11, the method includes extending a plane from the determined centre of the point in a direction perpendicular to the normal N1 of the point and then determining the coordinates of the corners C1, C2, C3, C4 of the point based on the intersections of this plane with the angular boundaries. It will be appreciated that performing a similar method for a plurality of points of the representation (e.g. for points in each angular bracket) enables a computer device to determine, and then render, an image or a video based on the three-dimensional representation, which image or video presents a number of continuous surfaces. Alternatives and modifications It will be understood that the present invention has been described above purely by way of example, and modifications of detail can be made within the scope of the invention. The scene typically comprises an extended reality (XR) scene, where this scene comprises a real or digital scene that contains one or more digital elements. The term extended reality (XR) covers each of virtual reality (VR), augmented reality (AR), and mixed reality (MR) and it will be appreciated that the disclosures herein are applicable to any of these technologies. In some embodiments, the scene and / or the points may be generated or processed using an artificial intelligence (Al) or machine learning (ML) model, where this may involve generating the Al or ML model on a powerful computer device before transmitting the model (e.g. the weights of a neural network) to a less powerful device. This can enable the generation of high quality representations on a device with limited 5 processing power. Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.

Claims

1. A method of determining a location of a point in a three-dimensional representation of a scene, the method comprising:identifying point data associated with the point;determining, from the point data:a capture device identifier associated with a capture device used to capture the point;a distance of the point from the capture device; andone or more angle identifiers that indicate an angle of the point from the capture device; and determining the location of the point based on the distance, the capture device identifier, and the one or more angle identifiers.

2. The method of any preceding claim, wherein the capture device identifier comprises a capture device index.

3. The method of any preceding claim, comprising determining a location of the capture device based on the capture device identifier and configuration information associated with the representation, preferably wherein determining the location of the capture device comprises comparing the capture device index to a table of capture devices located in the configuration information.

4. The method of any preceding claim, comprising determining a size and / or coverage of the point, preferably wherein the size and / or coverage of the point is dependent on an angular resolution of the three-dimensional representation.

5. The method of any preceding claim, wherein determining the coverage comprises determining coordinates for a plurality of corners of the point, preferably four corners.

6. The method of claim 4 or 5, comprising determining the coverage and / or the coordinates of the corners of the point in dependence on the angle of the point from the capture device and the distance of the point from the capture device.

7. The method of any of claims 4 to 6, comprising determining the coverage and / orthe coordinates of the corners of the point by extending a plane between two angular brackets associated with the capture device, the plane preferably being extended in a direction perpendicular to the angle of the point from the capture device.

8. The method of any of claims 4 to 7, comprising determining the coverage and / orthe coordinates of the corners of the point in dependence on a normal of the point from the capture device, preferably determining the coverage and / or the coordinates of the corners of the point by extending a plane between two angular brackets associated with the capture device, the plane preferably being extended in a direction perpendicular to the normal of the point.

9. The method of any preceding claim, comprising determining the coordinates of the corners of the point by:determining an angular bracket associated with the point and the angle of the point from the capture device, the point lying within this angular bracket, and the angular bracket being formed of a plurality of angular boundaries;determining a location of a centre of the point based on the angle and the distance;determining a normal of the point from the point data; andextending a plane from the location of the centre of the point to each of the angular boundaries, the plane being extended in a direction perpendicular to the normal of the point from the capture device; anddetermining the corners of the point as the intersections of the plane and the angular boundaries.

10. A method of determining the coordinates of corners of a point in a three-dimensional representation of a scene, the method comprising:determining a capture device used to capture the point;determining an angle of the point from the capture device;determining a distance of the point from the capture device;determining an angular bracket associated with the capture device and the angle, the point lying within this angular bracket, and the angular bracket being formed of a plurality of angular boundaries;determining a location of a centre of the point based on the angle and the distance;determining a normal of the point; andextending a plane from the location of the centre to each of the angular boundaries, the plane being extended in a direction perpendicular to the normal of the point from the capture device; anddetermining the corners of the point as the intersections of the plane and the angular boundaries.

11. The method of any preceding claim, wherein the three-dimensional representation is associated with a frame of a video, the video comprising a plurality of frames and each frame being associated with a separate three-dimensional representation of the scene, wherein the method comprises:receiving, at a first time, the configuration information;receiving, at a second time, point data for a plurality of three-dimensional representations; anddetermining the locations ofa plurality of points in different three-dimensional representations using the configuration information and the point data.

12. The method of any preceding claim, wherein:the one or more angle identifiers comprise an azimuth angle identifier and / or an elevation angle identifier; and / orthe one or more angle identifiers comprise an azimuth angle and / or an elevation angle.

13. The method of any preceding claim, wherein the one or more angle identifiers comprise a reference to an angular bracket of the three-dimensional representation, the angular bracket covering a volume of a space of the three-dimensional representation.

14. The method of claim 13, wherein the size of the angular bracket is dependent on an angular resolution of the three-dimensional representation.

15. The method of claim 13 or 14, wherein the angular bracket is associated with one or more of:a first azimuthal angular boundary and a second azimuthal angular boundary; anda first elevational angular boundary and a second elevational angular boundary.

16. The method of any preceding claim, wherein the capture device and / or the three-dimensional representation is associated with a plurality of angular brackets, preferably wherein the angle of each angular bracket is defined by configuration information for the three-dimensional representation.

17. The method of claim 16, wherein the configuration information identifies an angular resolution, an azimuthal resolution, and / or an elevational resolution for the capture device and / or the three-dimensional representation.

18. The method of claim 16 or 17, wherein the configuration information defines a relationship between the one or more angle identifiers and a corresponding angle, preferably between the identifier and each of an azimuth angle and an elevation angle.

19. The method of any preceding claim, wherein the point is associated with a quadrilateral in the three-dimensional representation.

20. The method of any preceding claim, comprising rendering an image based on: the location of the point and / or the location of the corners of the point; and / or point data for a plurality of points of the three-dimensional representation; preferably, wherein the image is associated with an extended reality (XR) and / or virtual reality (VR) experience.

21. The method of any preceding claim, wherein the point data comprises a size, the size defining a number of angular brackets covered by the point, preferably wherein the size comprises a size index that defines the size with reference to an index table.

22. The method of any preceding claim, wherein the three-dimensional representation is associated with a viewing zone, the viewing zone comprising a subset of the scene, wherein the viewing zone enables a user to move through a subset of the scene, preferably wherein the viewer is able to move within the viewing zone with six degrees of freedom (6DoF);preferably, wherein:the viewing zone has a volume of less than 50% of the volume of the scene, less than 20% of the volume of the scene, and / or less than 10% of the volume of the scene; and / orthe viewing zone has, or is associated with, a volume, preferably a real-world volume, of less than five cubic metres (5m3), less than one cubic metre (1 m3), less than one-tenth of a cubic metre (0.1m3) and / or less than one-hundredth of a cubic metre (0.01m3).

23. A method of determining point data for a point in a three-dimensional representation of a scene, the method comprising:identifying a capture device;determining a distance of a point from the capture device, preferably wherein the point is a point on a surface of an object;determining an angle of the point from the capture device; anddefining point data for the point, the point data including:a capture device identifier indicating the capture device used to capture the point; distance data indicating the distance of the point from the capture device; and one or more angle identifiers indicating the angle of the point from the capture device;preferably wherein the method comprises determining an attribute value for the point, the attribute value being included in the point data.

24. An apparatus for determining a location of a point in a three-dimensional representation of a scene, the apparatus comprising:means for identifying point data associated with the point;means for determining, from the point data:a capture device identifier associated with a capture device used to capture the point;a distance of the point from the capture device; andone or more angle identifiers that indicate an angle of the point from the capture device; andmeans for determining the location of the point based on the distance, the capture device identifier, and the one or more angle identifiers.

25. An apparatus for determining point data for a point in a three-dimensional representation of a scene,5 the apparatus comprising:means for identifying a capture device;means for determining a distance of a point from the capture device;means for determining an angle of the point from the capture device; andmeans for defining point data for the point, the point data including:10 a capture device identifier indicating the capture device used to capture the point;distance data indicating the distance of the point from the capture device; andone or more angle identifiers indicating the angle of the point from the capture device;preferably comprising means for determining an attribute value for the point, the attribute value being included in the point data.

Citation Information

Patent Citations

  • Backfilling Points in a Point Cloud

    US20140132733A1

  • Distance measurement system

    US20230341558A1

  • Encoding and decoding point data identifying a plurality of points in a three-dimensional space

    WO2024002462A1