Application of hierarchical coding in splitting calculation
By using hierarchical encoding technology in distributed rendering networks, the problems of finite frame rates and delays are solved, and efficient 3D video rendering and transmission are achieved.
Patent Information
- Application Number
- CN202380060588.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-07-15
- Filing Date
- 2023-06-30
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art faces the problems of limited frame rates and delays between image generation and display when generating and displaying 3D videos, especially in the case of limited bandwidth and insufficient computing power, making it difficult to achieve high-quality 3D video rendering and transmission.
Using a distributed rendering network, a first partial rendering frame sequence is generated and encoded through the first rendering node and transmitted to the second rendering node. The second rendering node generates a second partial or fully rendered frame sequence based on the received coded frame sequence, thereby achieving an increase in frame rate and a reduction in delay.
Through the combination of hierarchical encoding and distributed rendering network, the efficiency of video streams is improved, supported use cases are enhanced, the latency of the rendering process is reduced, and bandwidth usage is optimized, achieving high-quality 3D video rendering and transmission.
Smart Images

Figure CN120035841A_ABST
Abstract
Description
Technical Field
[0001] The following disclosure relates to a system for generating and displaying an image on a display device. The image may, for example, virtually represent an object in a 3D space. The display device may, for example, be a user extended reality ("XR", including virtual reality and / or augmented reality) headset, XR smart glasses, an autostereoscopic display, a television display, a mobile device, a PC, etc. Background Art
[0002] According to so-called "split computing" or "remote rendering", the image is usually generated remotely from the display device, and there are usually limitations on the communication speed and capacity between the image source and the display device. For example, the image source and the display device may be connected via a network, and the image source may be located, for example, on a separate device in the same room, or on a server in a nearby private data center, or on a server in the so-called cloud.
[0003] Due to the limitations of communication speed and capacity, as well as the processing speed and processing resources required for image generation, it is desirable to reduce factors such as frame rate and amount of data transmitted as much as possible without affecting the user experience. In addition, bandwidth may vary over time, and when bandwidth decreases, the amount of data transmitted may need to be reduced in a timely manner to avoid so-called "delay jitter". In the art, the resolution and / or frame rate of a video sequence transmitted to a display device is usually adjusted, although the response time of the changed resolution and frame rate usually lags behind the sudden drop in bandwidth (especially in the case of wireless transmission), and a relatively large I frame needs to be sent each time the resolution changes (e.g., an intraframe independent of the previous frame), which further increases the risk of delay jitter in unreliable data channels.
[0004] Additionally, because of the time required to generate the image and transmit it to the display device, there may be a difference between the display required when the image is generated (the "viewport") and the display required when the image is displayed on the display device. As a non-limiting example, when the display device is a user headset, the user may move his head intentionally or unintentionally.
[0005] These issues are particularly challenging when high quality 3D video (e.g., 3D meshes or point clouds) is desired, such as in the near-realistic graphics of modern video games, and when the headset is required to have low weight and low power consumption (e.g., 1 Watt target power consumption).
[0006] One known technique for dealing with limited frame rates and delays between image generation and image display is called reprojection, warping, or time warping.
[0007] Additionally, it is known to use depth maps to assist in representing 3D space. A depth map represents how far away different surfaces of a 3D object or different parts of an image should appear to be from the viewer. Depth maps can be used for temporal warping (parallax can be taken into account in such depth-assisted cases, and are therefore often referred to as "spatial warping") as well as other real-time display adjustments, such as zoom correction based on eye tracking. Although depth maps are often used to assist in spatial warping when local rendering is performed within the same device and accurate depth information is directly available, it is currently challenging to efficiently transmit compressed depth information with sufficient bit depth granularity when rendering is performed according to split computation methods, particularly due to scarce bandwidth, limited computing power, and the fact that existing hardware acceleration methods for compressing image and video data provide a maximum pixel data precision of 10-bit values, which are therefore relatively inaccurate for the purposes of depth information.
[0008] By improving the efficiency of the video streaming portion of the split computation, it is envisioned that the use cases supported by split computation will increase, highlighting the need to make the rendering process as efficient as possible. It has been proposed that the rendering process itself could be further split across multiple servers that are not necessarily co-located, under so-called "hybrid rendering", such that expensive computation is performed on one server and viewport rendering is then completed on another low-power device closer to the user. A problem known in the art is how to properly perform hybrid rendering and then efficiently compress the intermediate byproducts and transmit them to the device that performs the final rendering of the viewport. Summary of the invention
[0009] According to a first aspect, the present disclosure provides a networked system for generating a frame sequence for rendering a dynamic 3D scene, the system comprising a first rendering node and a second rendering node, wherein: the first rendering node is configured to: generate a first partial rendering frame sequence; perform encoding (specifically, as a non-limiting example, hierarchical encoding) on each first partial rendering frame to generate an encoded first partial rendering frame sequence; transmit the encoded first partial rendering frame sequence to the second rendering node; and the second rendering node is configured to: obtain the encoded first partial rendering frame sequence from the first rendering node; and generate a second partial or complete rendering frame sequence based on the encoded first partial rendering frame sequence. For example, the first rendering node can be a node of a distributed rendering network, and the second rendering node can be another node of the distributed rendering network, or can be a user device, such as a VR headset. The first rendering node can be a central node, and the second rendering node can be an edge node. The rendering node can serve multiple users, and the first rendering node can serve more users than the second rendering node.
[0010] Optionally, for the first aspect, the second rendering node is configured to decode the encoded first partial rendering frame sequence to obtain the first partial rendering frame sequence.
[0011] Optionally, for the first aspect, the second rendering node is configured to perform layered encoding on each second partially or fully rendered frame to generate an encoded sequence of second partially or fully rendered frames.
[0012] In one embodiment, the first rendering node is configured to perform layered encoding according to a first encoding scheme and the second rendering node is configured to perform layered encoding according to a second encoding scheme, the first encoding scheme being different from the second encoding scheme.
[0013] Optionally, for the first aspect, the frame rate of the second partially or fully rendered frame sequence is greater than the frame rate of the first partially rendered frame sequence. In a non-limiting embodiment, this is achieved by partially rendering a frame that includes data to be processed when performing frame rate interpolation to support more accurate interpolation.
[0014] Optionally, for the first aspect, the first rendering node generates a first partially rendered frame sequence including a first data type, and the second rendering node generates a second partially or completely rendered frame sequence including a second data type, wherein the first data type is different from the second data type. For example, the first data type may be point cloud data, and the second data type may be image data, wherein the image data is generated at least in part based on the point cloud data.
[0015] Optionally, for the first aspect, the system includes a display device, wherein a communication delay between the first rendering node and the display device is greater than a communication delay between the second rendering node and the display device.
[0016] In one embodiment: the first rendering node is configured to obtain a viewing position from a display device before generating a first partially rendered frame sequence; and the second rendering node is configured to obtain an updated viewing position from the display device before generating the second partially or fully rendered frame sequence.
[0017] Optionally, for the first aspect, generating the first partial rendered frame sequence requires more processing resources than generating the second partial or full rendered frame sequence.
[0018] Optionally, for the first aspect, each frame includes image data and depth map data, and the system includes a third rendering node configured to generate a third fully rendered frame sequence by performing time warping and / or depth correction on the second partially or fully rendered frame sequence using the depth map data. For example, the first rendering node and the second rendering node may be nodes of a distributed rendering network, and the third rendering node may be a user device, such as a VR headset. The first rendering node may be a central node, and the second rendering node may be an edge node. The rendering node may serve multiple users, and the first rendering node may serve more users than the second rendering node.
[0019] Optionally, for the first aspect, each frame comprises point cloud data, wherein each point of the plurality of points has a 3D position and one or more attributes.
[0020] When each frame includes point cloud data, the second rendering node or the third rendering node may be configured to calculate depth map data based on the 3D positions of the points.
[0021] Optionally, for the first aspect, the first rendering node is configured to generate one or more first partially rendered frame sequences for a first number of users or display devices; and the second rendering node is configured to generate one or more second partially or fully rendered frame sequences for a second number of multiple users or display devices, wherein the second number is less than the first number.
[0022] Optionally, for the first aspect, the networking system dynamically selects how many nodes and which specific nodes to use for the rendering process in response to at least one metric, the at least one metric including the complexity of the rendering task to be performed, the spare capacity available at each node, the location of the display device relative to the node network, the round-trip delay between the node and the display device, the available bandwidth between nodes and from the node to the display device, and a metric of the number of different display devices requesting to render the same 3D scene within a certain viewpoint range.
[0023] Optionally, for the first aspect, the partially rendered frame is encoded as volumetric data, including by way of non-limiting example point cloud data, mesh data, and / or texture data, allowing a subsequent rendering node to render multiple viewpoints within a range ("viewport"), and the partially rendered frame is used by one or more subsequent rendering nodes to generate a fully rendered frame for at least two users. This allows processing-intensive aspects of the scene to be computed only once for multiple users.
[0024] Optionally, for the first aspect, the partially rendered frame includes light field data, so that those environment characteristics need only be calculated once for multiple users in the same virtual environment.
[0025] Optionally, for the first aspect, the partially rendered frame includes spatial characteristics, enabling the behavior of sound to be calculated in 3D space, so that those environmental characteristics need only be calculated once for multiple users in the same virtual environment.
[0026] Optionally, for the first aspect, the partially rendered frame is encoded by at least partially using a lossy encoding method.
[0027] Optionally, for the first aspect, the partially rendered frame is encoded by at least partially using a layered encoding method.
[0028] Optionally, for the first aspect, the subsequent node receives only a portion of the data encoded by the first rendering node in response to a particular position of one or more viewpoints of the fully rendered frame to be calculated.
[0029] Optionally, for the first aspect, the subsequent node decodes only the subset of the encoded data produced by the first rendering node and received by the subsequent node that is necessary to fully render the particular field of view it is rendering at any point.
[0030] Optionally, for the first aspect, the partially rendered frame is encoded at least in part using a point cloud format that represents points according to one or more coordinate systems, each point being assigned one or more data attributes specifying visual characteristics, the visual characteristics including one or more of size, normal vector, motion information, color information, and transparency.
[0031] According to aspects related to the first aspect, the present disclosure provides a method for generating a frame sequence for rendering a dynamic 3D scene, wherein: a first rendering node of a networked system: generates a first partial rendering frame sequence; performs encoding (specifically, as a non-limiting example, hierarchical encoding) on each first partial rendering frame to generate an encoded first partial rendering frame sequence; and transmits the encoded first partial rendering frame sequence to a second rendering node of the networked system; and the second rendering node: obtains the encoded first partial rendering frame sequence from the first rendering node; and generates a second partial or full rendering frame sequence based on the encoded first partial rendering frame sequence. The optional features of the first aspect may be applied to related aspects.
[0032] According to aspects related to the first aspect, the present disclosure provides a bitstream comprising a first partially rendered frame sequence that is encoded. The bitstream is generated by a first rendering node by: generating a first partially rendered frame sequence; and performing encoding (specifically, as a non-limiting example, hierarchical encoding) on each first partially rendered frame to generate an encoded first partially rendered frame sequence. The bitstream is suitable for a second rendering node to generate a second partially or completely rendered frame sequence based on the encoded first partially rendered frame sequence. The optional features of the first aspect may be applied to related aspects.
[0033] According to a second aspect, the present disclosure provides a method for encoding a frame sequence representing a dynamic 3D scene, wherein each frame includes base layer image data and enhancement data, the method comprising: performing layered encoding on frames in the frame sequence to generate an encoded frame, the encoded frame including a base image layer and an enhancement image layer.
[0034] Optionally, for the second aspect, the enhancement picture layer comprises data used by the decoder device to reconstruct a higher resolution rendition of the frame sequence.
[0035] Optionally, for the second aspect, the enhancement picture layer comprises data used by said decoder means to reconstruct a higher bit depth rendition of the sequence of frames relative to the bit depth of the base picture layer.
[0036] Optionally, for the second aspect, the enhancement picture layer comprises data used by the decoder means to reconstruct distances of objects in the picture from a viewer.
[0037] Optionally, for the second aspect, the enhancement image layer comprises data used by the decoder device to reconstruct tactile feedback for the user.
[0038] Optionally, for the second aspect, the method further includes: in response to a decrease in the bandwidth of the transmission channel, discarding the enhancement data during transmission of the frame sequence and indicating to the encoder that the enhancement data has been discarded, and if the enhancement data has been discarded, refreshing the time buffer for encoding the enhancement data and performing an instantaneous decoder refresh (IDR) on the enhancement data (not necessarily on the base layer data) to account for the decoder losing some previous enhancement data.
[0039] Optionally, for the second aspect, the layered coding method used is MPEG-5 LCEVC (Low Complexity Enhancement Video Coding) coding or SMPTE VC-6 coding.
[0040] Optionally, for the second aspect, at least one of the enhancement data is transmitted as embedded user data within a coefficient of the LCEVC data. As a non-limiting example, one or more residual coefficients of an image include embedded depth information representing a depth of a corresponding object relative to a viewpoint, and the decoder processes the embedded data to reconstruct a depth map associated with the image frame based at least in part on the embedded data. In a non-limiting embodiment, the depth map is reconstructed by processing the embedded data and the image data.
[0041] Optionally, for the second aspect, or as a separate implementation, frames plus depth data sent at a given frame rate are used by the display device to increase the frame rate through spatial warping (depth-based reprojection) to match the display frame rate, thereby enabling rendering and video streaming processing to be performed at a frame rate lower than the display frame rate and reducing the required bandwidth and processing resources.
[0042] Optionally, for the second aspect, each frame comprises image data and depth map data, and the method comprises performing layered encoding on frames in the frame sequence to generate encoded frames comprising one or more of a base depth layer and an enhanced depth layer.
[0043] Optionally, for the second aspect, each frame includes an image layer and a depth layer, and the coded frame includes a base image layer, a base depth layer, an enhanced image layer and an enhanced depth layer.
[0044] Optionally, for the second aspect, each frame comprises depth map data embedded in the image data, the base depth layer is a base image layer having the embedded depth map data, and the enhancement depth layer is an enhancement image layer having the embedded depth map data.
[0045] Optionally, for the second aspect, the method further includes: receiving a depth map discard indication indicating whether to discard depth map data when transmitting the frame sequence; and if the depth map data is discarded and the available bandwidth is lower than a threshold, discarding the depth map data of the frame, and performing layered encoding on the image data of the frame to generate an encoded frame including a base image layer and an enhanced image layer.
[0046] According to aspects related to the second aspect, the present disclosure provides an encoder configured to encode a frame sequence representing a dynamic 3D scene, wherein each frame includes base layer image data and enhancement data, and the encoder is specifically configured to perform hierarchical encoding on frames in the frame sequence to generate an encoded frame, wherein the encoded frame includes a base image layer and an enhancement image layer. The optional features of the second aspect can be applied to related aspects.
[0047] According to an aspect related to the second aspect, the present disclosure provides a bitstream comprising a sequence of coded frames representing a dynamic 3D scene, wherein each frame comprises a base picture layer and an enhancement picture layer.
[0048] According to a third aspect, the present disclosure provides a method for decoding a frame sequence representing a dynamic 3D scene, wherein each frame includes base layer image data and enhancement data, the method comprising: performing layered decoding on the frames in the frame sequence to generate a decoded frame from a base image layer and an enhancement image layer.
[0049] Optionally, for the third aspect, the enhancement picture layer comprises data used in a decoding method to reconstruct a higher resolution rendition of the frame sequence.
[0050] Optionally, for the third aspect, the enhancement picture layer comprises data used in the decoding method to reconstruct a higher bit depth rendition of the frame sequence relative to the bit depth of the base picture layer.
[0051] Optionally, for the third aspect, the enhancement picture layer comprises data used in the decoding method to reconstruct the distance of objects in the image from the viewer.
[0052] Optionally, for the third aspect, the enhanced image layer includes data used in the decoding method to reconstruct tactile feedback of the user.
[0053] Optionally, for the third aspect, the method further includes: sending an indication that the enhancement data should be discarded.
[0054] Optionally, for the third aspect, the layered decoding method used is MPEG-5 LCEVC (Low Complexity Enhancement Video Coding) encoding or SMPTE VC-6 encoding.
[0055] Optionally, for the third aspect, at least one of the enhancement data is received as user data embedded within LCEVC data coefficients. As a non-limiting example, one or more residual coefficients of an image include embedded depth information representing a depth of a corresponding object relative to a viewpoint, and the decoder processes the embedded data to reconstruct a depth map associated with the image frame based at least in part on the embedded data. In a non-limiting embodiment, the depth map is reconstructed by processing the embedded data and the image data.
[0056] Optionally, for the third aspect, or as a separate implementation, the display device uses frames plus depth data sent at a given frame rate to increase the frame rate through spatial warping (depth-based reprojection) to match the display frame rate, thereby enabling rendering and video stream processing to be performed at a frame rate lower than the display frame rate and reducing the required bandwidth and processing resources.
[0057] Optionally, for the third aspect, each frame comprises image data and depth map data, and the method comprises performing layered decoding on frames in the frame sequence to generate an encoded frame comprising one or more of a base depth layer and an enhancement depth layer.
[0058] Optionally, for the third aspect, each frame includes an image layer and a depth layer, and the coded frame includes a base image layer, a base depth layer, an enhanced image layer, and an enhanced depth layer.
[0059] Optionally, for the third aspect, each frame comprises depth map data embedded in the image data, the base depth layer is a base image layer having the embedded depth map data, and the enhancement depth layer is an enhancement image layer having the embedded depth map data.
[0060] Optionally, for the third aspect, the method also includes: sending a depth map discard indication indicating that depth map data should be discarded when transmitting a frame sequence; and if the depth map data is discarded, performing layered decoding on the image data of the frame to generate a decoded frame from a base image layer and an enhanced image layer.
[0061] In aspects related to the third aspect, the following disclosure provides a decoder or a display device configured to perform the method according to the third aspect. Optional features of the third aspect may be applicable to related aspects.
[0062] According to a fourth aspect, the present disclosure provides a method for encoding a frame sequence representing a dynamic 3D scene, wherein each frame includes image data and other data, and the method includes: performing layered encoding on frames in the frame sequence to generate a coded frame, and the coded frame includes a base layer and an enhancement layer. Specifically, according to an embodiment of the fourth aspect, the present disclosure provides a method for encoding a frame sequence representing a dynamic 3D scene, wherein each frame includes image data and depth map data, and the method includes: performing layered encoding on frames in the frame sequence to generate a coded frame, and the coded frame includes a base depth layer and an enhancement depth layer.
[0063] In aspects related to the fourth aspect, the following disclosure provides an encoder or renderer configured to perform the method according to the fourth aspect. Optional features of the fourth aspect may be applicable to related aspects.
[0064] According to a fifth aspect, the present disclosure provides a method for decoding a sequence of frames representing a dynamic 3D scene, wherein each frame includes image data and depth map data, the method including identifying one or more object regions in the image from the image data, and assigning depth information to each of the one or more object regions in the image according to the depth map data. In some embodiments, the depth map data includes a base depth layer and an enhanced depth layer, and decoding each frame includes performing layered decoding using the image data, the base depth layer, and the enhanced depth layer.
[0065] In aspects related to the fifth aspect, the following disclosure provides a decoder or a display device configured to perform the method according to the fifth aspect. Optional features of the fifth aspect may be applied to related aspects.
[0066] According to a sixth aspect, the following disclosure provides a method comprising: receiving frames plus depth data at a given frame rate, the frame plus depth data being encoded by a layered coding method; and at a display device, using the depth data to increase the frame rate by depth-based reprojection to match the display frame rate.
[0067] Optionally, for the sixth aspect, the frame includes data representing a dynamic 3D scene, and each frame includes base layer image data and enhancement data. Optionally, for the sixth aspect, the method includes: performing layered decoding on frames in a frame sequence to generate a decoded frame from an encoded frame including a base image layer and an enhancement image layer.
[0068] In aspects related to the sixth aspect, the following disclosure provides a decoder or a display device configured to perform the method according to the sixth aspect. Optional features of the sixth aspect may be applied to related aspects.
[0069] According to a seventh aspect, the present disclosure provides a bit sequence representing the encoding of a frame sequence representing a dynamic 3D scene, the bit sequence comprising: encoding data of a base depth layer of the frame; and encoding data of an enhanced depth layer for the frame.
[0070] Optionally, for the seventh aspect, the bit sequence further includes: encoded data of a base image layer of the frame; and encoded data of an enhanced image layer of the frame.
[0071] Optionally, for the seventh aspect, the base depth layer is a base image layer having embedded depth map data, and the enhancement depth layer is an enhancement image layer having embedded depth map data.
[0072] According to an eighth aspect, the present disclosure provides a method for transmitting a frame sequence representing a dynamic 3D scene, wherein each frame includes image data and depth map data, the method comprising: obtaining a coded frame including a base depth layer and an enhanced depth layer; determining whether to discard the depth map data when transmitting the frame sequence; if the depth map data is to be discarded: discarding the depth map data from the coded frame to obtain a reduced coded frame including a base image layer and an enhanced image layer; and transmitting the reduced coded frame; and otherwise, transmitting the coded frame.
[0073] Optionally, for the eighth aspect: each frame includes an image layer and a depth layer, the encoded frame includes a base image layer, a base depth layer, an enhanced image layer and an enhanced depth layer, discarding depth map data includes discarding the enhanced depth layer, and preferably also includes discarding the base depth layer.
[0074] Optionally, for the eighth aspect: each frame includes depth map data embedded in the image data, the base depth layer is a base image layer having embedded depth map data, the enhanced depth layer is enhanced image layer data having embedded depth map data, and discarding the depth map data includes removing the embedded depth map data from the enhanced image layer, and preferably further includes removing the embedded depth map data from the base image layer.
[0075] Optionally, for the eighth aspect, the method further includes: generating a depth map discard indication, which indicates whether to discard depth map data when transmitting a frame sequence; and sending the depth map discard indication to an upstream encoder, or sending the depth map discard indication together with the encoded frame or the reduced encoded frame.
[0076] In aspects related to the eighth aspect, the following disclosure provides a transmitter configured to perform the method according to the eighth aspect. Optional features of the eighth aspect may be applicable to related aspects.
[0077] According to the ninth aspect, the following disclosure provides a method for encoding a frame sequence performed by an encoder, the method comprising: performing layered encoding on a first frame in the frame sequence to generate a first encoded frame including a base layer and an enhancement layer; storing components of the enhancement layer of the first encoded frame in a time buffer for time encoding of subsequent frames; sending the first encoded frame to a transmitter for transmission; receiving an enhancement discard indication, which indicates whether the enhancement layer is discarded when transmitting the first encoded frame; performing layered encoding on a second frame in the frame sequence to generate a second encoded frame including a base layer and an enhancement layer, wherein: if the enhancement discard indication indicates that the enhancement layer is not discarded, the enhancement layer of the second encoded frame is generated with reference to the time buffer, and if the enhancement discard indication indicates that the enhancement layer is discarded, the enhancement layer of the second encoded frame is generated without reference to the time buffer.
[0078] Optionally, for the ninth aspect, the method further includes clearing the time buffer if the enhancement discard indication indicates that the enhancement layer is discarded.
[0079] Optionally, for the ninth aspect, the method further includes: if no enhancement discard indication is received within a predetermined time limit, layered encoding a second frame in the frame sequence to generate a second coded frame including a base layer and an enhancement layer, wherein the enhancement layer of the second coded frame is generated by a reference time buffer.
[0080] Optionally, for the ninth aspect, the method further includes: if no enhancement discard indication is received within a predetermined time limit, layered encoding a second frame in the frame sequence to generate a second coded frame including a base layer and an enhancement layer, wherein the enhancement layer of the second coded frame is generated without referring to the time buffer.
[0081] Optionally, for the ninth aspect, generating an enhancement layer for a frame without referencing the temporal buffer includes: decoding a base layer of the encoded frame; and calculating a residual as a difference between the decoded base layer and the frame.
[0082] Optionally, for the ninth aspect, generating an enhancement layer of a frame with reference to the time buffer includes: decoding a base layer of the encoded frame; calculating a residual as the difference between the decoded base layer and the frame; and calculating the difference between the residual of the current frame and the corresponding residual of the previous frame stored in the time buffer.
[0083] Optionally, for the ninth aspect, the layered coding is LCEVC coding.
[0084] In aspects related to the ninth aspect, the following disclosure provides an encoder configured to perform the method according to the ninth aspect. Optional features of the ninth aspect may be applicable to related aspects.
[0085] Although various aspects (particularly, the first aspect) have been described with respect to one or more "frames", it should be understood that data set in a "frame" is optional, and therefore it should be noted that data can be set in a frame in other formats and still be considered part of the present disclosure.
[0086] Different variations of different aspects may be combined in any combination. Certain aspects may be modified by omitting and / or features. Certain aspects not described above may be provided by combining any single feature described above and / or below in connection with any aspect, variation or embodiment.
[0087] Any of the above methods can be implemented by one or more processors executing a computer program. The computer program includes computer readable instructions, and when the computer readable instructions are executed by the processor, the processor performs the corresponding method. The computer readable instructions can be stored in a non-transitory computer readable medium. The computer readable instructions can be encoded in a digital signal such as an optical signal or an electrical signal.
[0088] Additionally, for any of the above methods of generating or using a sequence of coded frames, the sequence of coded frames may be separated out as a bitstream that may be transmitted to a final destination for use by a user shortly after rendering or shortly after encoding, or may be stored indefinitely in a memory device. For example, the bitstream may be stored by a streaming service as part of a video-on-demand service. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 schematically illustrates the display of a frame sequence representing a dynamic 3D scene without using a depth map;
[0090] Figure 2 The display of a frame sequence representing a dynamic 3D scene using a depth map is schematically shown;
[0091] Figure 3 A networked system for generating a sequence of frames for rendering a dynamic 3D scene is schematically illustrated;
[0092] Figure 4A and 4B Schematically showing the features of LCEVC encoding and decoding related to the present invention;
[0093] Figure 5 is an illustrative example of layered decoding;
[0094] Figure 6 is a schematic example of a residual plot;
[0095] Figure 7 is a schematic diagram of parallel processing in LCEVC compared to a fully serial coding scheme;
[0096] Figure 8 Figure 1 is a schematic diagram of the key features of VC-6 decoding.
[0097] Fig. 9 is a diagram of an example configuration of a content rendering network. DETAILED DESCRIPTION
[0098] The non-limiting example embodiments disclosed herein provide methods for structuring rendering resources into a hierarchy of potentially available resources (defined herein as a "content rendering network" or "CRN", similar to a content delivery network that provides a hierarchical caching infrastructure for web content). Based on real-time metrics, including, as non-limiting examples, the expected processing load of the rendering to be performed, the status of available resources in the CRN, the network latency of resources in the CRN, and the number of other users requesting rendering of the same 3D space from a range of viewpoint angles, the CRN can dynamically select one or more resources to perform the rendering process, and - in the best case - can efficiently pool expensive rendering calculations so that the calculations are performed once for multiple users, and then the compressed byproducts are distributed to another resource (another node of the CRN or a user device), which performs the rendering of the final viewport and streams it to a display device.
[0099] Other non-limiting example embodiments disclosed herein provide methods for optimizing the transmission of ultra-low latency video frames through a layered coding method, i.e., including a data layer representing a signal at a lower quality level and one or more additional data layers, the one or more additional data layers including information for reconstructing a reproduction of the signal at a higher quality level. As non-limiting examples, the additional data layers allow a decoder device to enhance one or more quality aspects of the signal, including visual fidelity, resolution, image data bit depth (e.g., for high nits HDR), frame rate, stereoscopic vision, focal length (e.g., for zoom adjustment), haptics, etc. One or more additional data layers can be considered optional and can be safely discarded during encoding or even during transmission without affecting the quality of the base layer data and without the need to transmit expensive base layer I frames (because only the IDR of the enhanced data is required, which is typically much smaller). Therefore, layered coding allows enhanced video to be transmitted within the same bandwidth, but more importantly, it provides additional degrees of freedom in how to reduce quality when bandwidth is scarce and to dynamically adapt to bandwidth drops more quickly.
[0100] Figure 1 and Figure 2 Schematic showing the use of depth maps and warping when generating and displaying dynamic 3D scenes.
[0101] exist Figure 1 and Figure 2 In FIG. 1 , a plurality of image parts 1a, 1b, 1c are shown. These can be displayed simultaneously as part of an overall image.
[0102] Alternatively, they may be displayed in sequence as separate frames. In other words, image portion 1a may correspond to frame N, image portion 1b may correspond to frame N+1 and so on.
[0103] exist Figure 1 In the 3D virtual environment, all image parts are mapped to the same plane 3 when displayed. In other words, in a 3D virtual environment seen by a user through a 3D display such as a VR, XR or AR display, all image parts 1a, 1b, 1c are displayed as if they were at the same distance from the viewer.
[0104] On the other hand, Figure 2 In the embodiment of the present invention, each image part 1a, 1b, 1c is mapped to a corresponding virtual plane 3a, 3b, 3c in the 3D virtual environment, which may be located at different distances from the viewer.
[0105] In a 3D environment, planes at a fixed distance from the viewer are curved, and planes at different distances have different curvatures. Therefore, various visual effects such as time warping and focus adjustment (which can be local or zoom, and can be used in response to eye tracking) behave differently depending on the distance between the viewer and the image portion. Therefore, by incorporating information about the distance between the image portion and the viewer, the 3D environment can be rendered more realistically.
[0106] Such information about the distance from the viewer can be incorporated into the frame using a depth map. In the case of a planar image portion, a depth can be assigned to each pixel or block of the image portion. Alternatively, when the frame includes point cloud information or other information about the location of points or objects in the 3D virtual environment, this information can be used as the equivalent of a depth map when performing visual effect calculations.
[0107] Figure 1 and Figure 2 It also exhibits time-warping characteristics. Figure 1 and Figure 2 As shown, each of the displayed images 1a, 1b, 1c is a portion of the region of the warpable images 2a, 2b, 2c. Due to the delay between rendering and displaying frames, the warpable images 2a, 2b, 2c are rendered before the latest viewing direction of the user is known. The warpable images 2a, 2b, 2c can be transmitted to a display device, or to a rendering node close to the display device, which can perform time warping based on the corresponding warpable images 2a, 2b, 2c and the latest viewing direction of the user to generate display image portions 1a, 1b, Figure 1 c.
[0108] Figure 3 A system for generating a sequence of image frames is schematically shown.
[0109] Reference Figure 3 The system includes an image generator 31, an encoder 32, a transmitter 33, a network 34, a receiver 35, a decoder 36 and a display device 37.
[0110] The image generator 31 , the encoder 32 and the transmitter 33 together form a first rendering node.
[0111] The image generator 31 may, for example, include a rendering engine for initially rendering a virtual environment such as a game or a virtual conference room.
[0112] The image generator 31 is configured to generate a sequence of frames to be displayed. The frame may include one or more image parts 1a, 1b, 1c as described above. Additionally or alternatively, the frame may include point cloud data. In a point cloud, each point typically has a 3D position and one or more attributes, where the attributes may include, for example, surface color, transparency value, object size, and surface normal direction. Each attribute may have a value selected from a continuous range or may have a value selected from a discrete set.
[0113] Each frame may further include depth map data as described above. The depth map data may be provided as a depth layer layer separate from the image layer. In some contexts, such as MPEG immersive video (MIV), the image layer may alternatively be described as a texture layer. Similarly, in some contexts, the depth layer layer may alternatively be described as a geometry layer.
[0114] In addition, each frame may include a predicted display window position. The predicted display window position is the position of the portion of the generated image 2a, 2b, 2c that may be displayed by the display device 37. The predicted display window position may be based on the viewing position of the user obtained from the display device 37 (such as the user's virtual position and / or orientation in the 3D environment). The predicted display window position may be defined using one or more coordinates. For example, with reference to Figure 1 , the predicted display window position can be defined using the coordinates of the corners or center of the predicted display window, and the predicted display window position can be defined using the size of the predicted display window. The predicted display window position can be encoded as part of the metadata included in the frame.
[0115] Each frame may include further information, which may be provided as a separate layer. For example, each frame may further include audio information or haptic feedback information indicating audio or haptic that may accompany the displayed visual data. An audio layer or haptic layer may accompany each frame, and may be omitted for frames where accompanying audio or haptic is not required.
[0116] The frame can be based on the state of the virtual environment, the position of the user, or the viewing direction of the user. Here, the position and viewing direction can be physical attributes of the user in the real world, or the position and viewing direction can also be purely virtual, such as controlled using a handheld controller. The image generator 31 can obtain information indicating the position, viewing direction, or motion of the user, for example, from the display device 37. In other cases, the generated image can be independent of the user's position and viewing direction. The generation of such images usually requires a large amount of computer resources such as a powerful GPU, and can be implemented in a cloud service or on a local but powerful computer. For example, a cloud service (such as a cloud rendering service (CRN)) can reduce the cost per user, making image frame generation more accessible to a wider range of users. "Rendering" here refers at least to the initial stage of rendering to generate an image. Further rendering can be performed at the display device 37 based on the generated image to produce a final image for display.
[0117] The encoder 32 is configured to encode frames to be transmitted to the display device 37. The encoder 32 may be implemented using executable software or may be implemented on specific hardware such as an ASIC.
[0118] The encoder may apply inter-frame or intra-frame compression based on the current encoded frame and optionally one or more previously encoded frames.
[0119] The encoder 32 may be a multi-layer encoder, such as an LCEVC encoder (eg Figure 4A as shown) or VC-6 encoder.
[0120] For example, when a generated frame includes depth map data, the encoder can perform layered encoding on each frame to generate an encoded frame including a base depth layer and an enhanced depth layer. Similar to encoding image data, encoding depth maps in this manner can improve compression. In certain applications, such as HDR video, depth maps need to be very detailed, with bit depths as high as 12 or 14 bits, which significantly increases the data to be transmitted. Therefore, providing methods for improving depth map compression can make more realistic displays based on depth maps feasible when rendering is performed in real time or when transmitting rendered data. In addition, this type of layered encoding can easily discard (and then pick up again) one or more of the layers, providing flexibility and tools for bandwidth management.
[0121] Layered encoding also helps because the end decoder / user device (such as a user display device) can choose whether to process these additional layers. For example, in a non-layered approach, the best the end device (i.e., the receiver, decoder, or display device associated with the user who will view the frame) can do is determine that it does not have enough resources to meet a given quality (whether it is resolution, frame rate, inclusion of depth maps), and then signal to the controller / renderer / encoder that it does not have enough resources. The controller will then send future frames at a lower quality. In this alternative scenario, unfortunately, the end device still has to process the higher quality data until the lower quality data arrives (if it can process the received frame at all).
[0122] In some of the described embodiments, this situation is improved because when / if the end device determines, for example, that it does not have the processing capabilities to handle the highest quality level, it can discard and / or choose not to process certain layers. The end device can also signal to the controller that it requires a lower quality level, but in the meantime, the end device can only process the number of layers it can handle. Thus, the end device can react to conditions more quickly.
[0123] In some cases, the depth map data may be embedded in the image data. In this case, the base depth layer may be a base image layer with embedded depth map data, and the enhancement depth layer may be an enhancement image layer with embedded depth map data.
[0124] Alternatively, when the generated frames include a depth layer separate from the image layer and multi-layer coding is applied, the coded depth layer can be separated from the coded image layer. The advantage of this is that the coded depth layer can be discarded under certain conditions while still retaining the image layer that can be displayed (albeit with a lower level of realism). For example, when available communication resources are reduced, the coded depth layer can be discarded by the transmitter or encoder, or it can be discarded by a terminal device that lacks the processing resources to handle the highest quality level.
[0125] Similarly, if certain frames include an audio base layer, a haptic feedback base layer, an audio enhancement layer, or a haptic feedback enhancement layer, these layers can be flexibly processed or discarded.
[0126] Additionally or alternatively, in the case where the frame includes point cloud data, the encoder can apply point cloud data encoding techniques such as described in European patent application EP21386059.6, which is incorporated herein by reference. Such a point cloud encoder can serve as a base encoder for layered coding techniques such as LCEVC or VC-6. It is worth noting that LCEVC and VC-6 techniques encode and decode layered signals, but are agnostic to the content type of the data encoded in the signal. For example, a signal can include textures, video frames, geometry or depth data, meshes, point clouds, rendering attributes, or physical engine attributes.
[0127] The transmitter 33 may be any known type of transmitter for wired or wireless communication, including an Ethernet transmitter or a Bluetooth transmitter.
[0128] The transmitter 33 may be configured to make decisions about how to transmit the frame, and / or may provide feedback to the encoder 32 or the image generator 31. For example, the transmitter 33 may determine the available communication resources (e.g., bandwidth) for sending the frame, and may drop one or more layers from the encoded frame, or indicate to the image generator 31 and / or the encoder 32 that the frame should be generated and encoded with fewer layers when insufficient bandwidth is available to transmit all the generated data. As a specific example, the transmitter 33 may be configured to drop a depth layer, an LCEVC enhancement layer, or a VC-6 enhancement layer from a frame when insufficient communication resources are available.
[0129] The network 34 is used for communication between the transmitter 33 and the receiver 35, and can be any known type of network, such as a WAN or LAN or a wireless Wi-Fi or Bluetooth network. The network 34 can also be a combination of multiple different types of networks. Many users only have access to networks with a bandwidth of 30MBps, which can cause latency jitter in the streaming media. The required bandwidth and observed latency can be reduced by strategies such as forward-looking rendering and last-millisecond reprojection, which are achieved by improving compression.
[0130] Receiver 35 may be any known type of receiver for wired or wireless communication, including an Ethernet transmitter or a Bluetooth transmitter.
[0131] The decoder 36 is configured to receive and decode the encoded frames. The decoder 36 may be implemented using executable software or may be implemented on specific hardware such as an ASIC.
[0132] The display device 37 may be, for example, a television screen or a VR headset. The timing of the display may be associated with a configured frame rate so that the display device 37 may wait before displaying an image. The display device 37 may be configured to perform the warping, i.e., to adjust the warpable images 2a, 2b, 2c, the final images 1a, 1b, 1c corresponding to the final viewing direction of the user, and display the final image in order to obtain the final display window position.
[0133] As described above, the first rendering node may include an image generator 31, an encoder 32, and a transmitter 33. Additional similar rendering nodes may be included in the system and may work together to generate a sequence of frames.
[0134] In one case, multiple rendering nodes may each provide a portion of a frame sequence to a frame assembly node.
[0135] For example, the receiver 35, decoder 36 or display device 37 may be configured to combine portions of a sequence of frames from multiple sources to generate a final sequence of frames for display.
[0136] Alternatively, the frame assembly node may be separate from the receiver 35 , decoder 36 and display device 37 .
[0137] Additionally or alternatively, multiple rendering nodes may be chained. In other words, successive rendering nodes may add components to a partial sequence of frames as it is passed from rendering node to rendering node, and ultimately provide the complete sequence of frames to the receiver 35. Furthermore, each rendering node may obtain rendered components from multiple upstream rendering nodes and / or distribute rendered components to multiple downstream rendering nodes.
[0138] An example of this situation is Figure 3 The second rendering node includes a transceiver 33b, a decoder 36b, an image generator 31b and an encoder 32b. The second rendering node is configured to obtain a first partial rendering frame from the first rendering node, and generate a second partial or complete rendering frame based on the first partial rendering frame.
[0139] The image generator 31b may be similar to the image generator 31. The image generator 31b may also be referred to as an image processor because it processes a first partially rendered frame to generate a second partially or fully rendered frame. The frame may include one or more image portions 1a, 1b, 1c as described above. Additionally or alternatively, the frame may include point cloud data.
[0140] The data type of the second partially or fully rendered frame may be different from the data type of the first partially rendered frame. For example, the first partially rendered frame may include point cloud data, while the second partially or fully rendered frame may include image data generated from the point cloud data.
[0141] Alternatively, a first partially rendered frame may include first image data and a second partially or fully rendered frame may include second image data generated from the first image data.
[0142] Alternatively, a first partially rendered frame may include first point cloud data and a second partially or fully rendered frame may include second point cloud data generated from the first point cloud data.
[0143] The frames generated by each render node may include multiple data types, such as a mix of image data and point cloud data.
[0144] The second rendering node may operate on the first partially rendered frame in an encoded form (e.g., simply merging the first bitstream generated by the first rendering node with the second bitstream generated by the second rendering node. Alternatively, the decoder 36 b may decode the first bitstream so that the image generator 31 b may perform further rendering by referring to the first partially rendered frame. If the further rendering performed by the second rendering node involves only a portion of the 2D area or 3D volume rendered by the first rendering node, or if the further rendering performed by the second rendering node involves only partial attributes of the frame rendered by the first rendering node (e.g., if color information is not required for the further rendering), the decoder 36 b may decode only a portion of the data encoded in the first bitstream.
[0145] Similarly, the first rendering node may not be at the beginning of the rendering node chain. The first rendering node itself may obtain an earlier generated partial rendering frame from one or more upstream rendering nodes, and generate the first partial rendering frame based on the earlier generated partial rendering frame.
[0146] Chains of render nodes may be useful for performing different rendering tasks that require different amounts of processing resources or different frame rates. For example, a company may provide distributed processing in the form of centralized hubs where processing resources are plentiful but farther from the user, and edge locations where processing resources are more scarce but closer to the user. Expensive but fairly static rendering features, such as background lighting or the effects of the environment on sound, can be generated at the central hub (e.g., using ray tracing), while features that require fewer resources but are more responsive or have higher frame rates can be generated closer to the user. In other words, the more responsive a rendering feature is, the shorter the latency required between the rendering node that generates the feature and the user's display, and in a rendering node chain, the node that generates each rendering feature can be selected based on the maximum latency required for that feature. On the other hand, if a rendering feature is expensive to generate, it may be better to generate that feature less frequently and with a higher maximum latency. For example, static, high-quality background features can be generated early in a rendering node chain, while dynamic, but potentially lower-quality foreground features can be generated later in the rendering node chain, closer to the user's device. Here, the effect of the environment on the sound means, for example, that a set of surfaces can be constructed, each of which has different sound reflection and absorption characteristics depending on the material and shape. The frame rate can be matched by creating multiple frames with features generated at a lower frame rate and combining them with frames with features generated at a higher frame rate. In a non-limiting embodiment, the preliminary rendering generates volumetric object data including motion vectors at a first (lowest) frame rate, then generates a 2D rendered frame plus depth information for a specific user at a second (higher) frame rate, and then transmits the video plus depth data to the user device, which generates a final frame for display through spatial distortion (depth-based reprojection) at a third (highest) frame rate. One or more of these steps can be performed in conjunction with other described embodiments. When additional rendering tasks are performed at different rendering nodes in the chain, the user's viewing position may change. Each or any rendering node can obtain an updated viewing position before performing its respective rendering task.
[0147] In addition, the system can generate multiple frame sequences for different users or different displays at the same time. For example, in the context of a VR or AR experience, each user or display can view a different 3D environment, or can view different parts of the same 3D environment. When using a chain of rendering nodes, each node can serve multiple users or only one user.
[0148] For example, a starting rendering node (e.g., at a centralized hub) can serve a large group of users. For example, a group of users may be viewing a nearby portion of the same 3D environment. In this case, the starting node can present a wide field of view ("field of view") that is relevant to all users in the large group.
[0149] The starting node may send the wide field of view to a first intermediate rendering node, which renders additional aspects of the 3D environment. These additional aspects may be, for example, aspects that require less processing power to render, or may be aspects specific to individual users of a group. Additionally, the intermediate rendering node may render features in a smaller field of view than the starting node - this smaller field of view may be relevant to each user rather than to a group of users. Additionally, the first intermediate rendering node may serve only a smaller number of users (e.g., half of a large group of users), while the remaining users are served by a second intermediate rendering node, which also receives the wide field of view from the starting node.
[0150] Then, the (one or more) intermediate rendering nodes may send a second part or a fully rendered frame sequence to the terminal device for each user. The terminal device may perform further processing, such as warping or focal length adjustment, optionally using depth map data.
[0151] Preferably, each rendering node encodes the partially or fully rendered frame before transmitting it to the next rendering node or receiver 35. This means that when the rendering nodes are separated by one or more networks, or more generally when implemented in a distributed system such as the cloud, the communication resources required can be reduced.
[0152] However, each rendering node in the chain encodes different partially or fully rendered frames using different data. Therefore, it may be advantageous for different rendering nodes to use different rendering formats and / or encoding formats. For example, the output from the first rendering node may be point cloud data that logically describes the 3D scene. This point cloud data may be encoded using the techniques of EP21386059.6. The second rendering node may then operate on the point cloud data to generate image data that is more easily displayable by a general display device without the display device having to model the 3D environment. This image data may be encoded using video encoding techniques.
[0153] The link of rendering nodes may be extended to an arbitrary tree structure, where a rendering node obtains partially rendered frames from more than one previous rendering node and generates a further partially or fully rendered frame based on a sequence of the obtained multiple partially rendered frames.
[0154] For example, a content rendering network (CRN) comprising multiple rendering nodes can be used to provide volumetric events to a large number of simultaneous users, such as users participating in a shared virtual environment. Rendering the same event for each user is much more expensive in terms of computational time and power consumption than rendering the volumetric effect once and performing the equivalent rendering of the multicast volumetric effect for multiple users. For example, each user can have a second rendering node (such as a VR headset), and the network can include a central first rendering node. The first rendering node can render the volumetric event and distribute the partial rendered frames depicting the volumetric event to different second rendering nodes. The second rendering node of each user can then integrate the partial rendered frames depicting the volumetric event into the view of the virtual environment currently being displayed to each user based on parameters such as the user's virtual position.
[0155] The receiver 35, the decoder 36, and the display device 37 may be combined into a single device, or may be separated into two or more devices. For example, some VR headset systems include a base unit and a headset unit that communicate with each other. The receiver 35 and the decoder 36 may be incorporated into such a base unit.
[0156] In some embodiments, network 34 may be omitted. For example, a home display system may include a base unit configured as an image source and a portable display unit including display device 37.
[0157] If the decoder 36 or display device 37 does not process or cannot process one or more layers, the receiver 35 or another transmitter associated with the decoder 36 or display device 37 may send back corresponding layer discard indications via the network 34. Each rendering node may receive a layer discard indication. A rendering node that generates a partially or fully rendered frame for that particular decoder 36 or display device 37 may stop generating the discarded layers. On the other hand, a rendering node that generates partially or fully rendered frames for multiple terminal devices may ignore a layer discard indication received from one terminal device (because the other devices still need the discarded layers). Alternatively, a rendering node that serves multiple terminal devices may record the received layer discard indications and may stop generating the discarded layers only when all terminal devices served by the rendering node indicate that the layer is to be discarded.
[0158] In a preferred embodiment, the encoder or decoder is part of a layered coding scheme or format based on layers. Layered coding enables frames to be communicated at higher resolutions and / or higher frame rates than single-layer coding schemes. In layered coding, one or more enhancement layers are communicated with base data, wherein the enhancement layers can be used to upsample the base data at the decoder, for example, to provide upsampling in a spatial or temporal dimension. When combined with equivalent downsampling of the original frame and generation of enhancement layers at the encoder, layered coding can generally provide lossless compression of the data, with higher resolution and / or higher frame rates for a given transmission bit rate. Examples of layer-based hierarchical coding schemes include LCEVC: MPEG-5 Part 2 LCEVC ("Low Complexity Enhancement Video Coding") and VC-6: SMPTE VC-6ST-2117, the former described in PCT / GB2020 / 050695, published as WO 2020 / 188273 (and associated standard documents), and the latter described in PCT / GB2018 / 053552, published as WO 2019 / 111010 (and associated standard documents), all of which are incorporated herein by reference. However, the concepts shown herein are not necessarily limited to these specific layered coding schemes.
[0159] Another example is described in WO2018 / 046940, which is incorporated herein by reference. In this example, a set of residuals is encoded relative to the residuals stored in the time buffer.
[0160] LCEVC (Low Complexity Enhancement Video Coding) is a standardized coding method proposed in standard specification documents, including ISO / IEC 23094-2 Version 1 Low Complexity Enhancement Video Coding, released in November 2021, which is incorporated herein by reference.
[0161] Figure 4A and Figure 4B Selected features of an LCEVC encoder 402 and an LCEVC decoder 404 are schematically shown, which show how LCEVC can be used to efficiently encode and decode static image regions 3. Further implementation details of these types of encoders and decoders are set out in earlier published patent applications GB1615265.4 and WO2020188273, each of which is incorporated herein by reference.
[0162] In each of the encoder 402 and the decoder 404, items are shown on two logical layers. The two layers are separated by a dotted line. The items in the first, highest layer relate to data at a higher quality level. The items in the second, lowest layer relate to data at a lower quality level. The higher and lower quality levels relate to a hierarchical structure having multiple quality layers. In some examples, the hierarchical structure includes more than two quality levels. In such examples, the encoder 402 and the decoder 404 may include more than two different layers. Figure 4A and 4B There may be one or more other layers above and / or below the layers depicted in FIG.
[0163] Reference Figure 4A , the encoder 402 obtains input data 406 at a higher quality level. The input data 406 includes a first time sample t of the signal at the higher quality level. 1 The input data 406 may, for example, include an image generated by the image generator 31. The encoder 402 uses the input data 406 to derive downsampled data 412 at a lower quality level, for example by performing a downsampling operation on the input data 406. Where the downsampled data 412 is processed at a lower quality level, such processing generates processed data 413 at a lower quality level.
[0164] In some examples, generating processed data 413 involves encoding downsampled data 412. Encoding downsampled data 412 produces a lower quality level encoded signal. Encoder 402 can output, for example, an encoded signal for transmission to decoder 404. The encoded signal can be generated by an encoding device separate from encoder 402, rather than generated in encoder 402. The encoded signal can be an H.264 encoded signal. H.264 encoding can involve arranging a sequence of images into groups of pictures (GOPs). Each image in a GOP represents a different time sample of a signal. In a process known as "inter-frame prediction", a given image in a GOP can be encoded using one or more reference images associated with earlier and / or later time samples from the same GOP.
[0165] Generating processed data 413 at a lower quality level may further involve decoding the encoded signal at the lower quality level. The decoding operation may be performed to simulate the decoding operation at the decoder 404, as will become apparent below. Decoding the encoded signal produces a decoded signal at a lower quality level. In some examples, the encoder 402 decodes the encoded signal at a lower quality level to produce a decoded signal at a lower quality level. In other examples, the encoder 402 receives a decoded signal at a lower quality level, for example, from an encoding and / or decoding device separate from the encoder 402. The encoded signal may be decoded using an H.264 decoder. H.264 decoding produces a sequence of images (i.e., a sequence of time samples of a signal) at a lower quality level. After the H.264 decoding process is completed, no individual image indicates the temporal correlation between different images in the sequence. Therefore, any use of the temporal correlation between sequential images employed by H.264 encoding is eliminated during H.264 decoding because the sequential images are decoupled from each other. Therefore, subsequent processing is performed on an image-by-image basis, where the encoder 402 processes the video signal data.
[0166] In an example, generating processed data 413 at a lower quality level further involves obtaining correction data based on a comparison between the downsampled data 412 and a decoded signal obtained by the encoder 402, for example, based on a difference between the downsampled data 412 and the decoded signal. The correction data may be used to correct errors introduced when encoding and decoding the downsampled data 412. In some examples, the encoder 402 outputs the correction data, for example, for transmission to the decoder 404, along with the encoded signal. This allows a recipient to correct errors introduced when encoding and decoding the downsampled data 412.
[0167] In some examples, generating processed data 413 at the lower quality level further involves correcting the decoded signal using the correction data. In other examples, the encoder 402 uses the downsampled data 412 instead of correcting the decoded signal using the correction data.
[0168] In some examples, generating processed data 413 involves performing one or more operations in addition to the encoding, decoding, obtaining, and correcting actions described above.
[0169] However, in some examples, no processing is performed on the downsampled data 412 .
[0170] The data at the lower quality level is used to derive upsampled data 414 at a higher quality level, for example, by performing an upsampling operation on the data at the lower quality level. The upsampled data 414 includes a second reproduction of the first time sample of the signal at the higher quality level. The encoder 402 obtains a set of residual elements 416 that can be used to reconstruct the input data 406 using the upsampled data 414. The set of residual elements 416 is associated with the first time sample t of the signal. 1 The set of residual elements 416 is obtained by comparing the input data 406 with the upsampled data 414 .
[0171] In this example, the encoder 402 generates a set of temporal correlation elements 426. The term "temporal correlation element" is used herein to refer to a correlation element that indicates a degree of temporal correlation. A temporal correlation element may also be a spatiotemporal correlation element that indicates a degree of spatial correlation between residual elements. In this example, the set of temporal correlation elements 426 is associated with a first time sample t 1 The second time sample t of the sum signal 0 In the example described herein, the second time sample t 0 is an earlier time sample relative to the first time sample. However, in other examples, the second time sample t 0 is relative to the first time sample t 1 In some examples where the input data 406 includes a sequence of time samples, the earlier time sample is the one in the input data that is the first time sample t 1 The previous time sample. At the first time sample t 1 When the earlier time samples are arranged in the order of presentation, the earlier time samples are placed before the first time sample t 1 Before.
[0172] The second time sample t 0 It can be relative to the first time sample t 1 In some examples, the second time sample t 0 is relative to the first time sample t 1 , rather than relative to the first time sample t 1 The previous time sample.
[0173] In this example, the set of temporal correlation elements 426 indicates a degree of spatial correlation between a plurality of residual elements in the set of residual elements 416. The set of temporal correlation elements 426 also indicates a first reference data based on the input data 406 and a second temporal sample t based on a signal at, for example, a higher quality level. 0 Therefore, the first reference data is related to the first time sample t of the signal.1 and the second reference data is associated with the second time sample t of the signal 0 The first reference data and the second reference data are used to determine the first time sample t 1 The second time sample t of the sum signal 2 A reference or comparison of the relevant degree of temporal correlation. The first reference data and / or the second reference data may be at a higher quality level.
[0174] In some examples, the first reference data and the second reference data include a first set of spatial correlation elements and a second set of spatial correlation elements, respectively, the first set of spatial correlation elements being related to the first time sample t 1 and the second set of spatial correlation elements is associated with the second time sample t of the signal 0 Associated.
[0175] In other examples, the first reference data and the second reference data include a first rendition and a second rendition of the signal, respectively, the first rendition being associated with a first time sample t of the signal. 1 and the second reproduction is associated with the second time sample t of the signal 0 Associated.
[0176] The set of time correlation elements 426 will be referred to hereinafter as “Δ t Correlation element” because the temporal correlation is exploited using data from different time samples to generate Δ t Correlation element 426.
[0177] In this example, encoder 402 transmits the group Δ t Correlation element 426. Since the group Δ t The correlation element 426 exploits temporal redundancy at higher residual layers, so in the presence of strong temporal correlation, the set Δ t The correlation elements 426 may be smaller and may in some cases include more correlation elements with zero values. Thus, when applied to a static image region 3 (which is static and only changes internally), the set of Δ t Correlation element 426.
[0178] Now turn to Figure 4B , the decoder 404 receives the data 420 based on the downsampled data 412 and receives the group Δ t Correlation element 426.
[0179] Where the encoder 402 has processed the downsampled data 412 to generate the processed data 413, the decoder 404 processes the received data 420 to generate the processed data 422. The processing may include decoding the encoded signal to produce a decoded signal at a lower quality level. In some examples, the decoder 404 does not perform such processing on the received data 420. The data at the lower quality level, such as the received data 420 or the processed data 422, is used to derive the upsampled data 414. The upsampled data 414 may be derived by performing an upsampling operation on the data at the lower quality level.
[0180] Decoder 404 is based at least in part on the set of Δ t The correlation element 426 is used to obtain the set of residual elements 416. The set of residual elements 416 can be used to reconstruct the input data 406 using the upsampled data 414.
[0181] The present disclosure describes implementations for integrating hybrid backward-compatible coding techniques with existing decoders, optionally via software updates. In a non-limiting example, the present disclosure relates to implementations and integrations of MPEG-5 Part 2 Low Complexity Enhancement Video Coding (LCEVC). LCEVC is a hybrid backward-compatible coding technique that is a flexible, adaptable, efficient, and computationally inexpensive coding format that combines different video coding formats, base codecs (i.e., encoder-decoder pairs such as AVC / H.264, HEVC / H.265, or any other current or future codecs, as well as non-standard algorithms such as VP9, AV1, etc.) with coded data of one or more enhancement layers.
[0182] An example hybrid backward compatible coding technique uses a downsampled source signal encoded using a base codec to form a base stream. An enhanced stream is formed using a set of encoded residuals that correct or enhance the base stream, for example, by increasing the resolution or by increasing the frame rate. There may be multiple levels of enhanced data in the hierarchical structure. In some arrangements, the base stream may be decoded by a hardware decoder, while the enhanced stream may be suitable for processing using a software implementation. Therefore, the stream is considered to be a base stream and one or more enhanced streams, wherein there may typically be two enhanced streams but one enhanced stream is often used. It is worth noting that typically the base stream may be decoded by a hardware decoder, while the enhanced stream may be suitable for software processing implementations with suitable power consumption. The stream may also be considered to be a layer.
[0183] The combined intermediate image is then upsampled again to give a preliminary output image of the highest resolution. The second enhancement sublayer is combined with the preliminary output image to give a combined output image.
[0184] The second enhancement sublayer may be derived in part from a temporal buffer, which is a storage for the second enhancement sublayer of a previous frame. The use of a temporal buffer reduces the amount of data that needs to be included in the coded frame. The temporal buffer may also be used for the first enhancement sublayer.
[0185] An indication of whether the temporal buffer can be used for the current frame or which parts of the temporal buffer can be used for the current frame can be included in the coded frame. Alternatively, the decoder itself can determine that the temporal buffer cannot be used for the current frame, for example in the case where the first or second enhancement sublayer of the previous frame was dropped or incorrectly received and thus the temporal buffer expires.
[0186] Similarly, an indication of whether a temporal buffer should be used for a current frame may be sent to the encoder before encoding the current frame. For example, the decoder may send an enhancement layer (or sublayer) discard indication after the decoder fails to receive the first or second enhancement sublayer of the previous frame. If the encoder receives a discard indication, the encoder may transmit a set of residual elements 416 for the current frame instead of Δ t A collection of related elements 426 .
[0187] Compared to using a block-based approach in the MPEG family of algorithms, video frames are encoded in a hierarchical manner. Encoding frames in a hierarchical manner includes generating a residual for a full frame, and then generating a residual for a reduced or decimated frame, etc. In the examples described herein, the residual can be considered as an error or difference at a specific quality level or resolution.
[0188] For context purposes only, as the detailed structure of the LCEVC is known and is set out in the approved draft standard specification, Figure 4B How LCEVC operates on the decoding side is explained in the form of a logic flow chart, assuming H.264 as the base codec. Those skilled in the art will understand how the examples described herein are also applicable to other multi-layer coding schemes (e.g., those coding schemes using base layers and enhancement layers) based on the general description of LCEVC. The LCEVC decoder operates on a separate video frame layer. It uses the decoded low-resolution pictures and LCEVC enhancement data from the base (H.264 or other) video decoder as input to produce decoded full-resolution pictures ready to be rendered on the display view. LCEVC enhancement data is usually received in the supplementary enhancement information (SEI) of the H.264 network abstraction layer (NAL) or in the additional data packet identifier (PID) and is separated from the base coded video by a demultiplexer. Therefore, the base video decoder receives the demultiplexed encoded base stream and the LCEVC decoder receives the demultiplexed encoded enhancement stream, which is decoded by the LCEVC decoder to generate a set of residuals for combining with the decoded low-resolution pictures from the base video decoder.
[0189] LCEVC can be quickly implemented in existing decoders with software updates and is inherently backward compatible since devices that have not yet been updated to decode LCEVC can play video using the basic base codec, further simplifying deployment.
[0190] In this context, a decoder implementation is proposed herein to integrate decoding and rendering with existing systems and devices that perform basic decoding. The integration is easy to deploy. This also enables support for a wide range of encoding and player vendors, and can be easily updated to support future systems. Embodiments of the present invention specifically relate to how to implement LCEVC in a way that provides decoding of protected content in a secure manner.
[0191] The proposed decoder implementation can be provided through an optimized software library for decoding MPEG-5 LCEVC enhancement streams, thereby providing a simple and powerful control interface or API. This allows developers to flexibly deploy LCEVC at any level of the software stack, such as from low-level command line tools to integration with common open source encoders and players. Specifically, embodiments of the present invention generally relate to driver-level implementations and system-on-chip (SoC)-level implementations.
[0192] The terms LCEVC and enhancement may be used interchangeably herein, for example, an enhancement layer may include one or more enhancement streams, ie, residual data of the LCEVC enhancement data.
[0193] Figure 5 is a schematic diagram showing the process flow of LCEVC. In the first step, the base decoder decodes the base layer to obtain a low-resolution frame (i.e., the base layer). In the next step, the initial enhancement (a sublayer of the enhancement layer) corrects artifacts in the base. As another step, the final frame (for output) is reconstructed at the target resolution by applying another (e.g., other residual details) sublayer of the enhancement layer. This shows that LCEVC improves quality and reduces the overall computational requirements for encoding by making the best use of existing codecs and enhanced features. Embodiments of the present invention provide ways to achieve this in a secure manner (e.g., when processing protected content).
[0194] Figure 6 An enhancement layer is shown. As can be seen, the enhancement layer includes sparse, highly detailed information that would not be interesting (or valuable to the viewer) without the base video. However, subtle motion in the rendered XR view will increase the information stored in this enhancement layer.
[0195] Figure 7 A comparison between latency when using a layered coding scheme (such as LCEVC) and a full serial coding scheme within a VR or XR pipeline is shown.
[0196] like Figure 7 As shown, both methods begin with rendering at a source device 710. This can be any process that generates frame data such as image data or point cloud data, as described above.
[0197] In a fully serial scheme, each frame is encoded at full resolution 720. This can be achieved using any standard codec suitable for the frame data, such as h.264, HEVC, AV1, or VVC for video data (usually a single layer codec).
[0198] Then, the encoded frame is transmitted from the source device to the target device at step 730. For example, the encoded frame may be transmitted between nodes of a content rendering network or between a server and a user device.
[0199] At step 740, the encoded frame is received and stored in a jitter buffer until the receiving entity is ready to process the frame.
[0200] At step 750, the receiving entity decodes the full-resolution frame.
[0201] The full serial scheme ends with reprojection 760 of the frame data and display 770. The reprojection may include warping as described above.
[0202] The difference in layered coding schemes is that after rendering 710, each frame is pre-processed 721 to produce a base layer. This typically involves downsampling, i.e. reducing the spatial resolution of the image data.
[0203] Then, at step 723, the base layer is encoded using a base codec. This can be accomplished using any standard codec suitable for frame data, such as h.264, HEVC, AV1, or VVC (typically a single-layer codec) for video data. Alternatively, the base codec itself can be a layered codec. The base codec can be lossy, so encoding and then decoding the base layer does not accurately restore the original data of the base layer.
[0204] Finally, at step 725, the encoder generates at least one enhancement layer using data from the coded base layer. This typically involves a comparison between the original frame generated by rendering 710 and the decoded coded base layer generated at step 723. For example, this can be done using information about Figure 4A The LCEVC technique described in the previous embodiment can be implemented. Step 725 can also include performing a coding method on the enhancement layer to produce a coded enhancement layer. This can include inter-coding the difference between the enhancement layer of the current frame and the enhancement layer of the previous frame. In addition or alternatively, a general coding method for data transmission can be performed, such as adding a check bit.
[0205] Layered coding schemes also differ from serial coding schemes in that transmission and decoding can be separated into two parallel processes or streams.
[0206] When the encoded base layer is transmitted, the first parallel process begins at step 731. This may begin immediately after the base layer is encoded at step 723. The receiving entity stores the encoded base layer for each frame in a jitter buffer 741 until it is ready to decode the base layer at step 751. The jitter buffer 741 may be smaller than the jitter buffer 740 because the base layer is a downsampled version of the frame stored in the jitter buffer 740 for the serial coding scheme. Thus, the one or more encoded frames in the jitter buffer 741 take up less space than the one or more full-resolution frames in the jitter buffer 740 and take less time (and / or less computing resources) to decode than the one or more full-resolution frames in the jitter buffer 740.
[0207] When one or more enhancement layers are transmitted, a second parallel process begins at step 733. This may begin after one or more enhancement layers are generated and / or encoded at step 725. When the encoded enhancement layers are received at the target device, they are decoded 753 to recover the enhancement layers.
[0208] The first and second parallel processes are then merged at a compositing step 755. In compositing, the base layer and the enhancement layer are combined to produce a reconstructed reproduction of the original frame as rendered in step 710. In the case of a lossless layered codec, the reconstructed reproduction is equivalent to the original frame, while in the case of a lossy layered codec, the reconstructed reproduction should have similar characteristics to the original frame, at least in terms of resolution, etc. For example, compositing may involve upsampling the base layer and using values from the enhancement layer to correct for compression losses in the upsampled base layer caused by the base codec. Upsampling of the base layer may alternatively be incorporated at the end of the second parallel process before the compositing step 755.
[0209] The hierarchical scheme ends with reprojection 760 and display 770 of the frame data. These steps are the same as for the full serial scheme.
[0210] In a fully serial scheme, frames are encoded and decoded at full resolution, which takes longer than encoding and decoding at a lower base resolution. In addition, jitter buffering takes longer or consumes more processing resources and requires more space at full resolution than at a lower base resolution.
[0211] On the other hand, layered coding schemes require additional processing, including pre-processing and post-analysis at the encoder (generating one or more enhancement layers), as well as decoding of enhancement layers and combining of decoded layers at the decoder. However, the additional complexity of layered coding is offset by the time / resources saved by encoding and decoding at a lower base resolution. Therefore, the overall latency of layered coding is lower than full-resolution encoding, especially when the processing of enhancement layers and base layers is performed in parallel as much as possible.
[0212] Layered coding also provides new degrees of freedom when encoding and transmitting data. Enhancement data can be discarded in the event of a sudden drop in bandwidth, thereby reducing delay jitter. In addition, delay gains can be obtained by using specific constructs (including tiles and slices) in layered coding (such as the LCEVC standard). More specifically, the base layer encoding of a frame area can be decoded in a series of parallel slices. In implementations such as LCEVC, only one slice is required to calculate the enhancement layer, so once the encoder decodes one slice of the base layer, the calculation of the enhancement layer can begin. Similarly, the base layer can be transmitted as a separate coded slice instead of a complete coded frame, which enables the base layer of the frame to be transmitted earlier. In addition, the enhancement layer enables encoding at a higher bit depth. For example, LCEVC supports 14-bit depth maps and HDR, even if the base encoder only supports 8-bit or 10-bit base layers.
[0213] Figure 8 VC-6 encoding is schematically shown, which has been described in more detail in the standard document VC-6: SMPTE VC-6ST-2117 and WO2019 / 111010, all of which are incorporated herein by reference. Similar to LCEVC encoding, VC-6 uses downsampling to transmit the encoded base layer at a reduced resolution and uses one or more enhancement layers to restore the original data at the decoder.
[0214] Fig. 9 An example networked system for distributed rendering, which may be referred to as a "content rendering network," is schematically illustrated.
[0215] The first CRN node 910 performs a first pass volume rendering of the relevant view area in the virtual environment. The rendering is performed once for multiple viewing devices and includes the most computationally expensive parts of the rendering, such as ray tracing. The rendering can be performed at a relatively low frame rate (such as 25fps) and can include motion information for upsampling at downstream nodes in the network.
[0216] The frame rendered by the first CRN node 910 is transmitted to an edge CRN node 920 that is closer to a specific user or a specific user group. If there are multiple users or multiple groups of users, the frame rendered by the first CRN node 910 can be transmitted to multiple edge CRN nodes 920, each edge CRN node is associated with a corresponding one or more users.
[0217] Each edge CRN node 920 performs a second pass rendering of the user viewport (i.e., the area of the virtual environment visible to a particular user) and encodes a frame having one or more of an image layer and a depth layer for displaying a 3D environment as described above. This rendering can be performed once per user group, for example, if the data is in the MPEG immersive video (MIV) format defined in ISO / IEC FDIS 23090-12, which is incorporated herein by reference. Alternatively, if the data is in another video format, the rendering can be performed once per user. The rendering can be performed at a higher frame rate (e.g., 50fps) using interpolation based on the motion information generated by the first CRN node 910.
[0218] The frames rendered by the edge CRN node 920 are then streamed to the user device 940 via the network 930. The network 930 may have an air interface, such as between a user headset and a user base unit.
[0219] The user device 940 performs low complexity decoding at the display resolution and display frame rate. The depth layer can be used at this stage to apply spatial warp reprojection.
[0220] Even when network speeds are as low as 30Mbps, the inventors found that the encoding technique described above can support streaming that is resilient to packet loss and can compensate for any network limitations by dropping layers, thereby ensuring acceptable VR streaming at the highest definition possible, as supported by the network and user device.
[0221] The following forms part of the description.
[0222] XR and the Metaverse: How to achieve interoperability, wow effects, and mass adoption
[0223] Leaving aside the inevitable hype, there is no doubt: in one form or another, the Metaverse and Extended Reality (XR) are here to stay and will be the next generation of the Internet. It pays to better understand them and the corresponding requirements and technical implications.
[0224] However, although affordable XR devices can meet certain needs, achieving high-quality immersive XR experience still requires overcoming a bottleneck, which usually requires a large amount of data to support, such as through "split computing" or "cloud rendering."
[0225] “In this application, ‘split compute’ is the process of dividing computation across two or more devices, with 3D rendering performed on a different device (either nearby or remote) than the display device. Rendering can also be further split across multiple resources (‘hybrid rendering’). The resulting rendered frames are then streamed (‘cast’) to the display device.
[0226] The inventors considered several technical challenges related to what happens before and after 3D rendering:
[0227] Three key user requirements for mass market adoption: lighter devices, interoperability, and no compromise in quality.
[0228] Moore’s Law and Koomey’s Law cannot close the graphics processing power gap between XR devices and gaming hardware. Therefore, inventors have realized that high-quality interoperable XR applications can be run by splitting computing and ultra-low latency video projection.
[0229] The data compression challenges of delivering volumetric objects to rendering devices, and subsequently projecting high-quality ultra-low latency video to XR displays, will require a low-complexity multi-layer data encoding approach that makes efficient use of available resources and is particularly well suited for 3D rendering engines.
[0230] ●Two recently standardized low-complexity multi-layer codecs - MPEG-5 LCEVC and SMPTE VC-6 - enable high-quality Metaverse / XR applications within practical constraints. LCEVC enhances any video codec (e.g. h.264, HEVC, AV1, VVC) to enable ultra-low latency delivery of high-quality XR video (even video+depth) within strict wireless constraints (<30Mbps), while also reducing latency jitter due to the ability to dynamically drop top-level packets in the event of a sudden drop in network bandwidth. Applying the VC-6 toolkit to point cloud compression enables the large-scale distribution of realistic 6DoF volumetric video with unprecedented quality, with excellent feedback from end users and Hollywood studios.
[0231] ●XR rendering and video projection are 1:1 operations in many scenarios, rendering content for individual users. Therefore, unlike broadcasting and video streaming, power usage may grow linearly with the number of users. In this case, sustainability and energy consumption are important considerations, which strengthens the case for adopting low-complexity multi-layer encoding.
[0232] ●The ability to properly address data volume challenges and leverage existing methods and standards in a timely manner will be one of the key factors in driving the XR metaverse to large scale.
[0233] What exactly is the metaverse?
[0234] The Metaverse promises to break down geographic barriers by providing new virtual degrees of freedom for work, play, travel, and social interaction. But what is it actually, and why now?
[0235] John Riccitiello, CEO of Unity, has given a reasonable definition of the Metaverse (derived from “beyond the universe”): “It is the next generation of the Internet that is always live, primarily 3D, primarily interactive, primarily for social purposes, and primarily persistent”. In effect, it is a new type of Internet user interface that can interact with other parties and access data as a seamless and intuitive enhancement to the real world in front of us, inspired by the online 3D worlds that multiplayer video game users are already familiar with. Metaverse solutions generally aim to replicate the 6 degrees of freedom (6DoF) and inherent depth in the real world, rather than the flat 2D interfaces we are used to browsing online. The Metaverse is often used interchangeably with the term “Extended Reality” (XR, a combination of Virtual Reality and Augmented Reality), and it promises to be the biggest step change in the development of network computing since the introduction of the World Wide Web in the 1990s, taking us from relatively shallow text-based interfaces to intuitive “browsing” of hypertext with multimedia components.
[0236] This new paradigm (metaverse, cyberspace, 6DoF 3D) is different and powerful, and will replicate the inherent depth and intuitiveness of the real world rather than the flat hypertext interfaces we use today.
[0237] Why now?
[0238] Gamers have been playing in 6-degree-of-freedom (6DoF) 3D worlds for decades, so one might legitimately ask why in the last two decades we haven't been doing e-shopping in 3D and / or using Excel spreadsheets in 6DoF 3D? If GTA V lets us roam freely around a pretty decent simulation of Southern California, then why are Microsoft Office, Amazon, Instagram, and SAP still so boringly flat?
[0239] In fact, the transition of non-gaming applications to 6DoF 3D applications has until now been hampered by several technical barriers that are now close to being overcome.
[0240] Without a doubt, a notable barrier to immersive XR to date has been bulky VR headsets that were unable to meet the resolution and frame rates required for a smooth experience. Some may remember the feeling of motion sickness when trying to use virtual headsets in the early nineties. Others may remember comments about the uncool form factor of the products, or concerns about VR users being “isolated” in their digital worlds. Commercially available headsets are now lighter, enable non-isolated augmented reality, and finally work well enough from a visual perspective that the average consumer may find them appealing. Upcoming headsets will further address some of the remaining pain points that affect realism, such as projecting an image of the user’s face to reduce the sense of “physical separation” felt by people outside the headset, adding sensors for real-time eye tracking and facial tracking for more realistic social interactions in the metaverse, and including varifocal display technology to solve the focus problem - this is technically called “vergence-accommodation conflict”, where the display forces your focus to be fixed at about 1.5m away regardless of actual depth. In short, we are rapidly approaching headsets that are good enough for mass adoption.
[0241] However, suboptimal headsets aren’t the only obstacle to the rise of the Metaverse. 6DoF 3D digital worlds don’t necessarily require immersive XR, as 3D video games have thrived for decades on traditional flat screens. Why haven’t we seen office productivity tools, e-commerce sites, or other software applications follow in the footsteps of video games?
[0242] Another significant hurdle to non-gaming uses of 3D worlds so far has been the difficulty of integrating intuitive 6DoF controls with the peripheral input devices we typically use on laptops, tablets, and phones.
[0243] Keyboards, mice, trackpads, and touchscreens are great, but none of them are particularly well suited for 6DoF navigation of 3D environments. Beyond game controllers, the holy grail of tamed, intuitive 6DoF control has finally matured in recent years: gesture controls driven by cheap cameras and infrared proximity sensors now appear in many devices.
[0244] This type of immersive interface will not be limited to XR displays we wear on our faces. Even traditional flat screen devices such as laptops, phones and large screens are equipped with cameras and infrared sensors capable of gesture detection and head tracking. Today, for just a few thousand dollars, one can buy a glasses-free autostereoscopic 8K screen that tracks head and eye movements to display a stereoscopic scene that appears to pop out of the screen as we watch it, effectively acting as a window into a 6DoF 3D digital world.
[0245] Depending on the application, the development of immersive 3D interfaces with integrated gesture control will be either evolutionary or transformative, similar to what we have seen with hypertext and touchscreen interfaces over the past few decades. Hypertext and traditional (“rectilinear”) video sources will also continue to exist in the Metaverse: they will simply be displayed in the context of immersive user interfaces and manipulated intuitively in 6DoF.
[0246] In fact, why hasn’t it already?
[0247] People want more intuitive and immersive digital applications, and everything described here is already technically possible today with relatively inexpensive gear. Many people have already started buying those gear: more than 20 million homes will soon have at least one of the latest-generation headsets. As mentioned above, lighter AR glasses are already capable of augmented reality, and a new generation of improvements is expected soon, driven by the entry of other large tech giants into the field. Some autostereoscopic displays that don’t require glasses are approaching the cost of traditional displays. Almost all flat-screen devices are equipped with cameras and infrared sensors that enable gesture tracking and head tracking.
[0248] If there is consumer demand for more intuitive and immersive digital applications, and all of the technologies just described are available and affordable in the market today, then why haven’t we entered the Metaverse yet? What are the remaining obstacles?
[0249] Key user requirements for mass market adoption
[0250] In adoption cycle terms, XR consumers have progressed from innovators (pre-2020) to early adopters (2020-2022), with new compelling XR experiences appearing every day. Does it all work well enough and seamlessly enough to appeal to the proverbial “early majority” (which inevitably includes non-gaming use cases)?
[0251] Mass adoption requires us to meet three key user needs:
[0252] XR devices must be small and lightweight, ideally as light as a pair of glasses, with a maximum peak power consumption of 1-2 watts. This leaves very little space for electronics and batteries, which are not suitable for processing high-quality 3D rendering in real time. Small form factors are a prerequisite for achieving scale, as people cannot spend hours a day looking at the screen.
[0253] "Wearing a gaming laptop (or even a powerful phone) on your face."
[0254] Metaverse applications must be interoperable like web pages, meaning that different viewing devices with different computing power (including XR headsets / smart glasses and more traditional TVs, phones, tablets or laptops) must be able to access them. Interoperability is obviously important for service providers to ensure that their services have the widest possible user base.
[0255] · Quality of experience is non-negotiable: end users expect a visually stunning, realistic, smooth and immersive experience. This is not surprising given that video game players today take it for granted that they can see beads of sweat dripping from the foreheads of virtual football players during an average football match. While traditional 2D user interfaces may tolerate some imperfections, the key to the Metaverse is precisely the illusion of “presence”, which necessarily requires top-notch audio-visual quality, high-quality 3D objects, realistic lighting, high resolution, high frame rates and low-latency real-time. Grainy objects and slow frame rates are not enough and may even cause (motion) discomfort to the user. Unsurprisingly, the aforementioned McKinsey analysis shows a high correlation between the realism of the experience and the frequency of usage, which drives companies to create more realistic experiences.
[0256] Having identified the above requirements, the inventors have recognized that the first two require seamless “split compute,” meaning that most of the computation—especially 3D rendering—must be performed on another device, perhaps even in the cloud, so that, when needed, you can time-share a more powerful GPU than you would otherwise be able to equip. The resulting rendered graphics frames must be streamed to that device as high-resolution, high-frame-rate, ultra-low-latency stereoscopic video. As we’ll see in the next section, the limitations associated with XR projection are pretty much the worst nightmare for video encoding, especially when wireless connectivity (wi-fi or 5G) is involved. Fortunately, as we’ll also see, there’s a standard solution that makes this possible within practical limitations.
[0257] In addition to enabling complex graphics to be displayed on low-power display devices, remotely computing and then streaming video to the display device is the best way to ensure interoperability among Metaverse destinations, which is the second key requirement. Any device, regardless of available computing power, will be able to connect to the remote server and receive the video stream, including mobile phones, TVs, autostereoscopic displays, XR headsets, and XR smart glasses. The server can also seamlessly handle the differences in the quality each device is capable of displaying, adapting the format and quality of the video stream to the specific user device and network conditions: in a sense, video can act as the new HTML, with metadata, enhancement layers, and auxiliary data channels acting like hypertext objects that enrich baseline web pages.
[0258] The third requirement implies the emergence of many new large datasets on top of traditional audio, images, and video (which account for over 90% of current internet traffic) that need to be efficiently exchanged in the Metaverse. Examples include immersive audio, point clouds of various natures, meshes, textures, stereoscopic video, video + depth maps for reprojection, etc. The key difference between a Metaverse destination and a video game is that volumetric objects can be much harder to preload and have to rely on real-time access, and we will be streaming those large 3D assets in real-time much more frequently than games.
[0259] Separating rendering from display
[0260] Visiting the Metaverse won’t require just ultra-immersive XR headsets. XR will combine headsets (more than 23 million by 2023, according to IDC), lighter glasses-sized XR viewers, and autostereoscopic displays with billions of more traditional phones, tablets, PCs, and TVs. Traditional flat screens will display virtual destinations in a similar way to how 3D games are displayed today.
[0261] There’s a problem here. To experience realistic 3D graphics, hardcore gamers need to equip themselves with enough graphics processing power (i.e., “latest-generation discrete GPU with active cooling”), while the typical mass-market XR user might just own a headset with 50 times less graphics processing power than a typical gaming PC and a work laptop with equally less powerful integrated graphics.
[0262] In other words, since users won’t accept “wearing a Playstation 5 on their face” and we can’t assume everyone will be willing to buy a gaming PC, the wise thing to do is to design the Metaverse experience based on the lowest common denominator of compute and power consumption.
[0263] That lowest common denominator — a lightweight XR device that consumes 1 watt of power — will be far from achieving real-time decoding of realistic volumetric objects and high-frame-rate 3D rendering at stereoscopic 4K resolution. The laws of physics and silicon technology don’t leave us any hope: the latest generation of gaming GPUs are bulky and consume between 150 and 300 watts of power to perform these tasks. When your average phone consumes more than 4 watts, it starts to heat up. Add to that the weight of the cooling fans and batteries, and you can see why I joke about wearing a PS5 on your face. As a result, mobile GPUs cannot achieve the level of rendering quality that end users expect: you can render simple games like Beat Saber, but you can’t immerse users in realistic 6DoF experiences. Moreover, headsets with tier 1 mobile capabilities are still too bulky for the average user to accept wearing for several hours a day.
[0264] There is a more than 50x gap between the graphics processing power of lightweight devices and the visual quality that hardcore gamers are accustomed to experiencing.
[0265] Most modern mobile SoCs use 5nm silicon technology in terms of transistor size, and long-term R&D shows that 2nm may be the physical limit of silicon. 1nm means about 10 electrons, so we are reaching a scale where the quantum nature of electrons makes them jump through the silicon gate regardless of the applied voltage. In addition, in the process of further miniaturizing transistors from 5nm to 2nm, the transistor density will increase, but the power consumption of each transistor is unlikely to decrease further due to the quantum tunneling effect. We have seen that according to TSMC, with the transition from 7nm to 5nm, the transistor density has increased by 80%, but the computing performance per watt has only increased by 15%. This effect is also known as Koomey's Law, named after Stanford Professor Jonathan Koomey.
[0266] Unlike Moore’s Law, which tracks the evolution of peak computing performance, Coomey’s Law tracks the evolution of computing performance per watt. Unfortunately, Coomey’s Law measures a clear decline in our ability to perform more computations for the same unit of power: from a 100-fold improvement per decade in the 1970s (doubling every 18 months), we were only achieving about 15-fold improvements per decade in the first two decades of the 21st century (doubling every 2.6 years), and now—for the quantum physics reasons mentioned above—our rate of improvement has even further plateaued.
[0267] Because of this plateau, a 50x efficiency gain in processing power seems unlikely to be achieved in the next decade (or perhaps ever).
[0268] The difference in processing power also ties into interoperability, which is the second key user requirement. If we want to truly be a “new web” for Metaverse applications, we must ensure that any XR device can properly run them through seamless interoperability. How do we effectively deal with the orders of magnitude differences in available graphics rendering power that end-user devices may have?
[0269] As further improvements in chip technology acted as a placebo, the ideas of the solution inventors were based on mainframes and client-server archetypes. In the 1980s, a mainframe could control real-time user interfaces for dozens or even hundreds of user terminals. This concept can be adapted to enable immersive digital worlds with seamless interoperability: for simple tasks, the mainframe could be the phone in your pocket, and once you need powerful rendering power, you can seamlessly tap into the GPU resources of the nearest available shared computing node.
[0270] Performing rendering outside the XR device also means streaming the volumetric objects to the rendering device, which is more likely to have a good connection to the internet, especially if located in a data center. The XR device (which connects to the internet wirelessly via Wi-Fi or 5G, typically <50Mbps due to distance from repeaters, obstructions, concurrency, and packet loss caused by interference) only needs enough connection power to handle the video stream.
[0271] For Metaverse applications with high demands for 3D graphics, it is inevitable that rendering and display must be separated. A rendering device that can boot up and consume enough power (whether it is a handheld device, a nearby PC or a powerful server somewhere "in the cloud") will receive the compressed virtual objects, decode them, run the application and render the viewport, that is, what the end user is looking at at any one moment. Even the rendering itself may actually be split into different stages, performed on different machines. For example, a more powerful GPU can calculate lighting information, perhaps even using expensive algorithms such as ray tracing or path tracing, which require more than 1000 times more computing power than typical game engine lighting - while a less powerful nearby GPU can calculate the final rendering result. In other cases, for complex 3D scenes with many users, the expensive lighting calculations can be performed "centrally" for multiple users, generating intermediate volumetric data that is then transmitted to multiple local edge nodes to render the final viewport for each individual user. Rendering servers can thus be organized in a dynamic hierarchy, similar to the way cache servers of a content distribution network enable large-scale network content delivery: such a Metaverse infrastructure can be called a "content rendering network" (CRN). The resulting rendered viewport is then streamed to a lightweight XR display device via ultra-low latency video. As a result, the XR device only needs to manage its sensors, process / send data on what the user is doing, and decode the received high-frame-rate, high-resolution video.
[0272] This is certainly feasible to achieve at realistic quality within 1 watt, especially by using low-complexity video compression methods that fit within bandwidth and processing limitations.
[0273] Splitting the Computing Challenges of Making XR Viable at Scale
[0274] The importance of properly handling large amounts of data when it comes to mass adoption of XR cannot be underestimated. Data compression and processing that meets the quality, bandwidth, processing power, and latency constraints of solid, latest-generation network connections is key to the Metaverse, and the inventors identified three main technical challenges:
[0275] 1. Compressing and streaming volumetric objects to and between rendering devices in a suitable manner requires new encoding methods.
[0276] 3D objects can be efficiently compressed / decompressed using suitable lossy coding techniques to increase the chances that point clouds, textures, and meshes can be streamed efficiently. For hybrid rendering use cases, intermediate byproducts should preferably be efficiently transferred between different nodes of the content rendering network. Previously, occasional monolithic game world downloads were mostly losslessly encoded. A flexible software layered encoding approach suitable for efficient massively parallel execution would be advantageous.
[0277] 2. Ultra-low latency video encoding of 4K 72fps and above at bandwidths below 30-50Mbps to cope with actual Wi-Fi / 5G sustained rates.
[0278] In order to not make the experience “jerky,” the latency of video projection better be low and consistent. Some latency can be managed by rendering the viewport based on a prediction of where the system estimates the user will be looking after the delay, and then doing a last-millisecond reprojection based on where the user is actually looking. But we’re still talking about 20-30 milliseconds, with tight processing limits. In those ultra-low latency scenarios, many of the latest video compression efficiency tools can’t be used, which increases the bandwidth required to deliver a given quality. At the same time, wireless transmission means there are tight bandwidth limits (<50Mbps) before packet loss starts to produce unsustainable latency jitter. In addition to choosing the right protocol (with some level of forward error correction if possible), it’s particularly useful to combine the most efficient video codec available with low-complexity compression enhancement methods that can reduce bitrate and generate data layers that can be dropped on the fly in the event of bursty congestion. Note that an example of such an approach is already provided by the MPEG-5 LCEVC (Low Complexity Enhancement Video Coding) coding enhancement standard, which improves compression efficiency, reduces overall processing requirements, and generates a layer of enhancement data that can be discarded in the event of bursty network congestion without affecting the decoding of the base layer video, thereby reducing the risk of latency jitter.
[0279] 3. A strong network backbone / CDN / wireless network is established between the rendering and XR display devices, and sufficient cloud rendering resources are available (such as "CRN").
[0280] For split computing, the non-negotiable quality of experience may be at risk if the network and shared resources are not powerful enough. Telecom operators, cloud service providers, and CDNs will have to work overtime to ensure that as many people as possible are in the situation highlighted in point 2 above, that is, users have access to a powerful GPU server somewhere after a few milliseconds of network latency, and at least 30-50Mbps end-to-end reliable connection is available. Currently, less than 10% of the population in developed countries subscribes to fiber-to-the-home (FTTH) level connectivity services and installs the latest generation of Wi-Fi routers. In addition, the existing GPU server infrastructure in the cloud can only support a small fraction of them.
[0281] The Case for Low Complexity Layered Coding
[0282] LCEVC is a new hybrid multi-layer coding enhancement standard from ISO-MPEG. It is codec-agnostic in that it combines a lower resolution base layer encoded using any traditional codec (such as h.264, HEVC, VP9, AV1 or VVC) with a residual data layer that reconstructs the full resolution. The LCEVC coding tools are particularly well suited to effectively compress details from a processing and compression perspective, while making effective use of traditional codecs at lower resolutions, taking full advantage of the hardware acceleration available for that codec, making it more efficient.
[0283] LCEVC is proposed as a key enabler for ultra-low latency XR streaming, bringing benefits such as:
[0284] 1. Brings significant compression benefits to all available hardware codecs (h.264, HEVC or AV1). The LCEVC compression enhancement toolkit operates primarily in the spatial domain, so its benefits are very effective for ultra-low latency encoding, where, by definition, most temporal compression techniques (such as leveraging bidirectional predictive reference pyramids for subsequent frames) cannot be used. In addition to the efficiency gains relative to non-enhanced encoding being significant (typically in the range of 30-40%), the LCEVC compression enhancement toolkit is also very effective for ultra-low latency encoding.
[0285] In addition to the above, it is important that the absolute bandwidth is able to meet the constraints when using LCEVC. LCEVC can achieve high frame rate stereoscopic 4K in the 25-50Mbps range, making high-quality wireless projection and cloud XR streaming more feasible. At around 30Mbps, the difference in subjective quality can often mean the difference between "unwatchable" and "fit for purpose".
[0286] 2. Due to the inherent multi-layer structure of LCEVC, delay jitter is uniquely reduced. LCEVC data can be transmitted in a separate, lower-priority data channel and can be dynamically discarded in case of network congestion - even during frame transmission, after encoding - without affecting the base layer and with minimal impact on visual quality (i.e., image quality will look softer within a few frames without obvious spatial correlation impairments). This enables immediate handling of frequent oscillations and packet loss problems inherent in wireless transmission, avoiding the accumulation of tens of milliseconds of delay jitter every time the bandwidth suddenly drops. In addition, from the perspective of overall end-to-end system delay, LCEVC is able to transmit and decode the base layer while encoding the enhancement layer, providing additional degrees of freedom to reduce encoding delays.
[0287] 3. Low processing complexity, enabling LCEVC to be used immediately through efficient heterogeneous processing (e.g., 4K 72fps with limited resource overhead compared to native hardware encoding / decoding) and - with dedicated silicon acceleration, IP blocks already available - making it possible to efficiently encode / decode very high resolutions, frame rates, and bit depths (e.g., 16K 120fps 14-bit) with very small silicon area, enabling wireless delivery of retina display quality XR streams on lightweight devices.
[0288] 4. High bit depth high dynamic range (HDR). HDR is very important for realistic immersive experience, and according to the previous analysis, the higher the bit depth, the better. LCEVC utilizes a 16-bit acceleration pipeline, so it can encode HDR video in 12 or 14 bits without any processing and / or bit rate overhead, while still using available 8 or 10-bit base layer codecs.
[0289] 5. The ability to send 14-bit depth maps along with the video enables more realistic depth-based reprojection and eye-tracking-based zoom adjustments, resulting in a better overall visual experience. In the absence of depth maps, reprojection is performed immediately before display - to compensate for the difference between the actual viewport and the predicted viewport during rendering - via "flat" reprojection. This is one of the main visual obstacles caused by system latency. The availability of depth maps is also useful for eye tracking and upcoming zoom headsets, where the depth of VR objects viewed by the user can be understood on the local device and the display focus can be adjusted accordingly to avoid convergence accommodation conflicts (i.e., an important step towards visual realism). LCEVC's fast software-based encoding tools (performed through a 16-bit acceleration pipeline) can effectively pass depth maps along with the video, as long as sufficient resources and / or bandwidth are available. From a processing perspective, while making video+depth maps practical, the overall bandwidth reduction achieved by LCEVC is also critical to allowing depth map transmission within the available bandwidth constraints.
[0290] Importantly, these benefits can be achieved while keeping average system latency to a minimum. Since LCEVC inherently splits computations into separate independent processes (a kind of “vertical striping”) without compromising compression efficiency, end-to-end pipelines can be built to achieve maximum parallel execution, such as Figure 7 shown.
[0291] Therefore, a proper ultra-low latency implementation of LCEVC enhanced projection could produce a beneficial combination of visual quality enhancement, minimal average latency, and reduced latency jitter, making the case for local and / or cloud-based split compute with "retina quality" 120fps wireless XR streaming to lightweight devices. This would be an example of what we mean by "bringing the digital future to life," enabling creatives and digital businesses to design a whole new world of impactful experiences and user interfaces.
[0292] Elaborating further on LCEVC’s multi-layer and low-complexity coding toolkit, SMPTE VC-6 (SMPTE ST 2117-1) is another coding standard that is particularly well suited for Metaverse datasets, thanks to its recursive multi-layer scheme, the use of S-trees, and its scalable coding toolkit, which uniquely includes massively parallel entropy coding and in-loop neural networks. Figure 8 Demonstrating the VC-6 concept.
[0293] In addition to its applications in image / texture compression, professional video workflows, and AI media indexing acceleration, the VC-6 variant for point cloud compression is also used to compress the PresenZ volumetric video format, the only media format capable of rendering realistic 6DoF volumetric experiences. See WO 2023 / 047119, which is incorporated herein by reference. Several media examples that use the PresenZ player and benefit from this compression are available on the Steam store. Uniquely, the codec not only compresses huge data sets (more than 50 gigabits per second) to enable them to be delivered to end users, but also demonstrates ultra-fast software decompression capabilities, capable of real-time decoding while leaving enough idle processing resources for real-time rendering at high resolutions and high frame rates.
[0294] In the Metaverse, 3D object decompression, graphics rendering, and viewport compression will all be closely linked. In this case, hierarchical data structures are particularly advantageous. When using hierarchical data structures, you can get a higher quality texture or polygonal model from a lower quality version, adding some data that specifies details. As you get closer to an object, you get more data. Furthermore, if decoding is super fast and massively parallel, it may be better to operate directly on compressed data instead of uncompressed data, thereby maximizing the degree of realism that can be processed for a given memory bandwidth.
[0295] In addition to these benefits, the hierarchical multi-layer signal format can also help in the field of AI-based indexers, making them more accurate and faster in classification tasks, which is of great benefit to Metaverse Robotics. Lower quality signal reproduction allows for quick detection of areas that require further classification, and high-resolution decoding of regions of interest is completed only for areas of particular interest. Multiple classifiers operate in parallel in different ways on the same compressed dataset, which may be distributed across multiple nodes.
[0296] The types of data in the Metaverse are likely to increase: specific applications will need their own mix of point clouds, meshes, auxiliary data for physically based rendering engines, textures, and augmented video. They will need forms of progressive and region-of-interest partial decoding. They may have to deal with data with variable bit depths, sometimes higher than 10 bits (e.g., for depth maps or point cloud coordinates).
[0297] Low-complexity multi-layer encoding that can be efficiently executed in software on general-purpose parallel hardware is proposed as a scalable solution that can interoperate with a variety of devices with different displays and processing capabilities. Flexibility is desirable, for example, to enable a similar encoding scheme to run both as an ASIC / FPGA with very low gate count and as software on general-purpose graphics hardware with low processing power.
[0298] Such design standards are at the core of the video standards MPEG-5LCEVC and SMPTE VC-6, while the MPEG Immersive Video (“MIV”) and Volumetric Data Compression projects within MPEG may also help address some of these issues. Properly leveraging these latest coding standards could make high-quality XR feasible and interoperable.
[0299] Laying fiber optic cables, rolling out 5G and building new data centers will also be necessary, but these alone will not solve the problem. All available tricks will have to be used.
[0300] Higher quality with less energy
[0301] Sustainability was also a consideration. In short, we needed compelling, non-disruptive XR workflows that kept energy bills to a minimum.
[0302] Providing reliable numbers is difficult, and laudable initiatives like the “Greening of Streaming” continue to emerge to provide more clarity on what to measure and how to measure it. That said, the environmental impact of media delivery is certainly relevant and its impact is growing rapidly. About 2% of global greenhouse gas emissions come from data centers, which use about 200 terawatt hours of electricity, equivalent to the electricity used by the global airline industry. When this energy consumption is taken into account and added to the rest of the delivery elements of video streaming, plus we consider the 50% annual growth in usage, and remember that Moore’s Law is no longer helping much in terms of power consumption per unit of processing. Add in the Metaverse… and that number will soon be a double-digit percentage of total global energy consumption.
[0303] Since XR rendering and video projection are typically 1:1 operations, power consumption may grow linearly with the number of users, unlike broadcasting and video streaming: low-processing approaches may be important to the overall sustainability of the Metaverse, adding another brick to the case for low-complexity multi-layer encoding.
[0304] Making the Metaverse a Reality
[0305] The ever-expanding volume of diverse 3D data and the tight video streaming limitations of split computation are key challenges that need to be addressed to ensure that immersive worlds can be delivered to end users in an interoperable manner and at scale.
[0306] As highlighted above, rapid adoption of existing multi-layered encoding standards and development of new standards from the same IP / toolkit can greatly accelerate the development of high-quality and interoperable Metaverse destinations.
[0307] The ability to properly address data volume challenges and leverage existing methods and standards in a timely manner will be one of the key factors in driving the XR metaverse to large scale.
Claims
1. A networked system for generating a frame sequence for rendering a dynamic 3D scene, the system comprising a first rendering node and a second rendering node, in: The first rendering node is configured as follows: Generate the first part of the rendered frame sequence; as well as performing layered encoding on each first partially rendered frame to generate a sequence of encoded first partially rendered frames; as well as transmitting the encoded first partial rendering frame sequence to the second rendering node; as well as The second rendering node is configured as follows: Obtaining a first partial rendered frame sequence of the encoding from the first rendering node; as well as A second partial or full rendered frame sequence is generated based on the encoded first partial rendered frame sequence.
2. The networked system according to 1, wherein the second rendering node is configured to decode the encoded first partial rendering frame sequence to obtain the first partial rendering frame sequence.
3. The networked system of claim 1 or claim 2, wherein the second rendering node is configured to perform layered encoding on each second partially or fully rendered frame to generate an encoded sequence of second partially or fully rendered frames.
4. The networked system of claim 3, wherein the first rendering node is configured to perform layered encoding according to a first encoding scheme and the second rendering node is configured to perform layered encoding according to a second encoding scheme, the first encoding scheme being different from the second encoding scheme.
5. The networked system of any one of claims 1 to 4, wherein the frame rate of the second partially or fully rendered frame sequence is greater than the frame rate of the first partially rendered frame sequence. 6 . The networked system according to claim 1 , comprising a display device, wherein a communication delay between the first rendering node and the display device is greater than a communication delay between the second rendering node and the display device.
7. The networking system according to claim 6, in: The first rendering node is configured to obtain a viewing position from the display device before generating the first partial rendered frame sequence; and The second rendering node is configured to obtain an updated viewing position from the display device before generating the second partially or fully rendered frame sequence.
8. The networked system of any one of claims 1 to 7, wherein generating the first partial sequence of rendered frames requires more processing resources than generating the second partial or full sequence of rendered frames.
9. A networked system according to any one of claims 1 to 8, wherein each frame comprises image data and depth map data, and the system comprises a third rendering node, the third rendering node being configured to generate a third fully rendered frame sequence by performing time warping and / or depth correction on the second partially or fully rendered frame sequence using the depth map data.
10. The networked system of any one of claims 1 to 9, wherein each frame comprises point cloud data, wherein each point of the plurality of points has a 3D position and one or more attributes. 11 . The networked system of claim 10 , wherein the second rendering node or the third rendering node is configured to calculate depth map data based on the 3D position of the point.
12. A networked system according to any one of claims 1 to 11, in: The first rendering node is configured to generate one or more first partial rendering frame sequences for a first number of users or display devices; as well as The second rendering node is configured to generate one or more second partially or fully rendered frame sequences for a second number of multiple users or display devices, wherein the second number is less than the first number.
13. A networked system according to any one of claims 1 to 11, wherein the networked system dynamically selects how many nodes and which specific nodes to use for the rendering process in response to at least one metric, the at least one metric include: A measure of the complexity of the rendering task to be performed, Based on a measure of the spare capacity available at each node, Based on a measure of the position of the display device relative to the node network, Based on a measure of the round-trip delay between a node and a display device, A measure based on the available bandwidth between nodes and from the nodes to the display device, and A metric based on the number of different displays requesting to render the same 3D scene over a range of viewpoints.
14. A networked system according to any one of claims 1 to 13, wherein the partially rendered frame is encoded as volumetric data comprising one or more of: Point cloud data, Grid data, Texture data.
15. The networked system of any one of claims 1 to 14, wherein the partially rendered frame comprises light field data.
16. A networked system according to any one of claims 1 to 15, wherein the partially rendered frame comprises spatial characteristics allowing calculation of sound behaviour in 3D space.
17. The networked system of any one of claims 1 to 16, the partially rendered frame being encoded at least in part using a lossy encoding method.
18. The networked system of any one of claims 1 to 17, wherein the partially rendered frame is encoded using at least in part a layered encoding method.
19. The networked system of any one of claims 1 to 18, wherein a subsequent node receives only a portion of the data encoded by the first rendering node in response to a particular location of the one or more viewpoints of the fully rendered frame to be calculated.
20. A networked system according to any one of claims 1 to 19, wherein a subsequent node decodes only a subset of the encoded data generated by the first rendering node and received by the subsequent node that is necessary to fully render a particular field of view that it is rendering at any point.
21. A networked system according to any one of claims 1 to 20, wherein the partially rendered frame is encoded at least in part using a point cloud format that represents points according to one or more coordinate systems, each point being assigned one or more data attributes specifying visual characteristics, the visual characteristics including one or more of size, normal vector, motion information, color information, and transparency.
22. A method for encoding a sequence of frames representing a dynamic 3D scene, wherein each frame may be composed of base layer image data and enhancement data, the method include: Layered encoding is performed on frames in the sequence of frames to generate coded frames, the coded frames comprising one or more of a base picture layer and an enhancement picture layer.
23. A method according to claim 22, wherein the enhancement picture layer comprises data used by the decoder device to reconstruct a higher resolution rendition of the frame sequence.
24. A method according to claim 22 or claim 23, wherein the enhancement picture layer comprises data used by the decoder device to reconstruct a higher bit depth rendition of the sequence of frames relative to the bit depth of the base picture layer.
25. A method according to any one of claims 22 to 24, wherein the enhancement picture layer comprises data used by the decoder device to reconstruct the distance of objects in the picture from a viewer.
26. A method according to any one of claims 22 to 25, the enhancement image layer comprising data used by the decoder device to reconstruct tactile feedback for a user.
27. The method according to any one of claims 22 to 26, further comprising: include: In response to a decrease in the transmission channel bandwidth, enhancement data is discarded during transmission of the frame sequence and the encoder is indicated that the enhancement data has been discarded, and if the enhancement data has been discarded, a time buffer for encoding the enhancement data is flushed and an instantaneous decoder refresh (IDR) is performed on the enhancement data (not necessarily on the base layer data) to account for the decoder losing some previous enhancement data.
28. The method according to any one of claims 22 to 27, in, The layered coding method used is MPEG-5 LCEVC (Low Complexity Enhancement Video Coding) coding or SMPTE VC-6 coding.
29. The method of claim 28, wherein at least one of the enhancement data is transmitted as embedded user data within LCEVC data coefficients.
30. A method according to claim 29, wherein one or more residual coefficients of the image include embedded depth information, the embedded depth information represents the depth of the corresponding object relative to the viewpoint, and the decoder processes the embedded data to reconstruct a depth map associated with the image frame based at least in part on the embedded data.
31. The method of claim 30, wherein the depth map is reconstructed by processing the embedded data and the image data.
32. A method according to any one of claims 22 to 31, wherein a display device uses frames plus depth data sent at a given frame rate to increase the frame rate by depth-based reprojection to match the display frame rate.
33. A method according to any one of claims 22 to 32, wherein each frame comprises image data and depth map data, and the method include: Layered encoding is performed on frames in the sequence of frames to generate encoded frames including one or more of a base depth layer and an enhancement depth layer.
34. The method of claim 33, wherein each frame comprises an image layer and a depth layer, and the encoded frame comprises one or more of a base image layer, the base depth layer, an enhancement image layer, and the enhancement depth layer.
35. A method according to claim 33, wherein each frame includes the depth map data embedded in the image data, the base depth layer is a base image layer having embedded depth map data, and the enhancement depth layer is an enhancement image layer having embedded depth map data.
36. The method according to any one of claims 33 to 35, further comprising: include: receiving a depth map discard indication, the depth map discard indication indicating whether to discard the depth map data when transmitting the frame sequence; as well as If the depth map data is discarded, the depth map data of the frame is discarded, and the image data of the frame is layered encoded to generate an encoded frame including a base image layer and an enhanced image layer.
37. A method, the method comprising: include: receiving frames plus depth data at a given frame rate, the frames plus depth data being encoded by a layered encoding method; At a display device, the frame rate is increased through depth-based reprojection using the depth data to match a display frame rate.
38. The method of claim 37, wherein the frames include data representing a dynamic 3D scene and each frame includes base layer image data and enhancement data, the method include: Layered decoding is performed on frames in the sequence of frames to generate decoded frames from encoded frames including a base picture layer and an enhancement picture layer.
39. A bit sequence representing an encoding of a frame sequence representing a dynamic 3D scene, the bit sequence comprising one or more of the following: the encoded data of the frame's base depth layer; and The encoded data of the enhanced depth layer of the frame.
40. The bit sequence according to claim 39, further comprising one or more of the following: coded data of a base picture layer of the frame; and The encoded data of the enhancement picture layer of the frame.
41. The bit sequence of claim 39, wherein the base depth layer is a base image layer having embedded depth map data, and the enhancement depth layer is an enhancement image layer having embedded depth map data.
42. A method for transmitting a sequence of frames representing a dynamic 3D scene, wherein each frame comprises image data and depth map data, the method include: Obtaining a coded frame including one or more of a base depth layer and an enhanced depth layer; determining whether to discard the depth map data when transmitting the sequence of frames; If you want to discard the depth map data: discarding said depth map data from said coded frame to give a downscaled coded frame comprising one or more of a base picture layer and an enhancement picture layer; as well as transmitting the reduced coded frame; as well as Otherwise, the encoded frame is transmitted.
43. The method according to claim 42, in: Each frame consists of an image layer and a depth layer. The coded frame includes one or more of a base image layer, the base depth layer, an enhancement image layer, and the enhancement depth layer, and Discarding the depth map data includes discarding the enhanced depth layer, and preferably, further includes discarding the base depth layer.
44. The method according to claim 42, in: Each frame comprises said depth map data embedded in said image data, The base depth layer is a base image layer having embedded depth map data, and the enhancement depth layer is an enhancement image layer having embedded depth map data, and Discarding the depth map data comprises removing the embedded depth map data from the enhancement image layer, and preferably further comprises removing the embedded depth map data from the base image layer.
45. The method according to any one of claims 42 to 44, further comprising: include: generating a depth map discard indication, the depth map discard indication indicating whether to discard the depth map data when transmitting the frame sequence; as well as The depth map discarding indication is sent to an upstream encoder, or the depth map discarding indication is sent together with the coded frame or the reduced coded frame.
46. A method for encoding a frame sequence performed by an encoder, the method include: performing layered coding on a first frame in the sequence of frames to generate a first coded frame including a base layer and an enhancement layer; storing the enhancement layer components of the first coded frame in a temporal buffer for use in temporal coding of subsequent frames; sending the first coded frame to a transmitter for transmission; receiving an enhancement discard indication, the enhancement discard indication indicating whether to discard the enhancement layer when transmitting the first coded frame; performing layered encoding on a second frame in the sequence of frames to generate a second coded frame comprising a base layer and an enhancement layer, wherein: generating the enhancement layer of the second coded frame with reference to the temporal buffer if the enhancement discard indication indicates that the enhancement layer is not discarded, and If the enhancement discard indication indicates that the enhancement layer is discarded, generating the enhancement layer of the second coded frame without referring to the temporal buffer.
47. The method of claim 46, further comprising: include: If the enhancement discard indication indicates that the enhancement layer is discarded, the time buffer is flushed.
48. The method of claim 46 or claim 47, further comprising: include: If no enhancement discard indication is received within a predetermined time limit, a second frame in the frame sequence is layered encoded to generate a second encoded frame including a base layer and an enhancement layer, wherein the enhancement layer of the second encoded frame is generated with reference to the time buffer.
49. The method of claim 46 or claim 47, further comprising: include: If no enhancement discard indication is received within a predetermined time limit, a second frame in the frame sequence is layered encoded to generate a second coded frame including a base layer and an enhancement layer, wherein the enhancement layer of the second coded frame is generated without reference to the time buffer.
50. A method according to any one of claims 46 to 49, wherein an enhancement layer of a frame is generated without reference to the temporal buffer include: decoding the base layer of the coded frame; as well as The residual is calculated as the difference between the decoded base layer and the frame.
51. A method according to any one of claims 46 to 50, wherein an enhancement layer of a frame is generated with reference to the temporal buffer include: decoding the base layer of the coded frame; calculating a residual as a difference between a decoded base layer and the frame; as well as The difference between the residual of the current frame and the corresponding residual of the previous frame stored in the temporal buffer is calculated.
52. The method according to any one of claims 46 to 51, wherein the layered coding is LCEVC coding.
Citation Information
Patent Citations
Point cloud data frames compression
EP4156108A1
Data processing apparatuses, methods, computer programs and computer-readable media
GB201615265D0
Video compression using differences between a higher and a lower layer
WO2018046940A1
Methods and apparatuses for encoding and decoding a bytestream
WO2019111010A1
Low complexity enhancement video coding
WO2020188273A1