Decoupled rendering for increasing tolerance to latency variations in extended reality applications with remote rendering
By separating the rendering process, the visual scenes are divided into multiple graphics layers and grouped, solving the problem of uncontrollable delay caused by delay changes in remote rendering XR applications, achieving a more stable user experience.
Patent Information
- Application Number
- CN202080103730.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-04
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-09-04
AI Technical Summary
In extended reality (XR) applications of remote rendering, delay changes lead to uncontrollable delays, affecting the user experience.
Through the separate rendering process, the visual scene is divided into multiple graphics layers, and grouped and sorted according to the quality metrics, encoded into composite video frames, and transmitted to the client device. The client device decodes the received graphics layer group in time and renders it using an alternative graphics layer group before the decoding deadline.
It reduces the impact of latency changes on XR applications, and improves the system's delay stability and user experience.
Smart Images

Figure CN116018806B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to extended reality (XR) applications, and more particularly, to techniques for separating the rendering process to mitigate the impact of latency variations in XR applications with remote rendering. Background Art
[0002] Mobile devices, head-mounted displays (HMDs), set-top boxes, and similar client devices lack the graphics capabilities and processing capabilities for extended reality (XR) applications such as video games and training simulations. XR applications include both virtual reality (VR) and augmented reality (AR) applications. The limitations on processing capabilities and graphics capabilities can be overcome by implementing remote rendering in which the heavy work of three-dimensional (3D) graphics rendering is performed by a central server and the rendered video is transmitted over a network to the client device. For example, the rendering server can be located in a data center in the edge cloud. This approach requires real-time transmission of the encoded (compressed) video on the downlink to the client device and real-time transmission of sensor data (and camera captures for AR) on the uplink to the rendering server.
[0003] Video games and XR applications typically have very strict latency requirements. The interaction latency (i.e., the time from when the user presses a button or moves a controller until the video has been updated on the display (e.g., an HMD)) should be approximately 100 ms. The motion-to-photon latency requirement (i.e., the time from when the user's head moves until the video has been updated on the display) is even shorter, approximately 20 ms. In the case of remote rendering, these latency requirements are practically unachievable, and thus are mitigated by techniques such as time warping.
[0004] Regarding the problem of compound latency, AR / VR applications can react to changing network conditions by changing its encoding quality and thus the rate of the video stream. Such adaptive methods are supported in current real-time video streaming solutions such as Web Real-Time Communication (WebRTC) or Self-Clocking Rate Adaptation for Multimedia (SCReAM). Unstable throughput is a known problem in current (fixed and mobile) networks, and it will be even more unstable in fifth-generation (5G) networks, where with the new radio (NR), larger bit pipes are possible, but the fading and interference effects of high-frequency channels will also be more pronounced.
[0005] When the network throughput is lower than the video coding rate, video packets will start queuing at the bottleneck link and will either be delayed (delivered late) to the client or be lost. The potential retransmission of these frames can result in additional latency. Due to the strict latency requirements of XR applications, the result of queuing latency is that the delayed frames cannot be decoded and presented in time, which is effectively equivalent to being lost.
[0006] To prevent long queuing delays, AR / VR applications can react to changing network conditions by changing the encoding quality and thus the bitrate of the video stream. With bitrate adaptation, the system will attempt to match the encoding rate of the video to the network throughput to prevent queuing delays. However, when adapting the bitrate, there is a transition period between the time congestion is detected and the time the bitrate is adjusted. During this transition period, the video rate can exceed the throughput. By using a low-latency, low-loss, scalable throughput (L4S) mechanism, such as enabling low transmission queue establishment, to react to congestion as quickly as possible, the length of the transition period can be reduced. However, a quick reaction to congestion can lead to underutilization of available resources and ultimately a degraded user experience.
[0007] Eliminating queuing delays by dropping frames at the transmission queue or even at the server is not a viable solution because all frames must be available for decoding at the client. If a frame is lost, a new I-frame much larger than a P-frame will be required, so creating I-frames should be avoided if possible. SUMMARY OF THE INVENTION
[0008] The present disclosure generally relates to a method of split rendering for making video applications more tolerant of queuing delay variations. The rendered image is divided into layers called graphics layers to reduce the video stream size. The server groups and sorts the graphics layers based on quality of experience (QoE) importance. When bitrate adaptation is applied, the server first reduces the video quality of the layers with lower importance. The client device decodes and presents the graphics layers that have been received in time, i.e., until decoding should start. The client device also saves the last received instance of each graphics layer and uses them for presentation if the next instance of the graphics layer is not received in time for decoding. The client device provides feedback to the server indicating which layers it has received in time. Based on the feedback, the server controls its sorting and adaptation mechanism.
[0009] A first aspect of the disclosure includes a method of split rendering implemented by a server node for reducing the impact of delay variations in a remote rendering application. The method includes rendering a visual scene to create a plurality of graphics layers. Each graphics layer includes one or more objects in the visual scene. The method further includes grouping and sorting the graphics layers according to a quality metric associated with the graphics layer to create a plurality of graphics layer groups having different quality levels. The method further includes encoding the graphics layer groups into composite video frames such that each graphics layer group in the composite video frame is separately decodable. The method further includes transmitting each graphics layer group of the composite video frame to a client device based on the quality level, in the sorted order.
[0010] A second aspect of the disclosure includes a method of split rendering implemented by a client device to reduce the impact of latency variations. The method includes receiving at least a portion of a composite video frame from a server node. The composite video frame includes a plurality of separately decodable groups of graphics layers representing a visual scene. Each group of graphics layers includes one or more graphics layers representing an object in the visual scene. The method further includes decoding the received groups of graphics layers in the composite video frame received before a decoding deadline. The method further includes: if any group of graphics layers has not arrived before the decoding deadline, deriving an alternative group of graphics layers for each of one or more late-arriving groups of graphics layers in the composite video frame not received before the decoding deadline. The method further includes rendering the composite video frame from the graphics layers in the received groups of graphics layers and the alternative groups to reconstruct the visual scene for display. The method further includes sending feedback indicating the received groups of graphics layers to the server node.
[0011] A third aspect of the disclosure includes a server node configured to reduce the impact of latency variations in a remote rendering application. The server node is configured to split render a visual scene to create a plurality of graphics layers. Each graphics layer includes one or more objects in the visual scene. The server node is further configured to group and sort the graphics layers according to a quality metric associated with the graphics layer to create a plurality of groups of graphics layers with different quality levels. The server node is further configured to encode the groups of graphics layers into a composite video frame such that each group of graphics layers in the composite video frame is separately decodable. The server node is further configured to transmit each group of graphics layers of the composite video frame to the client device based on the quality level and in the sorted order.
[0012] A fourth aspect of the disclosure includes a client device configured to reduce the impact of latency variations. The client device is configured to receive at least a portion of a composite video frame from a server node. The composite video frame includes a plurality of separately decodable groups of graphics layers representing a visual scene. Each group of graphics layers includes one or more graphics layers representing an object in the visual scene. The client device is further configured to decode the received groups of graphics layers in the composite video frame received before a decoding deadline. The client device is further configured to, if any group of graphics layers has not arrived before the decoding deadline, derive an alternative group of graphics layers for each of one or more late-arriving groups of graphics layers in the composite video frame not received before the decoding deadline. The client device is further configured to render the composite video frame from the graphics layers in the received groups of graphics layers and the alternative groups to reconstruct the visual scene for display. The client device is further configured to send feedback indicating the received groups of graphics layers to the server node. The client device optionally displays the visual scene on a display.
[0013] A fifth aspect of the disclosure includes a server node configured to reduce the impact of latency variations in a remote rendering application. The server node includes communication circuitry for communicating with a client device and processing circuitry. The processing circuitry is configured to separately render a visual scene to create a plurality of graphics layers. Each graphics layer includes one or more objects in the visual scene. The processing circuitry is further configured to group and sort the graphics layers according to quality metrics associated with the graphics layers to create a plurality of groups of graphics layers having different quality levels. The processing circuitry is further configured to encode the groups of graphics layers into composite video frames such that each group of graphics layers in the composite video frames is separately decodable. The processing circuitry is further configured to transmit each group of graphics layers of the composite video frames to the client device based on the quality level and in the sorted order.
[0014] A sixth aspect of the disclosure includes a client device configured to reduce the impact of latency variations. The client device includes communication circuitry for communicating with a server node and processing circuitry. The processing circuitry is configured to receive at least a portion of a composite video frame from the server node. The composite video frame includes a plurality of separately decodable groups of graphics layers representing a visual scene. Each group of graphics layers includes one or more graphics layers representing objects in the visual scene. The processing circuitry is further configured to decode the received groups of graphics layers in the composite video frame received before a decoding deadline. The processing circuitry is further configured to, if any group of graphics layers has not arrived before the decoding deadline, derive an alternative group of graphics layers for each of one or more late-arriving groups of graphics layers in the composite video frame not received before the decoding deadline. The processing circuitry is further configured to render the composite video frame from the graphics layers in the received groups of graphics layers and the alternative groups of layers to reconstruct the visual scene for display. The processing circuitry is further configured to send feedback indicating the received groups of graphics layers to the server node. The client device optionally displays the visual scene on a display.
[0015] A seventh aspect of the disclosure includes a computer program for a server node or other network node configured to perform separate rendering to mitigate the impact of latency variations in a remote rendering application. The computer program includes executable instructions that, when executed by the processing circuitry in the server node, cause the server node to perform the method according to the first aspect.
[0016] An eighth aspect of the disclosure includes a carrier containing the computer program according to the seventh aspect. The carrier is one of an electronic signal, an optical signal, a radio signal, or a non-transitory computer-readable storage medium.
[0017] A ninth aspect of the disclosure includes a computer program for a client device or configured to perform split rendering to mitigate the impact of latency variations in a remote rendering application. The computer program includes executable instructions which, when executed by a processing circuit in a server node, cause the server node to perform the method according to the first aspect.
[0018] A tenth aspect of the disclosure includes a carrier containing the computer program according to the ninth aspect. The carrier is one of an electronic signal, an optical signal, a radio signal, or a non-transitory computer-readable storage medium. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a functional block diagram illustrating a communication system configured according to an embodiment of the present disclosure.
[0020] Figure 2 is a flowchart illustrating a method of split rendering implemented at a server node in an edge data network (EDN) to mitigate the impact of latency variations in a remote rendering system.
[0021] Figure 3 illustrates objects in a visual scene and their position information and velocities.
[0022] Figure 4 illustrates the objects of the visual scene after sorting and grouping.
[0023] Figure 5 illustrates an exemplary composite video frame.
[0024] Figure 6 is a flowchart illustrating a method of split rendering implemented at a client device connected to a server node in an edge data network (EDN) to mitigate the impact of latency variations.
[0025] Figure 7 is a schematic block diagram illustrating a server node configured for split rendering to mitigate the impact of latency variations in a remote rendering system.
[0026] Figure 8 is a schematic block diagram illustrating a client device configured for split rendering to mitigate the impact of latency variations in a remote rendering system.
[0027] Figure 9 is a flowchart illustrating a method of split rendering implemented at a server node in an edge data network (EDN) to mitigate the impact of latency variations in a remote rendering system.
[0028] Figure 10A flowchart illustrating a method of split rendering implemented at a client device connected to a server node in an edge data network (EDN) to mitigate the effects of latency variations.
[0029] Figure 11 A schematic block diagram of the main functional components of a server node configured to perform split rendering as described herein to mitigate the effects of latency variations.
[0030] Figure 12 A schematic block diagram of the main functional components of a client device configured to perform split rendering as described herein to mitigate the effects of latency variations. DETAILED DESCRIPTION
[0031] The improved split rendering process of the present disclosure mitigates the effects of latency variations in XR applications with remote rendering applications. Graphics applications such as game engines are executed on server nodes deployed in an edge data network (EDN). The game engine generates a visual scene for display to remote users. The visual scene is split rendered to generate graphic layers from 3D objects in the visual scene. The server node groups and sorts the graphic layers based on QoE importance to create graphic layer groups, encodes each graphic layer group into a composite video frame and attaches metadata to the composite video frame identifying the graphic layer group. The encoded video frames are then transmitted to a client device 200 (e.g., an HMD worn by a user) where the video frames are decoded and displayed, based on a quality level and in sorted order. When the deadline for decoding the video arrives, the client device reads the metadata, identifies the graphic layer groups and decodes each graphic layer group received before the decoding deadline. If a graphic layer group is not received in time before the decoding deadline, the client device 200 uses cached data from a previous frame to derive an alternative graphic layer group for the late or missing graphic layer group. The resulting graphic layers are then rendered to reconstruct the visual scene, and the resulting graphic layers are displayed to the user on the client device 200. The client device 200 further sends feedback to the server indicating the graphic layer groups that were received in time.
[0032] Based on feedback from the client, the split rendering techniques as described herein enable the server node to adapt the graphic layer groups separately depending on QoE importance and in response to changes in network conditions to ensure the best possible QoE for a given transmission condition. Additionally, the need for rapid adaptation to changing network conditions is eliminated by trading a slight degradation in video quality for reduced latency and latency variability.
[0033] The techniques as described herein can be combined with other techniques such as time warping techniques to mitigate the effects of missing information in the rendered image.
[0034] The solution is network agnostic and does not require any specific cooperation with the network in the form of an offline service level agreement (SLA) or in the form of interactions based on online network application programming interfaces (APIs). However, the proposed method can still benefit from some of the prior art that keeps the transmission queue short, such as explicit congestion notification (ECN) or L4S.
[0035] Figure 1 Illustrates a communication network 10 configured according to an embodiment of the present disclosure. As Figure 1 seen, network 10 includes an access network 12 that communicatively connects a client device 200 (such as an HMD) to a server node 100 disposed in a cloud network 14. In some embodiments, a computing device 20 is disposed between the client device 200 and the server node 100 and is configured to perform at least some of the processing functions of the client device 200.
[0036] The access network 12 can be any type of communication network (such as Wi-Fi, Ethernet, wireless local area network (WLAN), wideband code division multiple access (WCDMA), long term evolution (LTE), etc.) and the function of connecting a subscriber device such as the client device 200 to one or more service provider nodes such as the server node 100. The cloud network 14 provides "on-demand" availability of computer resources (such as memory, data storage devices, processing capabilities, etc.) to such subscriber devices without requiring the user to directly and actively manage those resources. According to embodiments of the present disclosure, such resources include, but are not limited to, one or more XR applications being executed on the server node 100. The XR applications can include, for example, game applications and / or simulation applications for training.
[0037] Figure 2 is a flowchart illustrating a method 300 for mitigating the effects of latency variations in video images due to transmission latency in a network. In this embodiment, method 300 is implemented by a server node 100 that executes an XR application. Method 300 is executed on a frame-by-frame basis.
[0038] Generally, the server node 100 groups and sorts the graphic layers in a video frame based on QoE importance and transmits different groups of graphic layers to the client device 200 in the sorted order based on the QoE level. More specifically, when the server node 100 receives a visual scene from the XR application, the server node 100 renders the visual scene to generate a plurality of graphic layers, each of the plurality of graphic layers including one or more 3D objects in the visual scene (block 310). Each graphic layer is associated with a motion vector and spatial position information.
[0039] The server node 100 optionally receives feedback from the client device 200 indicating the graphics layer groups received in the previous frame, which, as described below, can be used for bitrate adaptation of the graphics layer groups (block 320).
[0040] The server node 100 groups and sorts the graphics layers based on QoE importance (block 330). That is, graphics layers that are considered more important for the user experience are given a higher rank compared to those that are considered less important for the user experience. The server node 100 can consider many factors when determining the QoE rank of a graphics layer. The following is a non-exhaustive list of factors that can be considered when grouping and sorting the graphics layers:
[0041] ● Foreground objects and graphics layers for player objects generally have higher importance compared to the graphics of the background.
[0042] ● Some objects (such as objects representing the avatars of game opponents) are more important to be tracked precisely.
[0043] ● Newly emerging objects (such as bullet shots in a first-person shooter (FPS) game) are more important to be transferred because they cannot be inferred from the previous video frames at the decoder.
[0044] ● Moving objects are more important than static objects because the temporal warping from moving objects can cause viewing artifacts such as position jitter.
[0045] ● Client feedback on which layers can be presented in the last frame can be an input for re-prioritizing some layers.
[0046] The result of the grouping and sorting process is a set of graphics layer groups. Some graphics layer groups can include a single graphics layer, while other groups include two or more graphics layers. In the most extreme case, each group can include a single graphics layer.
[0047] Figure 2 Illustrates a group of objects rendered as graphics layers, Figure 2 which is used as an example to illustrate the grouping and sorting. The objects are labeled {A, B, C, D, E}. The position and speed of each object are given. Figure 3Describes the objects after grouping and sorting. The QoE metrics in this example are the speed values and Z-order assigned to the graphic layers. The graphic layers are initially grouped into three speed ranges V0 = 0, V1 = [0, .5], and V2 = [.5, 1]. The speed values are scaled relative to the position of the observer. The groups are sorted in decreasing speed order, resulting in {A}, {E}, {C, B, D}. Within each speed group, the graphic layers are grouped into three separate Z-layers, where the Z-ranges are Z0 = [0, .33], Z1 = [.33, .66], and Z2 = [.66, 1], respectively. Then, the top-level groups with more than one graphic layer are sorted by increasing Z-order, resulting in Figure 4 the order in. Alternatively, objects B, C, and D can be included in separate graphic layer groups, resulting in {A}, {E}, {C}, {B}, {D}.
[0048] Return to Figure 2 , the server node 100 determines whether congestion is detected (block 340). This step is shown after grouping and sorting, but can be performed earlier. The detection of congestion can be based on client feedback, network feedback, its own detection mechanism, or some combination of these methods. As described above, client feedback is provided to the server node 100 to indicate what graphic layers or graphic layer groups were actually received in the previous frame. The reception of fewer than all graphic layers in a graphic layer can be considered by the server node 100 as an indication of network congestion. Conversely, the reception of all graphic layers can be regarded by the server node 100 as an indication of the absence of network congestion. Network feedback can be in the form of lost or ECN-marked packets, or other information (such as available bandwidth) communicated via an explicit cooperation protocol. The own detection mechanism can include explicit or implicit feedback from the transport layer protocol.
[0049] The encoder at the server node 100 adapts to the network conditions. Depending on whether network congestion exists, the server node 100 determines the coding rate (also referred to herein as the target rate) for each graphic layer or graphic layer group. If congestion is detected, the target rate for one or more graphic layers or graphic layer groups is reduced (block 350). If no congestion is detected, the target rate for one or more graphic layers or graphic layer groups can be increased (block 360).
[0050] In one embodiment, the coding rate, also referred to as the target rate, is determined separately for each graphic layer group. The coding logic takes into account the QoE importance of the graphic layer or graphic layer group, i.e., more important graphic layers are encoded with higher quality. This QoE-aware coding ensures the best possible QoE at the client device 200 for a given network condition.
[0051] The sorted graphic layers, along with the determined target rates at boxes 350 and / or 360, are input into the video encoder. The encoder in server node 100 encodes each graphic layer or group of graphic layers (box 370) separately based on the determined target rates. Independent encoding eliminates dependencies between groups of graphic layers and ensures that each group of graphic layers is separately decodable independent of other groups of graphic layers.
[0052] As described above, a group of graphic layers can include one or more graphic layers. Return reference Figure 3 , the encoder can encode layer A and layer B separately in one example and then encode layers C, D, and E as a group of graphic layers. In this example, the final encoding includes three groups of graphic layers. In another example, the encoder can encode all graphic layers separately, i.e., as five separate 1-layer groups. The groups of graphic layers for encoding are assembled into a composite video frame in the order of QoE levels. For multi-layer groups, the group of graphic layers can be assigned the QoE level of the highest-level graphic layer in the group. During encoding, server node 100 appends metadata at the start of the composite frame, the metadata including the number of video streams or groups of graphic layers being transmitted.
[0053] In some embodiments, the encoder can use the feedback from client device 200 as a reference for the encoding process. That is, client device 200 can use the last received graphic layer or group of graphic layers to encode the current graphic layer of the group of graphic layers. If a graphic layer or group of graphic layers has not been received in a previous frame, the encoder can use the last received group of graphic layers as a reference for the current frame. For example, assume the encoder is encoding group C of the current frame and the client feedback indicates that group C has not been received in frame n - 1. In this case, the encoder can use group C from frame n - 2 as a reference to encode group C in the current frame.
[0054] After encoding, the composite video frame is transmitted to client device 200 (box 380). The metadata is transmitted first, followed by each group in order based on the group level.
[0055] Figure 5 Illustrates based on Figure 2 and Figure 3 An example of a composite video frame for the example shown in. The composite frame includes five groups of graphic layers labeled A - E respectively. The groups of graphic layers are organized in sorted order in the composite video frame and metadata is appended at the start of the composite video frame. In this example, the metadata is transmitted first, followed by groups A, E, C, B, and D in the order of group A, E, C, B, and D.
[0056] In some embodiments, the server node 100 may also encode the motion information of each graphics layer or group of graphics layers. The motion information can be used, for example, to perform time warping at the client device 200.
[0057] Figure 6 Illustrates a complementary method 400 implemented by the client device 200 (e.g., an HMD). At block 405, before the deadline for starting decoding arrives, the client device 200 receives at least a portion of the current video frame represented as frame n. The decoding deadline is reached at time Ti (block 410). When the decoding deadline is reached, the client device 200 decodes the portion of the current frame received before the decoding deadline and checks which portions of the corresponding frame have arrived (blocks 415, 420). For this determination, metadata placed at the start of the composite video frame is used. For example, if the decoding deadline Ti is the time shown by the vertical line in Figure 5 then the groups of graphics layers available for decoding include groups A and E.
[0058] The groups of graphics layers that have arrived in time are stored in a buffer and sent directly to the display (blocks 425, 430). For groups of graphics layers that have not arrived in time, the client device 200 checks whether it has a cached instance of the same group of graphics layers from the last received frame ( Figure 5 the frame at time Ti-1 in). If so, the client device 200 derives this group of graphics layers at time Ti from the cached instance of the group at time Ti-1 (block 440). Techniques such as time warping can be applied to obtain a better approximation of the graphics layers in this group of graphics layers. The derived group of graphics layers is then stored in the buffer and the derived group of graphics layers is sent downstream for display (blocks 445, 450).
[0059] When all groups of graphics layers have arrived or have been derived, the client device 200 renders the display image and outputs the image to the display (blocks 455, 460). The image is rendered from the graphics layers using 3D position information and layer textures.
[0060] Finally, the client device 200 sends feedback to the server node 100 indicating which groups of graphics layers were received before the decoding deadline (block 465). As described above, this feedback can be used to prioritize the graphics layers, to determine the coding rate of the groups of graphics layers, or to select a reference for encoding the graphics layers in the current frame.
[0061] In some embodiments, after reaching the decoding deadline, the client device 200 may continue to receive the missing portions of the current frame. If an alternative graphics layer group is stored in the buffer, when the missing portions are finally received, the alternative graphics layer group may be overwritten or replaced. In some embodiments, the client may wait until a predetermined time (Ti + t) after Ti to provide feedback to the server node 100. In this case, the feedback will indicate the graphics layer group received before Ti + t. In another embodiment, before sending feedback, the client device 200 waits until the first few bytes of the next frame arrive.
[0062] Figure 7 Illustrated is an exemplary server node 100 for an XR application configured to perform split rendering to reduce the impact of latency variations in a remote rendering application. The server node 100 includes a receiver 110, a processing unit 120, and a transmitter 190. The receiver 110 receives user input, sensor data, a camera feed, and client feedback from the client device 200. The processing unit 120 runs an application such as a video game or simulation and outputs a video stream to be displayed to the user of the client device 200. The transmitter 190 transmits the video stream to the client device 200.
[0063] The processing unit 120 in this example includes an optional game engine 130 or other application, a rendering engine 140, a packetizing and sorting unit 150, an encoding unit 160, a transmitting unit 170, and an optional feedback unit 180. The various units 120 - 180 may be implemented by hardware and / or by software code executed by a processor or processing circuitry. The game engine 130 receives user input, sensor data, and a camera feed, maintains the game context, and generates a visual scene to be presented on the display of the client device 200. The visual scene generated by the game engine 130 is input to the rendering unit 140. The rendering unit 140 renders the visual scene provided by the game engine 130 into a set of graphics layers. Each graphics layer includes one or more objects in the visual scene. The graphics layer includes position information for reconstructing the visual scene from the graphics layer and motion information that describes or characterizes the motion of the objects in the graphics layer. Typically, the objects in the same graphics layer will be characterized by the same motion.
[0064] The rendering unit 140 passes the graphics layers to the grouping and sorting unit 150. As previously described, the grouping and sorting unit 150 groups and sorts the graphics layers according to QoE importance. During the grouping and sorting process, some of the graphics layers in the graphics layers can be merged into a single graphics layer group representing a set of graphics layers with similar importance. The encoding unit 160 encodes each graphics layer group into a composite video frame respectively and attaches metadata at the start of the composite video frame to identify the graphics layer group. In the composite video frame, the groups are sorted according to importance, which can be derived from the importance of the graphics layers in the graphics layer group. The transmission unit 170 transmits the composite video frame to the client device 200 via the transmitter 290 over the network according to the communication protocol of the network. The processing unit 120 optionally includes a feedback unit 180 for processing client feedback from the client device 200. As previously described, the client feedback can be provided to the grouping and sorting unit 150 for prioritizing the graphics layers, or to the encoding unit 160 for encoding the graphics layers.
[0065] Figure 8 An exemplary client device 200 configured to perform split rendering to reduce the impact of latency variations in a remote rendering application is illustrated. The client device 200 includes a receiver 210, a processing unit 220, and a transmitter 290. The receiver 210 receives a video representing the rendering of a visual scene from the server node 100. The processing unit 220 decodes the received rendered video from the server node 100, reconstructs the visual scene based on the rendered video, and outputs the reconstructed visual scene for display to the user. The transmitter 290 sends user input, sensor data, camera feed, and client feedback received via a user interface (not shown) to the server node 100.
[0066] The processing unit 220 includes a decoding unit 230, an exporting unit 240, a buffer unit 250, a feedback unit 260, a rendering unit 270, and an optional display 280. The various units 220 - 280 can be implemented by hardware and / or by software code executed by a processor or processing circuitry. The decoding unit 230 decodes the rendered video received from the server node 100 to obtain graphic layers representing objects in the visual scene. If any of the graphic layers in the current frame are missing (either because they arrived too late or because they were lost), the exporting unit 240 exports an alternative graphic layer group for the missing graphic layer group from the information stored by the buffer unit 250. The buffer unit 250 stores the successfully decoded graphic layer groups and the alternative graphic layer groups generated by the exporting unit 240. The feedback unit 260 sends feedback to the server node 100 indicating that the graphic layer groups have been received in time for decoding. The rendering unit 270 renders the composite video frames received from the server node 100 to reconstruct the visual scene output by the game engine 130 or other applications running on the server node 100. The visual scene is output to the display 280 for presentation to the user.
[0067] Figure 9 Another exemplary method 500 of split rendering implemented by the server node 100 to reduce the impact of latency variations in a remote rendering application is illustrated. The server node 100 splits the rendering of the visual scene to create multiple graphic layers (block 510). Each graphic layer includes one or more objects in the visual scene. The server node 100 groups and sorts the graphic layers according to a quality metric associated with the graphic layers to create multiple graphic layer groups with different quality levels (block 520). The server node 100 encodes the graphic layer groups into composite video frames such that each graphic layer group in the composite video frame is separately decodable (block 530). The server node 100 transmits each graphic layer group of the composite video frame to the client device 200 in sorted order based on the quality level (block 540).
[0068] In some embodiments of method 500, the server node further adds metadata to the composite video frame identifying each graphic layer and its position in the video frame.
[0069] In some embodiments of method 500, grouping and sorting the graphic layers includes grouping M graphic layers into N graphic layer groups based on the quality metric of the graphic layers, where M > N.
[0070] Some embodiments of method 500 further include receiving feedback from the client device 200 indicating the graphic layers received in a previous frame and determining the quality metric of the graphic layers based on the feedback from the client device 200.
[0071] Some embodiments of method 500 further include receiving feedback from a client device indicating a graphics layer received in a previous frame and determining encoding of the graphics layer based on the feedback from client device 200.
[0072] In some embodiments of method 500, determining encoding of the graphics layer based on feedback from the client device includes reducing an encoding rate of a group of graphics layers when the feedback indicates that one of the graphics layers in the group was not received in the previous frame.
[0073] Some embodiments of method 500 further include detecting congestion in a communication link between a rendering device (e.g., server node 100) and client device 200 and independently changing an encoding rate of each group based on the detected congestion.
[0074] Some embodiments of method 500 further include reducing an encoding rate of at least one group of graphics layers when congestion is detected and increasing an encoding rate of at least one group when congestion is not detected.
[0075] Figure 10 An exemplary method 600 of split rendering implemented by client device 200 to reduce the impact of latency variations is illustrated. Client device 200 receives at least a portion of a composite video frame from server node 100 (block 610). The composite video frame includes a plurality of separately decodable groups of graphics layers representing a visual scene. Each group of graphics layers includes one or more graphics layers representing an object in the visual scene. Client device 200 decodes the received groups of graphics layers in the composite video frame received before a decoding deadline (block 620). If any group of graphics layers has not arrived before the decoding deadline, client device 200 derives an alternative group of graphics layers for each of one or more late-arriving groups of graphics layers in the composite video frame not received before the decoding deadline (block 630). Client device 200 renders the composite video frame from the graphics layers in the received groups of graphics layers and the alternative groups of layers to reconstruct the visual scene for display (block 640). Client device 200 sends feedback to server node 100 indicating the groups of graphics layers received in time (block 650). Client device 200 optionally displays the visual scene on a display (block 660).
[0076] In some embodiments of method 600, deriving alternative graphics layers includes: for each late-arriving group of graphics layers, retrieving a corresponding previous group of graphics layers from a buffer and deriving an alternative group of graphics layers based on the previous group of graphics layers.
[0077] Some embodiments of method 600 further include storing the groups of graphics layers received in time and the alternative groups of graphics layers in a buffer.
[0078] Some embodiments of method 600 further include: after a decoding deadline of a current frame, receiving a late-arriving graphics layer group, decoding the late-arriving graphics layer group, and storing the late-arriving graphics layer group in a buffer by replacing a corresponding replacement graphics layer group with the late-arriving graphics layer group.
[0079] Figure 11 FIG. is a schematic block diagram illustrating some exemplary components of server node 700 configured to perform split rendering as described herein to reduce the impact of latency variations. Server node 700 includes communication circuitry 710, processing circuitry 720, and memory 730. As described in more detail later, a computer program 740 that configures server node 700 to operate in accordance with this embodiment can be stored in memory 730.
[0080] Communication circuitry 810 includes network interface circuitry for communicating with other nodes in and / or communicatively coupled to communication network 10. Such nodes include, but are not limited to, one or more network nodes and / or functions arranged in cloud network 14, access network 12, and / or one or more client devices 200 (such as an HMD).
[0081] Processing circuitry 720 controls the overall operation of server node 16 and is configured to perform the steps of method 300 and method 500 shown respectively in Figure 2 and Figure 9 Such processing can include, for example, split rendering a visual scene to create graphics layers, grouping and sorting the graphics layers into graphics layer groups, encoding the graphics layer groups, and transmitting the graphics layer groups in sorted order based on QoE importance. Processing circuitry 720 can include one or more microprocessors, hardware, firmware, or a combination thereof.
[0082] Memory 730 includes both volatile and non-volatile memory for storing computer program code and data needed for operation by processing circuitry 720. Memory 730 can include any tangible, non-transitory computer-readable storage medium for storing data, including electronic, magnetic, optical, electromagnetic, or semiconductor data storage devices. Memory 730 stores a computer program 740 including executable instructions that configure processing circuitry 720 to implement respectively in Figure 2 and Figure 9Methods 300 and 500 shown therein. As described more fully below, computer program 740 may include one or more code modules in this regard. Generally, computer program instructions and configuration information are stored in non-volatile memories such as ROM, erasable programmable read-only memory (EPROM), or flash memory. Temporary data generated during operation may be stored in volatile memories such as random access memory (RAM). In some embodiments, computer program 740 for configuring processing circuit 720 as described herein may be stored in removable memories such as portable optical discs, portable digital video discs, or other removable media. Computer program 740 may also be included in carriers such as electrical signals, optical signals, radio signals, or computer-readable storage media.
[0083] Figure 12 is a schematic block diagram illustrating some exemplary components of client device 200 configured to perform split rendering as described herein to reduce the impact of latency variations. In the illustrated embodiment, client device 200 includes a head-mounted display 850 configured to display computer-generated images (CGIs), live images captured from a physical environment, or a combination of both. Client device 800 in this embodiment includes communication circuit 810, processing circuit 820, memory circuit 830, and user interface 850. As described in more detail later, computer program 840 for configuring client device 200 to operate according to this embodiment may be stored in memory 830.
[0084] Communication circuit 810 includes network interface circuitry for communicating with other nodes in and / or communicatively connected to communication network 10. Such nodes include, but are not limited to, one or more network nodes and / or functions arranged in cloud network 14, access network 12, and / or one or more server nodes 100.
[0085] Processing circuit 820 controls the overall operation of client device 800 and is configured to perform the steps of methods 400 and 600 shown respectively in Figure 6 and Figure 10 Such processing may include, for example, receiving a composite video stream including a plurality of separately decodable groups of graphic layers (each group of graphic layers including one or more graphic layers representing objects in a visual scene), decoding the received group of graphic layers in the composite video frame, deriving alternative groups of graphic layers for late or missing groups of graphic layers, rendering the composite video frame from the received group of graphic layers and the alternative layer groups to reconstruct the visual scene, and sending feedback to the server node indicating that the group of graphic layers received in time for decoding. Processing circuit 820 may include one or more microprocessors, hardware, firmware, or a combination thereof.
[0086] Memory 830 includes both volatile and non-volatile memories for storing computer program code and data required for operation by processing circuitry 820. Memory 830 may include any tangible, non-transitory computer-readable storage medium for storing data, including electronic, magnetic, optical, electromagnetic, or semiconductor data storage devices. Memory 830 stores computer program 840 including executable instructions that configure processing circuitry 820 to implement method 400 and method 600 shown in Figure 5 and Figure 10 respectively. As described more fully below, computer program 840 may include one or more code modules in this regard. Generally, computer program instructions and configuration information are stored in non-volatile memories such as ROM, erasable programmable read-only memory (EPROM), or flash memory. Temporary data generated during operation may be stored in volatile memory such as random access memory (RAM). In some embodiments, computer program 840 for configuring processing circuitry 820 as described herein may be stored in removable memory such as a portable compact disc, portable digital video disc, or other removable medium. Computer program 840 may also be embodied in a carrier such as an electronic signal, optical signal, radio signal, or computer-readable storage medium.
[0087] User interface 850 includes a head-mounted display and / or user controls such as buttons, actuators, and software-driven controls that facilitate the ability of a user to interact with and control the operation of applications running on server node 100. The head-mounted display is worn by the user and is configured to display rendered video images to the user. The head-mounted display may include sensors for sensing head movement and orientation and a camera for providing a video feed. Other types of displays may be used in place of the head-mounted display. Exemplary client devices include cathode ray tube (CRT), liquid crystal display (LCD), liquid crystal on silicon (LCos) display, and light emitting diode (LED) display. Other types of client devices not explicitly described herein may also be possible.
[0088] Those skilled in the art will also appreciate that the embodiments herein further include corresponding computer programs. The computer program includes instructions that, when executed on at least one processor of a device, cause the device to perform any of the corresponding processes described above. The computer program may include one or more code modules corresponding to the components or units described above in this regard.
[0089] The embodiments further include a carrier containing such a computer program. Such a carrier may include one of an electronic signal, optical signal, radio signal, or computer-readable storage medium.
[0090] In this regard, embodiments herein also include a computer program product stored on a non-transitory computer-readable (storage or recording) medium and including instructions that, when executed by a processor of a device, cause the device to perform as described above.
[0091] The embodiments further include a computer program product, the computer program product including a program code portion for performing the steps of any of the embodiments herein when the computer program product is executed by a computing device. Such a computer program product may be stored on a computer-readable recording medium.
Claims
1. A method (500) of remote rendering implemented by a server node (100, 700), the method comprises: Separating rendering (510) a visual scene to create a plurality of graphic layers, each graphic layer including one or more objects in the visual scene; Grouping and sorting (520) the graphic layers according to a quality metric associated with the graphic layers to create a plurality of graphic layer groups having different quality levels; Encoding (530) the plurality of graphic layer groups into a composite video frame, wherein each graphic layer group in the composite video frame is separately decodable; and Transmitting (540) each graphic layer group of the composite video frame to a client device (200, 800) based on the quality level and in the sorted order.
2. The method (500) according to claim 1, further comprising adding metadata to the composite video frame to identify each graphic layer and its position in the video frame.
3. The method (500) according to claim 1, wherein, Grouping and sorting the graphic layers includes grouping M graphic layers into N graphic layer groups based on the quality metric of the graphic layers, where M > N.
4. The method (500) according to claim 1, further comprises: Receiving feedback from the client device (200, 800) indicating the graphic layers received in a previous frame; Determining the quality metric of the graphic layers based on the feedback from the client device (200, 800).
5. The method (500) according to claim 1, further comprises: Receiving feedback from the client device (200, 800) indicating the graphic layers received in a previous frame; Determining the encoding of the graphic layers based on the feedback from the client device (200, 800).
6. The method (300) according to claim 5, wherein, Determining the encoding of the graphic layers based on the feedback from the client device (200, 800) includes: when the feedback indicates that one of the graphic layers in the graphic layer group was not received in the previous frame, reducing the encoding rate of the graphic layer group.
7. The method (500) according to any one of claims 1 to 5, further comprises: Detecting congestion in a communication link between the rendering device and the client device (200, 800); Independently changing the encoding rate of each graphic layer group based on the detected congestion.
8. The method (500) according to claim 7, wherein, Independently changing the encoding rate based on the detected congestion includes: When congestion is detected, reducing the encoding rate of at least one graphic layer group; and When no congestion is detected, increasing the encoding rate of at least one graphic layer group.
9. A method (600) implemented by a client device (200, 800) for displaying video, the method comprises: Receiving (610) at least a part of a composite video frame from a server node (100, 700), the composite video frame including a plurality of separately decodable graphic layer groups, each graphic layer group including one or more graphic layers representing objects in a visual scene; Decode (620) the received set of graphics layers in the composite video frame received before the decoding deadline; Derive (630) an alternative set of graphics layers for each of one or more late-arriving sets of graphics layers in the composite video frame not received before the decoding deadline; Render (640) the composite video frame from the graphics layers in the received set of graphics layers and the alternative set of graphics layers; And Send (650) feedback indicating the sets of graphics layers received in time to the server node (100, 700).
10. The method (600) according to claim 9, Wherein, Deriving alternative graphics layers includes, for each late-arriving set of graphics layers: Retrieving from a buffer a previous set of graphics layers corresponding to the late-arriving set of graphics layers; and Deriving the alternative set of graphics layers based on the previous set of graphics layers.
11. The method (600) according to claim 9 or 10, further comprising storing the sets of graphics layers received in time and the alternative sets of graphics layers in a buffer.
12. The method (600) according to claim 11, further Comprising: After the decoding deadline, after the decoding deadline of the current frame, receiving a late-arriving set of graphics layers; Decoding the late-arriving set of graphics layers; And Storing the late-arriving set of graphics layers in the buffer by replacing the corresponding alternative set of graphics layers with the late-arriving set of graphics layers.
13. A server node (100, 700) configured to perform remote rendering, the server node Comprising: A rendering unit configured to separately render a current video frame to create a plurality of graphics layers, each graphics layer including one or more objects; A grouping and sorting unit configured to group and sort the graphics layers according to a quality metric associated with the graphics layers to create a plurality of sets of graphics layers with different quality levels; An encoding unit configured to encode the plurality of sets of graphics layers into a composite video frame, wherein each set of graphics layers in the composite video frame is separately decodable; And A transmission unit configured to transmit each set of graphics layers of the composite video frame to a client device (200, 800) based on the quality level and in the sorted order.
14. The server node (100, 700) according to claim 13, further configured to perform the method according to any one of claims 2 to 8.
15. A server node (100, 700) configured to perform remote rendering, the server Comprising: A communication circuit (710) for communicating with a client device (200, 800); And A processing circuit (720) configured to: Separate and render a current video frame to create a plurality of graphics layers, each graphics layer including one or more objects; Group and sort the graphics layers according to a quality metric associated with the graphics layers to create a plurality of sets of graphics layers with different quality levels; Encode the plurality of sets of graphics layers into a composite video frame, wherein each set of graphics layers in the composite video frame is separately decodable; And Transmit each graphic layer group of the composite video frame to a client device (200, 800) based on a quality level and in a sorted order.
16. The server node (100, 700) according to claim 15, wherein, the processing circuit (720) is further configured to perform the method according to any one of claims 2 to 8.
17. A client device (200, 800) for displaying video, the client device (200, 800) comprising: a receiver configured to receive at least a portion of a composite video frame from a server node (100, 700), the composite video frame including a plurality of separately decodable graphic layer groups, each graphic layer group including one or more graphic layers; a decoding unit configured to decode the received graphic layer groups in the composite video frame received before a decoding deadline; a derivation unit configured to derive an alternative graphic layer group for each of one or more late-arriving graphic layer groups in the composite video frame not received before the decoding deadline; a rendering unit configured to render the composite video frame from the graphic layers in the received graphic layer groups and the alternative graphic layer groups; and a transmitter configured to send feedback to the server node (100, 700) indicating the graphic layer groups received in a timely manner.
18. The client device (200, 800) according to claim 17, further configured to perform the method according to any one of claims 10 to 12.
19. A client device (200, 800) for displaying video, the client device (200, 800) comprising: a communication circuit (810) for communicating with a server (100, 700); and a processing circuit (820) configured to: receive at least a portion of a composite video frame from a server node (100, 700), the composite video frame including a plurality of separately decodable graphic layer groups, each graphic layer group including one or more graphic layers; decode the received graphic layer groups in the composite video frame received before a decoding deadline; derive an alternative graphic layer group for each of one or more late-arriving graphic layer groups in the composite video frame not received before the decoding deadline; render and display the composite video frame from the graphic layers in the received graphic layer groups and the alternative graphic layer groups; and send feedback to the server node (100, 700) indicating the graphic layer groups received in a timely manner.
20. The client device (200, 800) according to claim 19, wherein, the processing circuit (820) is further configured to perform the method according to any one of claims 10 to 12.
21. A computer program product comprising executable instructions that, when executed by a processing circuit in a server node (100, 700), cause the server node (100, 700) to perform any one of the methods recited in claims 1 to 8.
22. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processing circuit in a server node (100, 700), cause the server node (100, 700) to perform any one of the methods recited in claims 1 to 8.
23. A computer program product comprising executable instructions that, when executed by a processing circuit in a client device (200, 800), cause the client device (200, 800) to perform any one of the methods recited in claims 9 to 12.
24. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processing circuit in a client device (200, 800), cause the client device (200, 800) to perform any one of the methods recited in claims 9 to 12.
Citation Information
Patent Citations
Method for supporting multiple layers in split rendering
US10446119B1