Implementing and evaluating split rendering over 5G networks
A framework for evaluating XR systems optimizes split rendering over 5G networks by simulating RAN and assessing media data at various levels, addressing configuration challenges and ensuring low latency and high quality video streaming.
Patent Information
- Application Number
- JP2022568887
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-17
- Filing Date
- 2021-05-18
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2041-05-18
AI Technical Summary
Existing technologies face challenges in efficiently evaluating and optimizing the configuration of components involved in split rendering for interactive media applications, particularly in extended reality (XR) systems, which affect the quality and latency of video streaming over 5G networks.
A comprehensive evaluation framework is implemented to assess the effectiveness of XR systems by simulating radio access networks (RAN) and evaluating media data at the video, slice, and packet levels, using techniques such as RAN simulation, encoding, and decoding to determine frame quality and overall system configuration parameters.
This framework allows for the optimization of system configurations to maintain low latency and high video quality in XR applications, ensuring effective split rendering over 5G networks by identifying optimal encoding and distribution parameters.
Smart Images

Figure 0007752135000002 
Figure 0007752135000003 
Figure 0007752135000004
Abstract
Description
[Technical Field]
[0001]
[0001] This application claims priority to U.S. Application No. 17 / 322,468, filed May 17, 2021, and U.S. Provisional Application No. 63 / 026,498, filed May 18, 2020, the entire contents of each of which are incorporated herein by reference. U.S. Application No. 17 / 322,468 claims the benefit of U.S. Provisional Application No. 63 / 026,498, filed May 18, 2020.
[0002] FIELD OF THE DISCLOSURE
[0002] This disclosure relates to the storage and transport of encoded media data. [Background technology]
[0003] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radiotelephones, video teleconferencing devices, etc. Digital video devices implement video compression techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), and extensions to such standards, for more efficiently transmitting and receiving digital video information.
[0004] After the video and other media data is encoded, the media data may be packetized for transmission or storage and assembled into a video file that conforms to any of a variety of standards, such as the International Organization for Standardization (ISO) Base Media File Format and its extensions. Summary of the Invention
[0005]
[0005] Generally, this disclosure describes techniques related to evaluating the configuration of various components involved in performing split rendering for interactive media, such as online cloud video gaming. These techniques may be implemented by a computing device or a system of computing devices. These techniques include an evaluation framework for evaluating the effectiveness of these devices and their configurations. For example, different parts of the system may be evaluated in different ways to determine different configuration settings. In particular, these techniques include evaluation at the video level, slice level, and packet level. Evaluation at the packet level may be used to determine whether a corresponding slice was properly received. Evaluation at the slice level may be used to determine the quality of the slice and corresponding frames that comprise the slice. Evaluation at the video level may be used to determine whether an overall set of configuration parameters for all components of the system is sufficient.
[0006]
[0006] In one example, a method for processing media data includes receiving tracking and sensor information from an extended reality (XR) client device, generating scene data using the tracking and sensor information, encoding the video frames to form encoded video frames, the scene data comprising one or more video frames, performing a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network, decoding the distributed encoded video frames according to the RAN simulation to form decoded video frames, calculating a value representing individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames, and determining an overall quality value from the value representing the individual frame quality for each of the video frames.
[0007]
[0007] In another example, a device for processing media data includes a memory configured to store video data and one or more processors implemented in circuitry, the one or more processors configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information; encode the video frames to form encoded video frames, the scene data comprising one or more video frames of video data; perform a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decode the delivered encoded video frames according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determine an overall quality value from the value representing the individual frame quality for each of the video frames.
[0008]
[0008] In another example, a computer-readable storage medium stores instructions that, when executed, cause a processor to receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information; encode the video frames to form encoded video frames, the scene data comprising one or more video frames; perform a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decode the delivered encoded video frames according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determine an overall quality value from the values representing the individual frame quality for each of the video frames.
[0009]
[0009] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0010] [Figure 1]
[0010] FIG. 1 is a block diagram illustrating an example computing system that may implement techniques of the present disclosure. [Figure 2]
[0011] 2 is a conceptual diagram illustrating an example breakdown for simulating and measuring the performance of the split rendering process implemented by the computing system of FIG. 1 in accordance with techniques of this disclosure. [Figure 3]
[0012] 3A-3C are exemplary graphs for various radio access network (RAN) simulations as described with respect to FIG. 2. [Figure 4]3A-3C are exemplary graphs for various radio access network (RAN) simulations as described with respect to FIG. 2. [Figure 5] 3A-3C are exemplary graphs for various radio access network (RAN) simulations as described with respect to FIG. 2. [Figure 6]
[0013] FIG. 1 is a block diagram illustrating an example vTrace system model in accordance with techniques of this disclosure. [Figure 7]
[0014] FIG. 1 is a block diagram illustrating an example system for performing a RAN simulation in accordance with techniques of this disclosure. [Figure 8]
[0015] FIG. 1 is a block diagram illustrating an example system for measuring the performance of XR content distribution configuration and modeling in accordance with techniques of this disclosure. [Figure 9]
[0016] FIG. 1 is a block diagram illustrating an example pipeline for real-time generation of v-trace data in accordance with techniques of this disclosure. [Figure 10]
[0017] 1 is a block diagram illustrating an exemplary rendering and encoding system in accordance with techniques of this disclosure. [Figure 11]
[0018] 1 is a graph illustrating an exemplary comparison between average PSRN and average PSNR in decibels (dB) for various encoder parameters. [Figure 12]
[0019] 1 is a flowchart illustrating an example method for processing media data in accordance with techniques of this disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011]
[0020] 1 is a block diagram illustrating an exemplary computing system 100 that may implement techniques of this disclosure. In this example, the computing system 100 includes an extended reality (XR) server device 110, a network 130, an XR client device 140, and a display device 152. The XR server device 110 includes an XR scene generation unit 112, an XR viewport pre-rendering rasterization unit 114, a 2D media coding unit 116, an XR media content distribution unit 118, and a 5G system (5GS) distribution unit 120. The network 130 may correspond to any network of computing devices that communicate according to one or more network protocols, such as the Internet. In particular, the network 130 may include a 5G radio access network (RAN) that includes access devices through which the XR client device 140 connects to the access network 130 and the XR server device 110. In other examples, other types of networks, such as other types of RANs, may be used. The XR client device 140 includes a 5GS delivery unit 150, a tracking / XR sensor 146, an XR viewport rendering unit 142, a 2D media decoder 144, and an XR media content delivery unit 148. The XR client device 140 also interfaces with a display device 152 for presenting XR media data to a user (not shown).
[0012]
[0021] In some examples, the XR scene generation unit 112 may correspond to an interactive media entertainment application, such as a video game, which may be executed by one or more processors implemented in the circuitry of the XR server device 110. The XR viewport pre-rendering rasterization unit 114 may format scene data generated by the XR scene generation unit 112 as pre-rendered two-dimensional (2D) media data (e.g., video data) for a user's viewport of the XR client device 140. The 2D media encoding unit 116 may encode the formatted scene data from the XR viewport pre-rendering rasterization unit 114 using a video encoding standard, such as ITU-T H.264 / Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 Versatile Video Coding (VVC), etc. In this example, the XR media content delivery unit 118 represents a content delivery sender. In this example, the XR media content delivery unit 148 represents the content delivery recipient, and the 2D media decoder 144 may perform error handling.
[0013]
[0022] Generally, the XR client device 140 may determine a user's viewport, e.g., the direction the user is looking and the user's physical location, which may correspond to the orientation of the XR client device 140 and the geographic location of the XR client device 140. The tracking / XR sensor 146 may determine such location and orientation data using, for example, a camera, an accelerometer, a magnetometer, a gyroscope, etc. The tracking / XR sensor 146 provides the location and orientation data to the XR viewport rendering unit 142 and the 5GS distribution unit 150. The XR client device 140 provides the tracking and sensor information 132 to the XR server device 110 via the network 130. The XR server device 110 receives the tracking and sensor information 132 and provides this information to the XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114. In this way, the XR scene generation unit 112 can generate scene data for the user's viewport and location, and then pre-render 2D media data for the user's viewport using the XR viewport pre-rendering rasterization unit 114. Thus, the XR server device 110 may deliver the encoded, pre-rendered 2D media data 134 to the XR client device 140 over the network 130, for example, using a 5G wireless configuration.
[0014]
[0023] The XR scene generation unit 112 may receive data representing a type of multimedia application (e.g., a type of video game), the state of the application, multiple user actions, etc. The XR viewport pre-rendering rasterization unit 114 may format the rasterized video signal. The 2D media encoding unit 116 may be configured with a specific decoder / codec, a bit rate for media encoding, a rate control algorithm and corresponding parameters, data for forming picture slices of video data, low-latency encoding parameters, error resilience parameters, intra-prediction parameters, etc. The XR media content delivery unit 118 may be configured with real-time transport protocol (RTP) parameters, rate control parameters, error resilience information, etc. The XR media content delivery unit 148 may be configured with feedback parameters, error concealment algorithms and parameters, post-correction algorithms and parameters, etc.
[0015]
[0024] Raster-based split rendering refers to when the XR server device 110 runs an XR engine (e.g., the XR scene generation unit 112) to generate an XR scene based on information coming from an XR device, e.g., the XR client device 140, and the tracking and sensor information 132. The XR server device 110 may rasterize the XR viewport and perform XR pre-rendering using the XR viewport pre-rendering rasterization unit 114.
[0016]
[0025] 1, the viewport is primarily rendered on the XR server device 110, but the XR client device 140 can perform up-to-date pose correction using, for example, asynchronous time-warping or other XR pose correction to address pose changes. The XR graphics workload can be split between a rendering workload on the powerful XR server device 110 (in the cloud or at the edge) and pose correction (such as asynchronous timewarp (ATW)) on the XR client device 140. Low motion-to-photon latency is maintained via on-device asynchronous time-warping (ATW) or other pose correction methods implemented by the XR client device 140.
[0017]
[0026] In some examples, the latency from rendering video data by the XR server device 110 and the XR client device 140 receiving such pre-rendered video data may be in the range of 50 milliseconds (ms). The latency for the XR client device 140 to provide location and position (e.g., pose) information may be less, e.g., 20 ms, but the XR server device 110 may perform asynchronous time warping to compensate for the latest pose at the XR client device 140.
[0018]
[0027] The following call flow is an example that highlights the steps to implement these techniques.
[0019] 1) The XR client device 140 connects to the network 130 and joins an XR application (e.g., executed by the XR scene generation unit 112).
[0020] a) The XR client device 140 sends static device information and capabilities (supported decoders, viewports).
[0021] 2) Based on this information, the XR server device 110 sets up the encoder and format.
[0022] 3) Loop: a) The XR client device 140 collects the XR pose (or predicted XR pose) using the tracking / XR sensor 146.
[0023] b) The XR client device 140 sends XR pose information in the form of tracking and sensor information 132 to the XR server device 110.
[0024] c) The XR server device 110 uses the tracking and sensor information 132 to pre-render the XR viewport via the XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114.
[0025] d) The 2D media encoding unit 116 encodes the XR viewport.
[0026] e) The XR media content delivery unit 118 and the 5GS delivery unit 120 send compressed media to the XR client device 140 along with data representing the XR pose for which the viewport was rendered.
[0027] f) The XR client device 140 decompresses the video data using the 2D media decoder 144.
[0028] g) The XR client device 140 uses the XR pose data comprising the video frames and the actual XR pose from the tracking / XR sensor 146 for improved prediction and to correct the local pose using, for example, ATW performed by the XR viewport rendering unit 142.
[0029] According to TR26.928, section 4.2.2, the relevant processing and delay components are summarized as follows:
[0030] User interaction delay is defined as the duration between the moment a user action is initiated and the time such action is taken into account by the content creation engine. In the gaming context, this is the time between the moment a user interacts with a game and the moment the game engine processes such player response.
[0031] The age of content is defined as the duration between the moment the content is created and the time it is presented to the user. In the gaming context, this is the time between the creation of a video frame by a game engine and the time the frame is finally presented to the player.
[0032]
[0029] Therefore, the roundtrip interaction delay is the sum of the content age and the user interaction delay. In raster-based split rendering in cloud gaming applications, where part of the rendering is done on the XR server and the service creates a framebuffer as a result of the rendering of the content state, the following processes contribute to such delay:
[0033] User interaction delays (pauses and other interactions) ○ Capturing user interactions in the game client, ○ Delivery of user interactions to the game engine, i.e. to the server (aka network latency), ○ User interaction handling by the game engine / server, Age of content o Creation of one or more video buffers (e.g., one for each eye) by the game engine / server, ○ Encoding the video buffer into video stream frames, ○ Delivery of video frames to game clients (aka network latency), ○ Decoding of video frames by the game client, ○ Presentation of video frames to the user (aka frame rate delay).
[0034] When the XR client device 140 applies ATW, the motion-to-photon latency requirement (of at most 20 ms) is met by the internal processing of the XR client device 140. What determines the network requirements for split rendering are the pose-to-render-to-photon time and the round-trip interaction delay. According to TR26.928, section 4.5, the allowed downlink latency is typically 50-60 ms.
[0035] The rasterized 3D scene available in the frame buffer (see TR26.928, section 4.4) is provided by the XR scene generation unit 112 and needs to be encoded, distributed, and decoded. According to TR26.928, section 4.2.1, the relevant format for the frame buffer is 2k x 2k per eye, potentially even higher. The frame rate is expected to be at least 60 fps, potentially as high as 90 fps. The format of the frame buffer is the usual textured video signals, which are then directly rendered. Since the processing is graphics-centric, formats other than the commonly used 4:2:0 and YUV signals can be considered.
[0036]
[0032] For practical considerations, NVIDIA encoding capabilities may be used, the parameters of such encoders are documented at developer.nvidia.com / nvidia-video-codec-sdk.
[0037] The techniques of this disclosure can be used to address several challenges and accomplish several tasks. For example, these techniques can be used to evaluate basic system design options and their performance, generate traffic models for evaluation of Radio Access Network (RAN) options, provide guidelines for good parameter settings for encoding, content delivery, and RAN configuration, identify capacity for such types of applications, and define potential optimizations. These techniques can simulate these various factors in a reasonable setup.
[0038] In a first example, a complete simulation for system 100 may be performed using split rendering. In this example, a 5G New Radio (NR) configuration and simulation may be performed for communication over network 130, including, for example, tracking and sensor information 132 and pre-rendered 2D media data 134. Tracking / XR sensors 146 may track and detect exemplary user movements. The quality of the video data presented on display device 152 during these various simulations may then be measured.
[0039] To implement this first example, a simulation may be performed using various models performing separate tasks. A source video model may include actions performed by the XR scene generation unit 112, the XR viewport pre-rendering rasterization unit 114, the XR viewport rendering unit 142, the 5GS delivery unit 120, and the display device 152. A content distribution model may include actions performed by the 2D media encoding unit 116, the XR media content distribution unit 118, the 5GS delivery unit 120, the 5GS delivery unit 150, and the 2D media decoder 144. An uplink model may include actions performed by the tracking / XR sensor 146 and the 5GS delivery unit 150. Tracking and sensor information 132 may be generated as part of the uplink traffic model, and a RAN simulation may be performed to generate exemplary pre-rendered 2D media data 134.
[0040] The various components of the XR server device 110, the XR client device 140, and the display device 152 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. The functionality attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that instructions for the software or firmware may be stored on a computer-readable medium and executed by the requisite hardware.
[0041] Based on the system design in S4-200771, for example, when some aspects can be addressed using some techniques of the present disclosure, there are some challenges for potential simulation and for generating traffic models. Such challenges include:
[0042] Evaluate basic system design options and their performance.
[0043] Generate traffic models for evaluation of RAN options.
[0044] · You do not violate any content license terms.
[0045] To provide guidelines for good parameter settings for coding, content distribution, and for RAN configuration.
[0046] Identifying capacity for such types of applications.
[0047] · Defining potential optimizations.
[0048] Simulate this in a reasonable setup in a repeatable manner.
[0049] Simulation and modeling for each system 100 can be separated into separate individual components.
[0050] Source video model Encoding and content delivery Wireless access distribution Content distribution receivers and decoders Display process Uplink traffic
[0039] Figure 2 is a conceptual diagram illustrating an example breakdown of components of the XR server device 110 for simulating and measuring the performance of the split rendering process implemented by the computing system 100 of Figure 1 in accordance with the techniques of this disclosure. The example breakdown of Figure 2 corresponds to the first example described above. In this example, the system 160 includes a game engine 162, a model encoding device 164, a content encoding and distribution model 166, a 5GS simulation unit 168, and a content distribution and decoding model 170. 1, the model encoding device 164 may correspond to the 2D media encoding unit 116 of FIG. 1, the content encoding and distribution model 166 may correspond to the XR media content delivery unit 118 and the 5GS delivery unit 120 of FIG. 1, the 5GS simulation unit 168 may correspond to the network 130 of FIG. 1, and the content distribution and decoding model 170 may correspond to the 5GS delivery unit 150, the XR media content delivery unit 148, the 2D media decoder 144, and the XR viewport rendering unit 142 of FIG. 1.
[0051]
[0040] Game engine 162 receives data 178 including pose model and trace data, game data, and game / engine configuration data to simulate an extended reality (XR) game using the received pose model and trace data. Game engine 162 generates video frames 180 from the received data and provides the video frames 180 to model encoding device 164. Model encoding device 164 encodes the video frames 180 to generate encoded video frames and v-trace data, and provides the encoded video frames and v-trace data 182 to content encoding and delivery model 166.
[0052] The content encoding and distribution model uses the encoded video frame and v trace data 182 to generate slice trace (s trace) data 184 and packets representing encapsulated, coded slices of the encoded video frame using exemplary configuration parameters 188. The content encoding and distribution model 166 provides the s trace data 184 and packets to a 5GS simulation unit 168. The 5GS simulation unit 168 performs a RAN simulation using the received s trace data 184, packets, and configuration parameters 190 to generate s′ trace data 186. The s trace data 184 represents a slice of video data before wireless transmission, and the s′ trace data 186 represents a slice of video data after wireless transmission (e.g., which may be corrupted by wireless transmission). The 5GS simulation unit 168 provides the s′ trace data 186 and the delivered packets, encapsulating the coded slices of video data, to a content distribution and decoding model 170. The content distribution and decoding model 170 performs an exemplary decoding process on the data in the packets.
[0053] Using the coded and decoded data from the content encoding and distribution model 166 and the content distribution and decoding model 170, a value representing per-frame quality 172 may be generated to test the performance of the RAN simulation performed by the 5GS simulation unit 168 and the configuration parameters 190. A combination / aggregation of the individual values of per-frame quality 172 may be used to generate an overall quality value 174.
[0054] Targets may be set for values for per-frame quality 172 and overall quality 174 to adjust various configuration parameters. In one example, the target may be to have no change in video quality as measured by per-frame quality 172 and overall quality 174. One or a few application metrics may be derived from RAN simulations performed by the 5GS simulation unit 168. For each scenario, several v-traces (e.g., 1-3) may be performed, but may be performed over a duration of several hours. The quantization parameter may be unchanged in the configuration parameters 188. Several different values for the configuration parameters 188 may be adjusted, which may cause bitrate changes and traffic characteristics as side effects. However, no quality changes should be made. A relatively small number of configurations, e.g., 3-10, may be tested.
[0055] The game engine 162, the model encoding device 164, the content encoding and distribution model 166, the 5GS simulation unit 168, and the content distribution and decoding model 170 may be implemented using one or more processors implemented in a circuit, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functionality attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that instructions for the software or firmware may be stored on a computer-readable medium and executed by the requisite hardware.
[0056] Based on the concepts explained above, several parameters are relevant for the overall system design.
[0057] Games: ○ Game Type ○ Game state, ○ Multi-user actions, etc. User interaction: ○ 6DOF pose based on head and torso movement, ○ Game interaction using a controller The format of the rasterized video signal. Typical parameters are:
[0058] ○ 1.5K x 1.5K per eye at 60, 90, 120fps ○ 2Kx2K at 60, 90, 120fps ○ YUV4:2:0 or 4:4:4 Encoder configuration ○ Codec: H.264 / AVC or H.265 / HEVC ○ Bitrate: Set the bitrate to a specific value (e.g. 50Mbit / s) Rate control: CBR, capped VBR, feedback-based, CRF, QP ○ Slice setting: 1 per frame, 1 per MB row, X per frame Intra configuration and error resilience: normal IDR, GDR pattern, adaptive intra, feedback-based intra, feedback-based predication and ACK-based, feedback-based prediction and NACK-based Latency settings: P-picture only, look-ahead unit Complexity settings for the encoder Content distribution Slice to IP mapping: fragmentation ○ RTP-based timecode and packet numbering ○ RTP / RTCP based feedback ACK / NACK ○ RTP / RTCP based feedback on bitrate 5G System / RAN Configuration: ○ QoS settings (5QI): GBR, latency, loss rate ○ HARQ transmission, scheduling, etc. Content Delivery Receiver Configuration: ○ Loss detection: Sequence number Delay / Latency Handling Error Resilience ATW Quality aspects ○ (encoded) video quality Lost data Immersiveness
[0046] Video quality can be affected by various factors. One factor is coding artifacts based on encoding. Such artifacts can be determined, for example, by the peak signal to noise ratio (PSNR). Another factor is artifacts due to lost packets and resulting error propagation.
[0059]
[0047] In accordance with the techniques of this disclosure, quality may be modeled for each simulation as a combination of:
[0060] The average PSNR delivered from the encoding, and Percentage of damaged video area where a damaged macroblock is defined as ◆ If it is part of the slice that was lost for this transmission ◆ If it is received correctly, but it is predicted from the wrong macroblock ○ The macroblock (or coding tree unit) is correct ◆ If it is received correctly and it predicts for an undamaged macroblock ◆ Predicting from intact macroblocks means that · Correct spatial predictions · Time predictions are correct o Reconstructing the MB is done by intra-refresh or by predicting from the same correct MB.
[0061] Generally, references to macroblocks above and throughout this disclosure refer to macroblocks in ITU-T H.264 / AVC. However, it should be understood that the concept of macroblock can be interchanged with coding unit (CU) or coding tree unit (CTU) in ITU-T H.265 / HEVC or ITU-T H.266 / VVC without loss of generality.
[0062] In the example of FIG. 2, a source model may be provided based on inputs describing the game being tested, a pause model, and a statistical model of the video signal that may be used to apply content distribution and encoding procedures. Based on these traces, 3GPP SA4 may define a content encoding and distribution model based on TR26.928 and taking into account system design as documented in S4-200771. The system model includes encoding, distribution, decoding, and may also include quality definitions. The system may also provide an interface to RAN simulation. RAN groups may then use this model to simulate different traffic characteristics and evaluate different performance options.
[0063] The following elements may be used by system 160 and other systems described herein.
[0064] 1) V-Trace: A video trace that provides sufficient information from the model encoding to understand the complexity and data rate of the encoding model, and also takes into account delivery options such as bitrate control, intra-refresh, and feedback-based error resilience. Details are forthcoming, but we are currently investigating parameters provided by the model encoding from x265 encoding runs.
[0065] 2) S-trace: A sequence of slices as the output of the model encoder. Each slice is assigned:
[0066] a. Related frames (timestamps) b. Slice size c. Slice quality (possibly some more information from the V-trace, e.g. complexity) d. Number of macroblocks (area covered) e. The timing of the slice, e.g., deadline or at least origin timestamp.
[0067] 3) S'trace: a. All information from the S trace b. Slice Loss Indicator c. Slice delay information
[0051] Figure 3 is an example graph for various RAN simulations such as those described with respect to Figure 2. In this example, the graph represents the percentage of corrupted video as a function of a number of supported users. The graph in Figure 3 represents an example graph that may result from the example above in which the quantization parameter is invariant to prevent changes in video quality. To allow for each number of supported users, various configuration parameters 190 may be modified. As the number of supported users increases, the amount of corrupted video data is expected to increase. However, a set of configuration data may be identified that allows for a relatively large number of supported users and a relatively small percentage of corrupted video, such as "Configuration 3" as shown in Figure 3.
[0068] Referring again to FIG. 2, in another example, video quality may be allowed to vary as measured by frame-by-frame quality 172 and overall quality 174. In this example, one or a few application metrics may be derived from RAN simulations performed by the 5GS simulation unit 168. For each scenario, several v-traces (e.g., 1-3) may be performed, but may be performed over a duration of several hours. Several different values for configuration parameters 188, including quantization parameters, may be adjusted, which may cause bitrate changes and changes in traffic characteristics as side effects. However, quality changes should not be made. A relatively small number of configurations may be tested, e.g., 3-10.
[0069] 4 and 5 are example graphs for various RAN simulations as described with respect to FIG. 2. These example graphs correspond to the above example where the quantization parameter was allowed to vary. In these examples, the graphs plot delivered quantizer / video quality data against the percentage of corrupted video for a fixed number of users in various scenarios. As shown in FIG. 5, the number of users and configuration parameters can be selected to achieve high video quality, low bitrate, and a low percentage of lost data, e.g., 0.1% corrupted video.
[0070] To implement the various test scenarios described above, representative source data can be generated to implement v-traces. S-traces can also be implemented, for example, using content coding, distribution, and quality, including distribution parameters, radio mapping, modeling, and quality definitions, on various system designs. RAN distribution simulations can be based on the S-traces mentioned above.
[0071]
[0055] Figure 6 is a block diagram illustrating an example v-trace system model 200 in accordance with the techniques of this disclosure. The v-trace system model 200 represents a model for evaluating source video encoding complexity. In this example, the v-trace system model 200 includes a game engine 202, a predictive model encoding unit 204, an intra-model encoding unit 206, and a trace combining unit 208. The v-trace system model 200 may generally correspond to a portion of the system 160 of Figure 2. For example, the game engine 202 may generally correspond to the game engine 162 of Figure 2, and the predictive model encoding unit 204, the intra-model encoding unit 206, and the trace combining unit 208 may correspond to the model encoding device 164. The game engine 202 receives input data 178, which includes pose and model trace data, game data, and game configuration data.
[0072]
[0056] The game engine 202 generates video frames 210 and provides the video frames 210 to each of the prediction model encoding unit 204 and the intra model encoding unit 206. The video frames 210 may be, for example, extended reality (XR) split data at 60 fps, with one frame per eye to achieve the XR effect. In this example, the v-trace system model 200 includes two tracks: one for intra-prediction and one for inter-prediction. In video coding, intra-prediction is performed by predicting a block of video data with a neighboring previously decoded block of the same frame, and inter-prediction is performed by predicting a block of video data using a reference block of a previously decoded frame.
[0073] The predictive model coding unit 204 may generate VP trace data 212, and the intra-model coding unit 206 may generate VI trace data 214. The trace combining unit 208 receives the VP trace data 212 and the VI trace data 214 and generates V trace data 216. The V trace data 216 may be formatted according to the FFMPEG passlogfile format after a coding pass. Such formats are described, for example, at ffmpeg.org / wiki / Encode / H.264 and slhck.info / video / 2017 / 03 / 01 / rate-control.html. The trace combining unit 208 may also use the configuration parameters "-pass[:stream_specifier] n (output,per-stream)" and "-passlogfile." Pseudocode for an exemplary algorithm for generating the v trace data 216 in this format is shown below.
[0074]
number
[0075]
[0058] As an alternative to ITU-T H.264 / AVC, the video data may be coded using ITU-T H.265 / HEVC or other video coding standards. For the H.265 example, an exemplary log file is described at x265.readthedocs.io / en / default / cli.html#input-output-file-options. Such a log file may include parameters such as:
[0076] · Encoding order: The order in which the encoder encodes the frames.
[0077] · Type: The slice type of the frame.
[0078] · POC: Picture Order Count - The display order of the frames.
[0079] · QP: The quantization parameter determined for the frame.
[0080] Bits: The number of bits consumed by the frame.
[0081] Scenecut: 1 if the frame is a scenecut, 0 otherwise.
[0082] RateFactor: Only applicable when CRF is enabled. The rate factor depends on the CRF given by the user. This is used to determine the QP to target a certain quality.
[0083] · BufferFill: Bits available for the next frame, including bits carried over from the current frame.
[0084] · BufferFillFinal: The buffer bits available after removing frames from the CPB.
[0085] · Latency, expressed in number of frames, between when a frame is put in and when it is put out.
[0086] · PSNR: Peak signal to noise ratio for Y, U and V planes.
[0087] · SSIM: A quality metric that indicates the structural similarity between frames.
[0088] · Ref List: POC of reference pictures in list 0 and 1 for the frame.
[0089] Some statistics about the encoded bitstream and encoder performance are available when --csv-log-level is greater than or equal to 2.
[0090] In yet other examples, other video codecs may be used, such as ITU-T H.266 / Versatile Video Coding (VVC), etc. Similar sets of data may be generated for analysis using VVC or other such video codecs.
[0091]
[0060] The trace combining unit 208 may generate the v-trace data 216 according to several v-trace generation principles. Some v-trace generation principles are explained at slhck.info / video / 2017 / 02 / 24 / crf-guide.html. The trace combining unit 208 may use H.264 / AVC using x264 in FFMPEG and H.265 / HEVC using x265 in FFMPEG. The trace combining unit 208 may use a near-lossless constant quality mode, where options include:
[0092] · Fixed QP Constant rate factor Recommendations from slhck.info / video / 2017 / 02 / 24 / crf-guide.html for QP / CRF settings: For x264, typical values are between 18 and 28. The default value is 23. 18 should be visually transparent.
[0093] ○ For x265, typical values are between 24 and 34. The default value is 28. 24 should be visually transparent.
[0094] A change of ±6 should result in approximately halving / doubling the file size, but results may vary.
[0095] o Multiple CRFs can be run in parallel from one output.
[0096] It can be determined whether the discussion in slhck.info / video / 2017 / 02 / 24 / crf-guide.html also applies to XR and video games.
[0097] The game engine 202, the predictive model encoding unit 204, the intra-model encoding unit 206, and the trace combination unit 208 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. The functionality attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that instructions for the software or firmware may be stored on a computer-readable medium and executed by the requisite hardware.
[0098] 7 is a block diagram illustrating an example system 220 for performing a RAN simulation in accordance with the techniques of this disclosure. The system 220 includes a content coding and distribution model 222, a slice-to-Internet Protocol Packet / Radio Link Control (IP / RLC) mapping unit 224, a RAN simulation unit 226, and an IP / RLC-to-slice mapping unit 228. The system 220 may generally correspond to a portion of the system 160 of FIG. 2. For example, the content coding and distribution model 222 may correspond to the content coding and distribution model 166 of FIG. 2, and the slice-to-IP / RLC mapping unit 224, the RAN simulation unit 226, and the IP / RLC mapping unit 228 may correspond to the 5GS simulation unit 168 of FIG. 2.
[0099] In this example, the content encoding and distribution model 222 receives v-trace data 216 from the trace combination unit 208 (FIG. 6), which includes data representing encoding of frames of video data. The content encoding and distribution model 222 also receives global configuration data 230. Using the global configuration data 230 (which may correspond to the configuration parameters 188 of FIG. 2), the content encoding and distribution model 222 may encode the frames of video data to form encoded slices of the video data and to form s-trace data 232 representing the encoded slices. The content encoding and distribution model 222 may produce slices over time, where each slice corresponds to a sequence of macroblocks (ITU-T H.264 / AVC) or maximal coding units / coding tree units (ITU-T H.265 / HEVC).
[0100] The slice-IP / RLC mapping unit 224 may use the Real-time Transport Protocol (RTP) to form IP packets and RLC fragments from slices of the s-trace data 232. The slice-IP / RLC mapping unit 224 may be configured with IP packet and payload size information for forming packets and fragments. Generally, if a packet is lost, the entire slice to which the packet corresponds is also lost. The slice-IP / RLC mapping unit 224 produces p-trace data 234, which the slice-IP / RLC mapping unit 224 provides to the RAN simulation unit 226. The RAN simulation unit 226 is configured with delay and loss requirements for each slice / packet, and may, through the RAN simulation, determine whether each slice was received or lost, as well as a latency value or a received timestamp value for each slice, at the end of the simulation.
[0101] The RAN simulation unit 226 may provide P′ trace data 236, which includes data representing whether slices were received or lost and the latency / reception time for received slices, to the IP / RLC to slice mapping unit 228. The IP / RLC to slice mapping unit 228 provides the s′ trace data 238 to the content coding and delivery model 222. The content coding and delivery model 222 may determine an overall quality value 240 by comparing video data decoded prior to the RAN simulation with video data decoded following the RAN simulation.
[0102] The RAN simulation unit 226 may be configured according to the maximum delay requirements for each slice downlink. For example, there may be a 10 ms MAC-to-MAC deadline and an external lossless requirement. For the uplink, there may be tracking, sensor, and pause information, content delivery uplink information, and pause update traffic (e.g., at a periodicity of 1.25 ms or 2 ms). The system may determine how the pause frequency affects quality, which may depend on the XR service type. For gaming, the pause frequency may have a large impact. Also, the pause and game action may be synchronized.
[0103]
[0067] RAN simulation options may include open-loop and closed-loop options. In the open-loop case, for the entire V trace, the system may generate s trace data, p trace data, RAN simulation one-way, p' trace data, and s' trace data, and then perform a quality assessment in the form of, for example, an overall quality value 240. In the closed-loop case, for each v trace entry, the system may generate s trace data, p trace data, RAN simulation one-way, p' trace data, and s' trace data, and then feed the resulting data back to the next s trace and generate a quality assessment in the form of an overall quality value 240.
[0104] The content encoding and distribution model 222, the slice-to-IP / RLC mapping unit 224, the RAN simulation unit 226, and the IP / RCL-to-slice mapping unit 228 may be implemented using one or more processors implemented in a circuit, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functionality attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that instructions for the software or firmware may be stored on a computer-readable medium and executed by the requisite hardware.
[0105] FIG. 8 is a block diagram illustrating an example system 250 for measuring the performance of an XR content distribution configuration and modeling in accordance with the techniques of this disclosure. The system 250 includes a video content model 252, a content encoding and distribution model 254, a radio access network (RAN) simulation unit 256, a content distribution and decoding model 258, and a content distribution and decoding model 260. These elements of the system 250 may correspond to the respective components of the system 160 of FIG. 2. For example, the video content model 252 may correspond to the game engine 162 and the model encoding device 164 of FIG. 2, the content encoding and distribution model 254 may correspond to the content encoding and distribution model 166 of FIG. 2, the RAN simulation unit 256 may correspond to the 5GS simulation unit 168, and the content distribution and decoding model 258 may correspond to the content distribution and decoding model 170 of FIG. 2.
[0106] In this example, the decoded media data is compared before and after passing through the RAN simulation unit 256. The video data from the content distribution and decoding model 258 that has passed through the RAN simulation unit 256 is assessed to determine a value for per-frame quality 262. The decoded media data from the content distribution and decoding model 260 and the value for per-frame quality 262 are used to determine overall quality data 264.
[0107] In particular, the video content model 252 provides the v-trace data 268 and the encoded video data to be transmitted to the content encoding and delivery model 254. The content delivery model 254 receives the global configuration data 266 and uses the global configuration data 266 to encode and deliver the video data corresponding to the v-trace data 268. The global configuration data 266 may include, for example, bitrate control information (e.g., for selecting one or more of constant quality, constant bitrate, feedback-based variable bitrate, or fixed rate factor), slice settings (number of slices, maximum slice size), error resilience (frame-based, slice-based, regular intra refresh, feedback-based intra refresh, or feedback-based prediction), and feedback data (off, statistical, or operational). The maximum slice size setting may depend on various statistics, including the number of slices, which may vary.
[0108] The content encoding and distribution model 254 outputs s trace data 270 to the RAN simulation unit 256 and outputs s′ trace data 270′ to the content distribution and decoding model 260. The RAN simulation unit 256 performs a RAN simulation on the s trace data 270 and outputs s′ trace data 272 to the content distribution and decoding model 258. The content distribution and decoding model 260 outputs q trace data 274, and the content distribution and decoding model 258 outputs a value for per-frame quality 262, resulting in q′ trace data 276. The overall quality data 264 may be one or more quality measures for the entire video sequence (e.g., an aggregation of the quality for individual video frames).
[0109] Various content coding considerations may influence the encoding of video data corresponding to v trace data 268 to form s trace data 270. In one example, the objectives are to have consistent quality, to use ITU-T H.265 / HEVC encoding, to form a fixed number of slices, where one slice is coded using intra prediction in opportunistic mode, and to have no feedback. In this example, bitrate control may use a fixed rate factor of 28, a number of slices of 10, a periodic intra refresh value of 10, periodic intra refresh and reference picture invalidation, and feedback is off. To create the slice size, the content encoding and delivery model 254 may generate 10 slices per video frame and use a fixed rate factor of 28 (where ±6 results in approximately halving / doubling the frame size, with a 12% decrease / increase for each + / -), and based on the VI trace, one intra slice is assumed (i.e., 10% of the adjusted I-frame size), and based on the VP trace, 9 inter slices are generated (where two aspects can be taken into account: the size of the frame (10%) and statistical variation).
[0110] The RAN simulation unit 256 may maintain state data for each macroblock (ITU-T H.264 / AVC) or coding tree unit (ITU-T H.265 / HEVC), such as damaged (predicted from the lost or damaged area, whether temporal or spatial) or correct. The RAN simulation unit 256 may provide feedback to the content encoding and distribution model 254, including data representing the number of lost or damaged blocks of video data. The content encoding and distribution model 254 may integrate the feedback for subsequent encoding. For example, the content encoding and distribution model 254 may react to slice loss, whether statistical or operational. In response, the content encoding and distribution model 254 may adjust the bitrate (e.g., obtain the encoding bitrate and adjust the quantization parameter up or down, where + / - 1 may result in a 12% impact). The content coding and delivery model 254 may add an intra-predicted slice if a slice is lost (in the case of a reported loss, significantly more intra-predicted data may be added. An intra-predicted slice may cover a large area depending on the motion vector activity. In some cases, the content coding and delivery model 254 may predict from only the acknowledged region, which may cause a statistical increase in frame size for the lost slice since the most recent slice may not be used for prediction.
[0111] The S trace data 270 may be formatted to include data about a timestamp representing the associated frame, the size of the slice, the quality of the slice (which may include more information, e.g., complexity, from the v trace), the number of macroblocks (in the case of ITU-T H.264 / AVC) or coding tree units (in the case of ITU-T H.265 / HEVC), and / or the timing of the slice (e.g., deadline for receipt of the slice). The S' trace data 270' and 272 may be formatted to include all of the information from the format for the s trace data 270 described above, and additionally data representing slice loss and / or slice delay.
[0112] There may be a fixed deadline for each slice, which may be set to a higher number. For example, for staggered 60 fps video data for each eye, i.e., 120 fps video data, there may be a desired latency set to 7 ms. Drop deadlines may be set differently, which may be useful, but should not complicate the RAN simulation.
[0113]
[0077] Content decoding considerations for producing the Q trace data 274 from the s' trace data 270' may take results from a RAN simulation performed by the RAN simulation unit 256 and may be used to map these simulation results to an estimate for the per-frame quality 262. When determining the value for the per-frame quality 262, slice considerations may include, for example, the quality of the received encoded slices, the number of lost slices that are degraded in quality enough that the slice is considered lost, and delayed slices that may be considered lost. Other considerations may include error propagation (e.g., prediction of slices from slices of degraded quality or lost slices), and the complexity of the frame / slice may determine the error propagation (how fast incorrect data spreads across a frame), the new intra-reset quality, and the percentage of correct / incorrect data for each frame and for different configurations.
[0114] A metric for determining q trace data 274 may be the percentage of inaccurate data, which may affect lost / error propagation. The Q trace data 274 format may include data representing encoding quality, resulting quality after delivery (percentage of corrupted data), and / or data rate of the frame.
[0115]
[0079] System 250 may maintain either a "damaged" or "correct" status for each macroblock (ITU-T H.264 / AVC) or coding tree unit (CTU) (ITU-T H.265 / HEVC). If the macroblock / CTU is damaged, system 250 may determine whether the macroblock / CTU is part of a slice that was lost due to this transmission, or whether the macroblock / CTU was received correctly by being predicted from an incorrectly decoded / lost region of another slice / frame. If the macroblock / CTU is correct, system 250 determines that the macroblock / CTU was received correctly and that it was predicted from an undamaged region of another slice / frame. Predicting from an undamaged region of another slice / frame means that the spatial prediction is correct, the temporal prediction is correct, or the macroblock / CTU is correctly reconstructed using intra-refresh and prediction from a correct region of another slice / frame following intra-refresh.
[0116] The system 250 may determine the overall quality value 264 by averaging the coding quality (e.g., averaging over a quantization parameter and possibly converting to a Peak Signal-to-Noise Ratio (PSNR)), the average incorrect video data (e.g., the average incorrect number of video frames and / or slices), and / or by multiplication / combination / aggregation, if necessary. 3GPP SA4 defines a constant rate factor (CRF)-based coding model with a specific quality factor (e.g., using the FFMPEG default value of 28). Different configurations for error resilience may be applied. A criterion for simulations may be the percentage of incorrect video areas such that at most one macroblock / CTU is incorrect every X (e.g., 60) seconds. For example, if 4096x4096 is used at 60 fps, this results in an average of 10e-6 damaged areas.
[0117] The above-described model may work for pixel-based split rendering and / or cloud gaming. Conversational applications, video streaming applications (e.g., Twitch.tv), and the like may also be used.
[0118] The video content model 252, the content encoding and distribution model 254, the RAN simulation unit 256, the content distribution and decoding model 258, and the content distribution and decoding model 260 may be implemented using one or more processors implemented in a circuit, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. The functionality attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that instructions for the software or firmware may be stored on a computer-readable medium and executed by the requisite hardware.
[0119] System 250 represents an exemplary content modeling based on the discussion in S4-200771. This modeling includes v-model inputs, global configuration for the encoder, statistical or dynamic feedback from content distribution receivers, a decoding model, and a quality model.
[0120]
[0084] Content encoding according to the content encoding and distribution model 254 can be modeled as follows:
[0121] For frame i from the V trace (based on sample time) Read frame i from the V trace (timestamp) ○ Read the latest dynamic information from the dynamic status information ○ Model coding (based on input parameters) ○ For slices s=1, 2, ..., S ◆ Drop slice s with associated parameters ◆ A new slice becomes available and creates an IP packet. ◆ IP packets p=1, 2, ..., P Drop IP packets using the relevant parameters and using the timestamp for S trace 1, but also using parameters such as the number of slices
[0085] Configuration parameters of the global configuration data 266 may include the following:
[0122] Input: Global config ○ Bitrate control: Constant quality, Constant bitrate, Feedback based variable bitrate, Fixed rate factor 28 ○ Slice settings (number of slices: 10, maximum slice size) Error resilience, frame-based, slice-based, periodic intra-refresh10, feedback-based intra-refresh, feedback-based prediction (NVIDIA: periodic intra-refresh, reference picture invalidation) Input: Dynamic for each slice information ○ Off ○ Statistically, the loss with bit rate or some delay ○ Works for bitrate or loss with any delay
[0086] Aspects of model encoding may include the following:
[0123] QP settings and the effect of intra ratio and slice settings Feedback ○ Bitrate adjustment: The encoder obtains the encoding bitrate and adjusts the QP (considered as + / - 1 to 12%) ○ Add Intra in case of lost slices: In case of reported loss, significantly more Intra is added. Intra covers a larger area (depending on motion vector activity) Predicting from ACKs only: statistical increase in frame size for lost slices, so the latest one cannot be used Slice Settings Decoding emulation by the content delivery and decoding model 258 may be based on slice delay and / or loss. Late and lost slices may be considered unavailable and may cause errors (e.g., in decoding subsequent frames / slices that reference the lost or late slice for inter prediction).
[0124]
[0088] The RAN simulation unit 256 may emulate a RAN simulation based on the existing 5Qis.
[0125]
[0089] The quality assessment may be based on two aspects, including the encoding quality and the quality degradation due to lost slices. The following simulation may be used to identify damaged macroblocks (or coding tree units or coding units in ITU-T H.265 / HEVC).
[0126] Keeping state for each macroblock (or CU / CTU) ○ Damaged Correct Damaged macroblocks ○ If it is part of the slice that was lost for this transmission If it was received correctly, but it was predicted from the wrong macroblock Macroblocks are correct If it is received correctly and it predicts for an undamaged macroblock Predicting from intact macroblocks means: ○ Spatial prediction is correct ○ Time prediction is correct o Reconstructing the MB is done by intra-refresh or by predicting from the same correct MB.
[0127]
[0090] Depending on the configuration and settings of the delivered video quality, different results may be obtained. The quality threshold may be, for example, having at most 0.1% of the video area damaged. Also, the quality of the original content may be the threshold.
[0128] 9 is a block diagram illustrating an example pipeline 280 for real-time generation of v-trace data in accordance with the techniques of this disclosure. The pipeline 280 includes a game engine 282, an encoder 284, a decoder 286, raw data storage 288, a predictive encoder 290, an intra-encoder 292, and a v-trace-to-3GPP unit 294. The components of the pipeline 280 may correspond to the respective components of the system 160 of FIG. 2. For example, the game engine 282 may correspond to the game engine 162 of FIG. 2, the encoder 284 may correspond to the model encoding device 164 of FIG. 2, and the decoder 286, raw data 288, predictive encoder 290, intra-encoder 292, and v-trace-to-3GPP unit 294 may correspond to the content encoding and distribution model 166 of FIG. 2.
[0129] The game engine 282 receives input data 296, including pose model and trace data, game data, and game configuration data. The game engine 282 generates video frames 298 from the input data 296. The encoder 284 (which may be an NVIDIA NVENC encoder) encodes the video frames 298 using configuration data 302 to generate a video bitstream 300. The encoder 284 may operate at 600 gigabytes per hour and may run for several hours. The decoder 286, which may be an FFMPEG decoder, decodes the video bitstream 300 to generate decoded video data that is stored as raw data 288 on a computer-readable medium, such as a hard disk or flash drive. The predictive encoder 290 and intra-encoder 292 receive configuration data 304 and encode the raw data 288 to form respective VP trace data and VI trace data. The V trace-to-3GPP unit 294 generates v trace data from the VP trace data and VI trace data.
[0130] 9, as described above, the encoder 284 may generate the bitstream 300, and the V-trace to 3GPP unit 294 may generate the v-trace data. V-trace generation begins by using the high-quality output of the game engine 282, and may be encoded by the encoder 284 using ITU-T H.264 (AVC) for the model game and repeatable trace. The decoder 286 may decode this information into 2K x 2K in 120 fps source frames. This sequence may be sent through a model encoder, including a predictive encoder 290 and an intra encoder 292, to identify the effects of different parameter settings. The following H.265 parameters may be used initially:
[0131] --input viki.yuv(4:2:0) --profile, -P main10 --input-res 2048x2048 · --fps 120fps, 60fps, 30fps --psnr --ssim --frames 500 (first test) --bframes 0 · --crf 22, 25, 28, 31, 34 --csv logfile.csv --csv-log-level 2 --log-level 4 --numa-pools “8” --keyint, -I 1 and -1 --slices 1, 128 (for 2048), 8 (for 2048) --output rvrviki.h265 --rc-lookahead 0 / 1
[0094] Based on these 180 encoding runs, good encoder modeling is expected, such that for longer game sequences only a subset of runs are needed.
[0132] Parameters for generation of V trace data in this example may include the duration of the P trace (e.g., hours), the various games being run by game engine 282, the pause trace, settings for encoder 284 (the goal is, for example, simple, high-quality FFMPEG-decodable video data), and FFMPEG configuration data. Four FFMPEG decoder configurations may be tested and, in some examples, pipelined in parallel.
[0133] HEVC, I-PPP in CRF28 HEVC in CRF28, I-IIIIII AVC and I-PPP using CRF23 AVC with CRF23, I-IIIIII x265.readthedocs.io / en / default / cli.html# describes the possible configuration parameters. These parameters can be used to configure rvrplugin.ini with 2048x2048 resolution, 5ps, and a bitstream generated using an ITU-T H.264 / AVC encoder on the server side, which can run at 60 or 120fps. This data can be re-encoded using crf-24 and updated PSNR and SSIM values. The initial command line arguments can include:
[0134] · ABR (average bitrate): x265.exe --input viki.yuv --input-res 1440x1440 --fps 120 --psnr --ssim --frames 500 --bframes 0 --bitrate 600000 --csv logfile.csv --csv-log-level 2 --log-level 4 --numa-pools “8” --keyint 1 --output rvrviki.h265 CRF (Constant Quality Variable Bitrate): x265.exe --input viki.yuv --input-res 1440x1440 --fps 120 --psnr --ssim --frames 500 --bframes 0 --crf 24 --csv logfilecrf.csv --csv-log-level 2 --log-level 4 --numa-pools “8” --keyint 1 --output rvrviki.h265 The game engine 282, the encoder 284, the decoder 286, the predictive encoder 290, the intra-encoder 292, and the vTrace to 3GPP unit 294 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functionality attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that instructions for the software or firmware may be stored on a computer-readable medium and executed by the requisite hardware.
[0135] 10 is a block diagram illustrating an exemplary rendering and encoding system 310 in accordance with techniques of this disclosure. In this example, the system 310 includes a renderer 312 and an encoder 314. The renderer 312 receives content changes 318 and pose information 316 to generate rendered video data, e.g., at a frame rate of 60 or 120 fps. The pose information 316 may be provided at very high frequency with a 2.5 ms delay. Every 1 / 60 Hz, the renderer 312 may take the latest pose from the pose information 316, render the video data from the content changes 318, and send the rendered image to the encoder 314 for encoding. The encoder 314 may queue an encoding call and then render the image.
[0136]
[0099] The encoder 314 may encode the video data using ITU-T H.265 / HEVC, for example, "x265.exe." The encoder 314 may operate with the following command line options:
[0137] --input viki.yuv(4:2:0) --profile, -P main10 · --input-res 1440x1440(2048x2048) -fps 120fps, 60fps, 30 --psnr --ssim --frames 500 (first test) --bframes 0 · --crf 22, 25, 28, 31, 34 --csv logfile.csv --csv-log-level 2 --log-level 4 --numa-pools “8” --keyint, -I 1 and -1 --slices 1, 90 (for 1440), 128 (for 2048), 6 (for 1440), 8 (for 2048) --output rvrviki.h265 · --analysis-save-reuse-level 10 --rc-lookahead 0 / 1
[0100] The encoder 314 does not need to use "-- frame-dup disabled", "-- constrained-intra", and "-- no-deblock".
[0138] The renderer 312 and the encoder 314 may be implemented using one or more processors implemented in circuitry such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. The functionality attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that instructions for the software or firmware may be stored on a computer-readable medium and executed by the requisite hardware.
[0139] 11 is a graph showing an exemplary comparison between the average PSRN and the average PSNR in decibels (dB) for various encoder parameters. In this example, there is an approximately 25% increase from 0 to 1, an approximately 50% increase from 0 to 2, an approximately 75% increase from 0 to 3, and an approximately 100% increase from 0 to 4. Such an increase may depend on the picture type. For example, if a video sequence generates many intra-predicted frames regardless of the configuration, the increase may be smaller, and very static content may not cause much of an increase. This can be used to model disabled reference behavior.
[0140]
[0103] Error propagation can also be modeled. As a basic principle, an error in a slice may destroy the slice in the current picture (e.g., X% damaged). When the next frame references this frame, the error may propagate both temporally and spatially, the region may be damaged, and the error may propagate spatially to the next frame, until an intraframe is received or an intraslice for the region is received. This may depend on, for example, the amount of motion vectors in the next frame, unless intra is applied, if more than X% is damaged. In the next frame, the size of the damaged area may increase. Intraprediction may eliminate error propagation, but in the case of multiple reference frames, the damaged area may be returned.
[0141]
[0104] Figure 12 is a flowchart illustrating an example method for processing media data in accordance with the techniques of this disclosure. The method of Figure 12 is described with reference to the system 100 of Figure 1, and in particular, the XR server device 110, by way of example. However, it should be understood that other devices may be configured to perform this or similar methods.
[0142]
[0105] First, the XR server device 110 may receive tracking and sensor information 132 from the XR client device 140 (350). For example, the XR client device 140 may use the tracking / XR sensors 146 to determine the orientation in which the user is looking. The XR client device 140 may send tracking and sensor information 132 representing the orientation in which the user is looking. The XR server device 110 may receive the tracking and sensor information 132 and provide this information to the XR scene generation unit 112. The XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114 may then generate video frames of scene data (352). In some examples, a video game engine may generate the scene data using the tracking and sensor information 132, video game data, and video game configuration data.
[0143] The 2D media encoding unit 116 may then encode the generated video frames of the scene data (354). In some examples, as shown in and described with respect to FIG. 6, such encoding may include intra-prediction, which encodes all of the frames in one run, and both inter-prediction and intra-prediction, which encodes the frames in another run. The intra-prediction encoding may generate vi trace data, and the inter-prediction encoding may generate vp trace data. The XR server device 110 may combine the vi trace data and the vp trace data to form v trace data, which may represent the encoding complexity of the video frames. In some examples, the 2D media encoding unit 116 may encode all frames of a particular v trace using a common, invariant quantization configuration.
[0144] The XR media content delivery unit 118 may then packetize the slices of the encoded video frames (356). Generally, the 2D media encoding unit 116 may be configured with a particular maximum transmission unit (MTU) size for radio access network (RAN) packets and produce slices having an amount of data less than or equal to the MTU size. Thus, in some examples, each slice may be capable of transmission within a single packet.
[0145] The XR server device 110 may then, for example, perform a RAN simulation of the RAN to forward packets using the configuration of the network 130 (358). Examples of performing a RAN simulation in this manner are described above with respect to FIGS. 2-4 and 7-9. For example, when performing the RAN simulation, the XR server device 110 may determine that a slice is lost when a packet for the slice is lost or corrupted. The XR server device 110 may also calculate the latency for forwarding the packets. The XR server device 110 may also generate S trace data and P trace data as part of the RAN simulation, as described above.
[0146] The XR server device 110 may then assemble the simulated received packets into encoded video frames (360) and decode the video frames (362). The XR server device 110 may then calculate individual frame quality (364). The individual frame quality may represent the difference between the video frame before or after encoding and the decoded video frame, for example, as described with respect to FIGS. 2 and 8.
[0147] The XR server device 110 may then determine an overall quality for the system from the individual frame qualities (366). For example, the XR server device 110 may determine the overall quality as the average encoding quality and / or the average number of incorrect video frames (e.g., frames with missing data or corrupted data). The XR server device 110 may perform this method for various different types of configurations, for example, to determine the number of users that can be supported for a given configuration. For example, the XR server device 110 may calculate a percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations. Additionally or alternatively, the XR server device 110 may calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more numbers of supported users.
[0148]
[0111] Thus, the method of Figure 12 represents one example of a method for processing media data, including receiving tracking and sensor information from an extended reality (XR) client device, generating scene data using the tracking and sensor information, encoding the video frames to form encoded video frames, the scene data comprising one or more video frames, performing a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network, decoding the delivered encoded video frames in accordance with the RAN simulation to form decoded video frames, calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames, and determining an overall quality value from the values representing the individual frame quality for each of the video frames.
[0149]
[0112] The following clauses represent some examples of the techniques of this disclosure.
[0150]
[0113] Clause 1: A method for sending media data, the method comprising receiving tracking and sensor information from an extended reality (XR) client device, generating scene data using the tracking and sensor information, rendering an XR viewport to form video data from the scene data, encoding the video data, and sending the encoded video data to the XR client device over a 5G network.
[0151]
[0114] Clause 2: A method for retrieving media data, the method comprising sending tracking and sensor information to an extended reality (XR) server device, receiving encoded video data corresponding to the tracking and sensor information from the XR server device, decoding the encoded video data, and rendering an XR viewport using the decoded video data.
[0152]
[0115] Clause 3: A method comprising a combination of the method of clause 1 and the method of clause 2.
[0153]
[0116] Clause 4: A method for processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information; encoding the video data to form v-trace data, the scene data comprising one or more video frames; performing a radio access network (RAN) simulation of distributing the encoded video data over a 5G network; decoding the distributed encoded video data to form decoded video data; calculating a value representing quality for each of the video frames from the generated one or more video frames and the decoded video data; and determining an overall quality value from the value representing quality for each of the video frames.
[0154]
[0117] Clause 5: The method of clause 4, wherein generating the scene data comprises running a video game engine using tracking and sensor information, video game data, and video game configuration data.
[0155]
[0118] Clause 6: The method of any of clauses 4 and 5, wherein encoding the video data comprises encoding the video data without modifying the quantization configuration.
[0156]
[0119] Clause 7: A method as described in any of clauses 4 to 6, wherein determining the overall quality value comprises generating a graph representing the percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations.
[0157]
[0120] Clause 8: A method as described in any of clauses 4 and 5, wherein determining the overall quality value comprises generating a graph representing the percentage of corrupted video frames as a function of quantization configuration for one or more numbers of supported users.
[0158]
[0121] Clause 9: A method according to any of clauses 4 to 8, wherein encoding the video data to form v trace data comprises encoding at least some portions of the video data using inter prediction to form vp trace data, encoding at least some portions of the video data using intra prediction to form vi trace data, and combining the vp trace data and the vi trace data to form the v trace data.
[0159]
[0122] Clause 10: The method of any of clauses 4 to 9, wherein the v-trace data is formatted according to the FFMPEG passlogfile format.
[0160]
[0123] Clause 11: A method according to any one of clauses 4 to 10, wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture-bits value, an inter-texture-bits value, a motion vector bits value, a miscellaneous bits value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bits value.
[0161]
[0124] Clause 12: A method as described in any of clauses 4 to 11, wherein performing a RAN simulation comprises receiving one or more slices of encoded video data corresponding to the v trace data, packetizing and fragmenting the slices to form packets in the p trace data using Real-time Transport Protocol (RTP) in accordance with IP packet and payload size configuration data, simulating the transfer of the packets, determining that one of the slices is lost when a packet for one of the slices is lost during the simulation of the transfer of the packets, and outputting data for each slice indicating whether the slice was received or lost during the simulation of the transfer and the latency for the transfer of the received packets.
[0162]
[0125] Clause 13: A method according to any of clauses 4 to 12, wherein performing a RAN simulation comprises receiving one or more slices of coded video data corresponding to v trace data, and forming s trace data from the slices in accordance with configuration data indicating a bitrate control technique, a slice setting, an error resilience technique, and a feedback technique.
[0163]
[0126] Clause 14: The method of clause 13, wherein the bitrate control technique comprises one of a fixed quality, a constant bitrate, a feedback-based variable bitrate, or a fixed rate factor.
[0164]
[0127] Clause 15: The method of any of clauses 13 and 14, wherein the slice settings comprise one or more of a number of slices or a maximum slice size.
[0165]
[0128] Clause 16: A method as described in any of clauses 13 to 15, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, normal intra-refresh, feedback-based intra-refresh, or feedback-based prediction.
[0166]
[0129] Clause 17: The method of any of clauses 13 to 16, wherein the feedback technique comprises one of no feedback, statistical feedback, or operational feedback.
[0167]
[0130] Clause 18: A method as described in any of clauses 4 to 17, wherein calculating a value representing quality for each video frame comprises, for each received coded slice, determining the quality of the slice, and for slices that have not been received, determining whether the unreceived slice was lost or of degraded quality.
[0168]
[0131] Clause 19: A method as described in any of clauses 4 to 18, wherein determining the overall quality value comprises determining the overall quality value using one or more of average encoding quality or average inaccurate video data.
[0169]
[0132] Clause 20: A device for processing media data, the device comprising one or more means for performing the method according to any of clauses 1 to 19.
[0170]
[0133] Clause 21: The device of clause 20, wherein the apparatus comprises at least one of an integrated circuit, a microprocessor, and a wireless communication device.
[0171]
[0134] Clause 22: A computer-readable storage medium storing instructions that, when executed, cause a processor to perform a method according to any one of clauses 1 to 19.
[0172]
[0135] Clause 23: A device for sending media data, the device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information; means for rendering an XR viewport to form video data from the scene data; means for encoding the video data; and means for sending the encoded video data to the XR client device via a 5G network.
[0173]
[0136] Clause 24: A device for retrieving media data, the device comprising: means for sending tracking and sensor information to an extended reality (XR) server device; means for receiving encoded video data corresponding to the tracking and sensor information from the XR server device; means for decoding the encoded video data; and means for rendering an XR viewport using the decoded video data.
[0174]
[0137] Clause 25: A device for processing media data, the device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information; means for encoding the video data to form v-trace data, the scene data comprising one or more video frames; means for performing a radio access network (RAN) simulation of distributing the encoded video data over a 5G network; means for decoding the distributed encoded video data to form decoded video data; means for calculating a value representing quality for each of the video frames from the generated one or more video frames and the decoded video data; and means for determining an overall quality value from the value representing quality for each of the video frames.
[0175]
[0138] Clause 26: A method for processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information; encoding the video frames to form encoded video frames, the scene data comprising one or more video frames; performing a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decoding the delivered encoded video frames in accordance with the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the values representing the individual frame quality for each of the video frames.
[0176]
[0139] Clause 27: The method of clause 26, wherein generating the scene data comprises running a video game engine using tracking and sensor information, video game data, and video game configuration data.
[0177]
[0140] Clause 28: The method of clause 26, wherein encoding the video frames comprises encoding the video frames without changing the quantization configuration.
[0178]
[0141] Clause 29: The method described in clause 26, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations.
[0179]
[0142] Clause 30: The method of clause 26, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of quantization configuration for one or more numbers of supported users.
[0180]
[0143] Clause 31: The method of clause 26, wherein encoding the video frame further comprises forming v trace data, the v trace data representing the complexity of encoding the video frame, and forming the v trace data comprises encoding at least some portions of the video frame using inter prediction to form vp trace data, encoding at least some portions of the video frame using intra prediction to form vi trace data, and combining the vp trace data and the vi trace data to form the v trace data.
[0181]
[0144] Clause 32: The method of clause 26, wherein encoding the video frame further comprises forming v trace data, the v trace data representing the complexity of encoding the video frame, wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.
[0182]
[0145] Clause 33: The method described in Clause 26, wherein performing a RAN simulation comprises receiving one or more slices of an encoded video frame, packetizing and fragmenting the slices to form packets in p-trace data using Real-time Transport Protocol (RTP) in accordance with IP packet and payload size configuration data, performing a simulation of packet forwarding, determining that one of the slices is lost when a packet for one of the slices is lost during the simulation of packet forwarding, and outputting, for each slice, data indicating whether the slice was received or lost during the simulation of forwarding and the latency for forwarding the received packets.
[0183]
[0146] Clause 34: The method described in Clause 26, wherein performing a RAN simulation comprises receiving one or more slices of an encoded video frame and forming s-trace data from the slices in accordance with configuration data indicating bit rate control techniques, slice settings, error resilience techniques, and feedback techniques.
[0184]
[0147] Clause 35: The method described in clause 34, wherein forming the trace data comprises assigning to each slice data representing the frame associated with the slice, the size of the slice, the quality of the slice, the area of the frame covered by the slice, and timing information for the slice.
[0185]
[0148] Clause 36: The method of clause 34, wherein the bitrate control technique comprises one of a fixed quality, a constant bitrate, a feedback-based variable bitrate, or a fixed rate factor.
[0186]
[0149] Clause 37: The method of clause 34, wherein the slice settings comprise one or more of a number of slices or a maximum slice size.
[0187]
[0150] Clause 38: The method of clause 34, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, normal intra-refresh, feedback-based intra-refresh, or feedback-based prediction.
[0188]
[0151] Clause 39: The method of clause 34, wherein the feedback technique comprises one of statistical feedback or behavioral feedback.
[0189]
[0152] Clause 40: The method described in Clause 26, wherein calculating a value representing individual frame quality for each of the video frames comprises determining, for each received encoded slice of the video frame, the quality of the slice, and determining, for unreceived slices of the video frame, whether the unreceived slice was lost or of degraded quality.
[0190]
[0153] Clause 41: The method of clause 26, wherein determining the overall quality value comprises determining the overall quality value using one or more of an average encoding quality or an average inaccurate video frames.
[0191]
[0154] Clause 42: The method described in clause 26, wherein the radio access network comprises a 5G network.
[0192]
[0155] Clause 43: The method of clause 26, wherein the RAN simulation comprises a first RAN simulation of a plurality of RAN simulations, wherein the overall video quality comprises a first overall video quality, and wherein performing the first RAN simulation comprises performing the first RAN simulation using a first configuration parameter, and the method further comprises determining a respective configuration parameter for each of the plurality of RAN simulations, performing each of the RAN simulations using the respective configuration parameter, determining an overall video quality value for each of the RAN simulations, and configuring the network to deliver XR data using the configuration parameter corresponding to the best overall video quality value for the RAN simulation.
[0193]
[0156] Clause 44: A device for processing media data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry, the one or more processors configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information; encode the video frames to form encoded video frames, the scene data comprising one or more video frames of video data; perform a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decode the delivered coded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determine an overall quality value from the values representing the individual frame quality for each of the video frames.
[0194]
[0157] Clause 45: A device as described in clause 44, wherein one or more processors are configured to execute a video game engine using tracking and sensor information, video game data, and video game configuration data to generate scene data.
[0195]
[0158] Clause 46: A device as described in Clause 44, wherein one or more processors are configured to calculate the percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations to determine the overall quality value.
[0196]
[0159] Clause 47: A device as described in Clause 44, wherein one or more processors are configured to calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more numbers of supported users to determine the overall quality value.
[0197]
[0160] Clause 48: The device described in Clause 44, wherein the one or more processors are further configured to form v trace data, the v trace data representing the encoding complexity of the video frames, and wherein, to form the v trace data, the one or more processors are configured to encode at least some portions of the video frames using inter prediction to form vp trace data, encode at least some portions of the video frames using intra prediction to form vi trace data, and combine the vp trace data and the vi trace data to form the v trace data.
[0198]
[0161] Clause 49: The device of Clause 44, wherein the one or more processors are further configured to form v trace data, the v trace data representing the encoding complexity of the video frame, wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.
[0199]
[0162] Clause 50: The device described in Clause 44, wherein to perform a RAN simulation, one or more processors are configured to receive one or more slices of an encoded video frame, packetize and fragment the slices to form packets in p-trace data using Real-time Transport Protocol (RTP) in accordance with IP packet and payload size configuration data, perform a simulation of forwarding of the packets, determine that one of the slices is lost when a packet for one of the slices is lost during the simulation of the packet forwarding, and output data for each slice indicating whether the slice was received or lost during the simulation of forwarding and the latency for forwarding the received packets.
[0200]
[0163] Clause 51: A device as described in Clause 44, wherein one or more processors are configured to receive one or more slices of an encoded video frame and form s-trace data from the slices in accordance with configuration data indicating bit rate control techniques, slice settings, error resilience techniques, and feedback techniques, in order to perform a RAN simulation.
[0201]
[0164] Clause 52: A device as described in Clause 51, configured such that, to form trace data, one or more processors assign to each slice data representing a frame associated with the slice, the size of the slice, the quality of the slice, the area of the frame covered by the slice, and timing information for the slice.
[0202]
[0165] Clause 53: A device as described in Clause 44, wherein, to calculate a value representing individual frame quality for each of the video frames, the one or more processors are configured to determine, for each received encoded slice of the video frame, the quality of the slice, and for unreceived slices of the video frame, determine whether the unreceived slice was lost or of degraded quality.
[0203]
[0166] Clause 54: A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information; encode the video frames to form encoded video frames, the scene data comprising one or more video frames; perform a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decode the delivered encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determine an overall quality value from the values representing the individual frame quality for each of the video frames.
[0204]
[0167] Clause 55: A computer-readable storage medium as described in Clause 54, wherein the instructions that cause the processor to generate scene data include instructions that cause the processor to run a video game engine using tracking and sensor information, video game data, and video game configuration data.
[0205]
[0168] Clause 56: A computer-readable storage medium as described in Clause 54, wherein the instructions that cause the processor to determine the overall quality value include instructions that cause the processor to calculate the percentage of corrupted video frames as a function of the number of users supported for one or more sets of configurations.
[0206]
[0169] Clause 57: A computer-readable storage medium as described in Clause 54, wherein the instructions that cause the processor to determine the overall quality value include instructions that cause the processor to calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more numbers of supported users.
[0207]
[0170] Clause 58: A computer-readable storage medium as described in Clause 54, wherein the instructions for causing the processor to encode the video frame further comprise instructions for causing the processor to form v trace data, the v trace data representing the complexity of the encoding of the video frame, and wherein the instructions for causing the processor to form the v trace data include instructions for causing the processor to encode at least some portions of the video frame using inter prediction to form vp trace data, encode at least some portions of the video frame using intra prediction to form vi trace data, and combine the vp trace data and the vi trace data to form the v trace data.
[0208]
[0171] Clause 59: A computer-readable storage medium as described in Clause 54, wherein the instructions causing the processor to encode a video frame further comprise instructions causing the processor to form v-trace data, the v-trace data representing the complexity of encoding the video frame, wherein the v-trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.
[0209]
[0172] Clause 60: A computer-readable storage medium as described in Clause 54, wherein the instructions for causing a processor to perform a RAN simulation include instructions for causing the processor to receive one or more slices of an encoded video frame, packetize and fragment the slices to form packets in p-trace data using Real-time Transport Protocol (RTP) in accordance with IP packet and payload size configuration data, perform a simulation of the forwarding of the packets, determine that one of the slices is lost when a packet for one of the slices is lost during the simulation of the forwarding of the packets, and output data for each slice indicating whether the slice was received or lost during the simulation of the forwarding and the latency for forwarding the received packets.
[0210]
[0173] Clause 61: A computer-readable storage medium as described in Clause 54, wherein the instructions for causing a processor to perform a RAN simulation include instructions for causing the processor to receive one or more slices of an encoded video frame and form s-trace data from the slices in accordance with configuration data indicating bit rate control techniques, slice settings, error resilience techniques, and feedback techniques.
[0211]
[0174] Clause 62: A computer-readable storage medium as described in Clause 61, wherein the instructions that cause the processor to form s trace data include instructions that cause the processor to assign to each slice data representing a frame associated with the slice, a size of the slice, a quality of the slice, an area of the frame covered by the slice, and timing information for the slice.
[0212]
[0175] Clause 63: A computer-readable storage medium as described in Clause 54, wherein the instructions that cause the processor to calculate a value representing individual frame quality for each video frame include instructions that cause the processor to determine, for each received encoded slice of the video frame, the quality of the slice, and for unreceived slices of the video frame, determine whether the unreceived slice was lost or of degraded quality.
[0213]
[0176] Clause 64: A device for processing media data, the device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information; means for encoding video frames to form encoded video frames, the scene data comprising one or more video frames; means for performing a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; means for decoding the encoded video frames delivered in accordance with the RAN simulation to form decoded video frames; means for calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and means for determining an overall quality value from the values representing the individual frame quality for each of the video frames.
[0214]
[0177] Clause 65: A method for processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information; encoding the video frames to form encoded video frames, the scene data comprising one or more video frames; performing a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decoding the delivered encoded video frames in accordance with the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the values representing the individual frame quality for each of the video frames.
[0215]
[0178] Clause 66: The method of clause 65, wherein generating the scene data comprises executing a video game engine using tracking and sensor information, video game data, and video game configuration data.
[0216]
[0179] Clause 67: The method of any of clauses 65 to 66, wherein encoding the video frames comprises encoding the video frames without changing the quantization configuration.
[0217]
[0180] Clause 68: A method as described in any of clauses 65 to 67, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations.
[0218]
[0181] Clause 69: A method as described in any of clauses 65 to 68, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the quantization configuration for one or more numbers of supported users.
[0219]
[0182] Clause 70: A method according to any of clauses 65 to 69, wherein encoding the video frame further comprises forming v trace data, the v trace data representing the complexity of encoding the video frame, and forming the v trace data comprises encoding at least some portions of the video frame using inter prediction to form vp trace data, encoding at least some portions of the video frame using intra prediction to form vi trace data, and combining the vp trace data and the vi trace data to form the v trace data.
[0220]
[0183] Clause 71: A method according to any of clauses 65 to 70, wherein encoding the video frame further comprises forming v trace data, the v trace data representing the complexity of encoding the video frame, wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.
[0221]
[0184] Clause 72: A method as described in any of clauses 65 to 71, wherein performing a RAN simulation comprises receiving one or more slices of an encoded video frame, packetizing and fragmenting the slices to form packets in the p-trace data using Real-time Transport Protocol (RTP) in accordance with IP packet and payload size configuration data, performing a simulation of forwarding of the packets, determining that one of the slices is lost when a packet for one of the slices is lost during the simulation of the forwarding of the packets, and outputting data for each slice indicating whether the slice was received or lost during the simulation of the forwarding and the latency for forwarding the received packets.
[0222]
[0185] Clause 73: A method as described in any of clauses 65 to 72, wherein performing a RAN simulation comprises receiving one or more slices of an encoded video frame and forming s-trace data from the slices in accordance with configuration data indicating bit rate control techniques, slice settings, error resilience techniques, and feedback techniques.
[0223]
[0186] Clause 74: The method described in clause 73, wherein forming the trace data comprises assigning to each slice data representing the frame associated with the slice, the size of the slice, the quality of the slice, the area of the frame covered by the slice, and timing information for the slice.
[0224]
[0187] Clause 75: The method of any of clauses 73 and 74, wherein the bit rate control technique comprises one of a fixed quality, a constant bit rate, a feedback based variable bit rate, or a fixed rate factor.
[0225]
[0188] Clause 76: The method of any of clauses 73 to 75, wherein the slice settings comprise one or more of a number of slices or a maximum slice size.
[0226]
[0189] Clause 77: The method of any of clauses 73 to 76, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, normal intra-refresh, feedback-based intra-refresh, or feedback-based prediction.
[0227]
[0190] Clause 78: The method of any of clauses 73 to 77, wherein the feedback technique comprises one of statistical feedback or behavioral feedback.
[0228]
[0191] Clause 79: A method described in any of clauses 65 to 78, wherein calculating a value representing individual frame quality for each of the video frames comprises determining, for each received coded slice of the video frame, the quality of the slice, and determining, for unreceived slices of the video frame, whether the unreceived slice was lost or of degraded quality.
[0229]
[0192] Clause 80: A method as described in any of clauses 65 to 79, wherein determining the overall quality value comprises determining the overall quality value using one or more of an average encoding quality or an average inaccurate video frames.
[0230]
[0193] Clause 81: A method according to any one of clauses 65 to 80, wherein the radio access network comprises a 5G network.
[0231]
[0194] Clause 82: A method according to any of clauses 65 to 81, wherein the RAN simulation comprises a first RAN simulation of a plurality of RAN simulations, the overall video quality comprises a first overall video quality, and performing the first RAN simulation comprises performing the first RAN simulation using first configuration parameters, and the method further comprises determining respective configuration parameters for each of the plurality of RAN simulations, performing each of the RAN simulations using the respective configuration parameters, determining an overall video quality value for each of the RAN simulations, and configuring the network to deliver XR data using the configuration parameters corresponding to the best overall video quality value for the RAN simulation.
[0232]
[0195] Clause 83: A device for processing media data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry, the one or more processors configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information; encode the video frames to form encoded video frames, the scene data comprising one or more video frames of video data; perform a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decode the delivered coded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determine an overall quality value from the values representing the individual frame quality for each of the video frames.
[0233]
[0196] Clause 84: A device as described in clause 83, wherein one or more processors are configured to execute a video game engine using tracking and sensor information, video game data, and video game configuration data to generate scene data.
[0234]
[0197] Clause 85: A device described in any of clauses 83 and 84, wherein one or more processors are configured to calculate, for one or more sets of configurations, the percentage of corrupted video frames as a function of the number of supported users, in order to determine the overall quality value.
[0235]
[0198] Clause 86: A device described in any of clauses 83 to 85, wherein one or more processors are configured to calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more numbers of supported users to determine the overall quality value.
[0236]
[0199] Clause 87: A device described in any of clauses 83 to 86, wherein the one or more processors are further configured to form v trace data, the v trace data representing the encoding complexity of the video frames, and wherein, to form the v trace data, the one or more processors are configured to encode at least some portions of the video frames using inter prediction to form vp trace data, encode at least some portions of the video frames using intra prediction to form vi trace data, and combine the vp trace data and the vi trace data to form the v trace data.
[0237]
[0200] Clause 88: A device described in any of clauses 83 to 87, wherein the one or more processors are further configured to form v trace data, the v trace data representing the encoding complexity of the video frame, and wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.
[0238]
[0201] Clause 89: A device described in any of clauses 83 to 88, wherein to perform a RAN simulation, one or more processors are configured to receive one or more slices of an encoded video frame, packetize and fragment the slices to form packets in p-trace data using Real-time Transport Protocol (RTP) in accordance with IP packet and payload size configuration data, perform a simulation of forwarding of the packets, determine that one of the slices is lost when a packet for one of the slices is lost during the simulation of the packet forwarding, and output data for each slice indicating whether the slice was received or lost during the simulation of forwarding and the latency for forwarding the received packets.
[0239]
[0202] Clause 90: A device described in any of clauses 83 to 89, configured to perform a RAN simulation, wherein one or more processors receive one or more slices of an encoded video frame and form s-trace data from the slices in accordance with configuration data indicating bit rate control techniques, slice settings, error resilience techniques, and feedback techniques.
[0240]
[0203] Clause 91: A device as described in Clause 90, configured such that, to form trace data, one or more processors assign to each slice data representing a frame associated with the slice, the size of the slice, the quality of the slice, the area of the frame covered by the slice, and timing information for the slice.
[0241]
[0204] Clause 92: A device described in any of clauses 83 to 91, wherein, in order to calculate a value representing individual frame quality for each of the video frames, the one or more processors are configured to determine, for each received coded slice of the video frame, the quality of the slice, and for unreceived slices of the video frame, determine whether the unreceived slice was lost or of degraded quality.
[0242]
[0205] Clause 93: A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information; encode the video frames to form encoded video frames, the scene data comprising one or more video frames; perform a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; decode the delivered encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determine an overall quality value from the values representing the individual frame quality for each of the video frames.
[0243]
[0206] Clause 94: A computer-readable storage medium as described in Clause 93, wherein the instructions that cause the processor to generate scene data include instructions that cause the processor to run a video game engine using tracking and sensor information, video game data, and video game configuration data.
[0244]
[0207] Clause 95: A computer-readable storage medium described in any of clauses 93 and 94, wherein the instructions that cause the processor to determine the overall quality value include instructions that cause the processor to calculate the percentage of corrupted video frames as a function of the number of users supported for one or more sets of configurations.
[0245]
[0208] Clause 96: A computer-readable storage medium described in any of clauses 93 to 95, wherein the instructions that cause the processor to determine the overall quality value include instructions that cause the processor to calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more numbers of supported users.
[0246]
[0209] Clause 97: A computer-readable storage medium according to any of clauses 93 to 96, wherein the instructions causing the processor to encode the video frame further comprise instructions causing the processor to form v trace data, the v trace data representing the complexity of the encoding of the video frame, and wherein the instructions causing the processor to form the v trace data include instructions causing the processor to encode at least some portions of the video frame using inter prediction to form vp trace data, encode at least some portions of the video frame using intra prediction to form vi trace data, and combine the vp trace data and the vi trace data to form the v trace data.
[0247]
[0210] Clause 98: A computer-readable storage medium as described in any of clauses 93 to 97, wherein the instructions causing the processor to encode a video frame further comprise instructions causing the processor to form v trace data, the v trace data representing the complexity of encoding the video frame, wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.
[0248]
[0211] Clause 99: A computer-readable storage medium as described in any of clauses 93 to 98, wherein instructions for causing a processor to perform a RAN simulation include instructions for causing the processor to receive one or more slices of an encoded video frame, packetize and fragment the slices to form packets in p-trace data using Real-time Transport Protocol (RTP) in accordance with IP packet and payload size configuration data, perform a simulation of the forwarding of the packets, determine that one of the slices is lost when a packet for one of the slices is lost during the simulation of the forwarding of the packets, and output data for each slice indicating whether the slice was received or lost during the simulation of the forwarding and the latency for forwarding the received packets.
[0249]
[0212] Clause 100: A computer-readable storage medium described in any of clauses 93 to 99, wherein the instructions for causing a processor to perform a RAN simulation include instructions for causing the processor to receive one or more slices of an encoded video frame and form s-trace data from the slices in accordance with configuration data indicating bit rate control techniques, slice settings, error resilience techniques, and feedback techniques.
[0250]
[0213] Clause 101: A computer-readable storage medium as described in clause 100, wherein the instructions that cause the processor to form s trace data include instructions that cause the processor to assign to each slice data representing a frame associated with the slice, a size of the slice, a quality of the slice, an area of the frame covered by the slice, and timing information for the slice.
[0251]
[0214] Clause 102: A computer-readable storage medium described in any of clauses 93 to 101, wherein the instructions that cause the processor to calculate a value representing individual frame quality for each video frame include instructions that cause the processor to determine, for each received encoded slice of the video frame, the quality of the slice, and for unreceived slices of the video frame, determine whether the unreceived slice was lost or of degraded quality.
[0252]
[0215] Clause 103: A device for processing media data, the device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information; means for encoding video frames to form encoded video frames, the scene data comprising one or more video frames; means for performing a radio access network (RAN) simulation of distributing the encoded video frames over a radio access network; means for decoding the encoded video frames delivered in accordance with the RAN simulation to form decoded video frames; means for calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and means for determining an overall quality value from the values representing the individual frame quality for each of the video frames.
[0253] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or may include communication media, including any medium that enables transfer of a computer program from one place to another, for example, according to a communications protocol. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transitory, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0254]
[0217] By way of example, and not limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory, tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0255]
[0218] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor," as used herein, may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Also, the techniques may be implemented entirely within one or more circuits or logic elements.
[0256] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chipsets). Although various components, modules, or units have been described in this disclosure to highlight functional aspects of devices configured to implement the disclosed techniques, those components, modules, or units do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors described above, along with suitable software and / or firmware.
[0257]
[0220] Various examples have been described. These and other examples are within the scope of the following claims. The inventions described in the claims of the present application as originally filed are set forth below. [C1] A method for processing media data, said method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames over a radio access network; decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the values representing the individual frame quality for each of the video frames. A method comprising: [C2] The method of C1, wherein generating the scene data comprises running a video game engine using the tracking and sensor information, video game data, and video game configuration data. [C3] The method of C1, wherein encoding the video frame comprises encoding the video frame without changing a quantization configuration. [C4] The method of C1, wherein determining the overall quality value comprises calculating a percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations. [C5] The method of C1, wherein determining the overall quality value comprises calculating a percentage of corrupted video frames as a function of quantization configuration for one or more numbers of supported users. [C6] The encoding of the video frames further comprises forming v-trace data, the v-trace data representing the complexity of the encoding of the video frames, and forming the v-trace data comprises: encoding at least some portions of said video frames using inter prediction to form vp trace data; encoding at least some portions of the video frames using intra prediction to form vi trace data; combining the vp trace data and the vi trace data to form the v trace data; The method of C1, comprising: [C7] The method of C1, wherein encoding the video frame further comprises forming v trace data, the v trace data representing the complexity of the encoding of the video frame, wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value. [C8] performing the RAN simulation, receiving one or more slices of the encoded video frame; packetizing and fragmenting the slices to form packets in p-trace data using a Real-time Transport Protocol (RTP) according to IP packet and payload size configuration data; performing a simulation of the forwarding of said packets; determining that one of the slices is lost when a packet for one of the slices is lost during the simulation of the forwarding of the packet; outputting, for each slice, data indicating whether the slice was received or lost during the simulation of the transfer and the latency for transfer of received packets; The method of C1, comprising: [C9] performing the RAN simulation, receiving one or more slices of the encoded video frame; forming s-trace data from the slices in accordance with configuration data indicative of bit rate control techniques, slice settings, error resilience techniques, and feedback techniques; The method of claim C1, comprising: [C10] The method of C9, wherein forming the s-trace data comprises assigning to each of the slices data representing a frame associated with the slice, a size of the slice, a quality of the slice, an area of the frame covered by the slice, and timing information for the slice. [C11] The method of C9, wherein the bit rate control technique comprises one of a fixed quality, a constant bit rate, a feedback based variable bit rate, or a fixed rate factor. [C12] The method of C9, wherein the slice settings comprise one or more of a number of slices or a maximum slice size. [C13] The method of C9, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, normal intra-refresh, feedback-based intra-refresh, or feedback-based prediction. [C14] The method of C9, wherein the feedback technique comprises one of statistical feedback or behavioral feedback. [C15] calculating the value representative of the individual frame quality for each of the video frames, for each received coded slice of the video frame, determining a quality of the slice; For unreceived slices of the video frame, determining whether the unreceived slices were lost or of degraded quality; The method of claim C1, comprising: [C16] The method of C1, wherein determining the overall quality value comprises determining the overall quality value using one or more of an average encoding quality or an average number of inaccurate video frames. [C17] The method of C1, wherein the radio access network comprises a 5G network. [C18] The method further comprises: the RAN simulation comprising a first RAN simulation of a plurality of RAN simulations; the overall video quality comprising a first overall video quality; and performing the first RAN simulation comprising performing the first RAN simulation using first configuration parameters; determining respective configuration parameters for each of the plurality of RAN simulations; performing each of the RAN simulations using the respective configuration parameters; 4. The method of claim 1, further comprising: determining an overall video quality value for each of the RAN simulations; and configuring the network to deliver XR data using the configuration parameters corresponding to the best overall video quality value for the RAN simulations. [C19] A device for processing media data, said device comprising: a memory configured to store video data; and one or more processors implemented in circuitry, the one or more processors comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames of the video data; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames over a radio access network; decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; determining an overall quality value from the values representing the individual frame qualities for each of the video frames; and A device configured to: [C20] The device of C19, wherein the one or more processors are configured to execute a video game engine using the tracking and sensor information, video game data, and video game configuration data to generate the scene data. [C21] The device of C19, wherein, to determine the overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames as a function of a number of supported users for one or more sets of configurations. [C22] The device of C19, wherein, to determine the overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames as a function of quantization configuration for one or more numbers of supported users. [C23] The one or more processors are further configured to form v-trace data, the v-trace data representing the encoding complexity of the video frames, and wherein to form the v-trace data, the one or more processors: encoding at least some portions of said video frames using inter prediction to form vp trace data; encoding at least some portions of the video frames using intra prediction to form vi trace data; combining the vp trace data and the vi trace data to form the v trace data; 20. The device of claim 19, configured to: [C24] The device of C19, wherein the one or more processors are further configured to form v trace data, the v trace data representing the encoding complexity of the video frame, and wherein the v trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f code value, a b code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value. [C25] To perform the RAN simulation, the one or more processors: receive one or more slices of the encoded video frame; packetizing and fragmenting the slices to form packets in p-trace data using a Real-time Transport Protocol (RTP) according to IP packet and payload size configuration data; performing a simulation of the forwarding of said packets; determining that one of the slices is lost when a packet for one of the slices is lost during the simulation of the forwarding of the packet; outputting, for each slice, data indicating whether the slice was received or lost during the simulation of the transfer and the latency for transfer of received packets; 20. The device of claim 19, configured to: [C26] To perform the RAN simulation, the one or more processors: receive one or more slices of the encoded video frame; forming s-trace data from the slices in accordance with configuration data indicative of bit rate control techniques, slice settings, error resilience techniques, and feedback techniques; 20. The device of claim 19, configured to: [C27] The device described in C26, wherein, to form the s trace data, the one or more processors are configured to assign to each of the slices data representing a frame associated with the slice, a size of the slice, a quality of the slice, an area of the frame covered by the slice, and timing information for the slice. [C28] To calculate the value representing the individual frame quality for each of the video frames, the one or more processors: for each received coded slice of the video frame, determining a quality of the slice; For unreceived slices of the video frame, determining whether the unreceived slices were lost or of degraded quality; 20. The device of claim 19, configured to: [C29] A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames over a radio access network; decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the values representing the individual frame quality for each of the video frames. A computer-readable storage medium that causes the [C30] The computer-readable storage medium of C29, wherein the instructions that cause the processor to generate the scene data comprise instructions that cause the processor to run a video game engine using the tracking and sensor information, video game data, and video game configuration data. [C31] A computer-readable storage medium as described in C29, wherein the instructions that cause the processor to determine the overall quality value comprise instructions that cause the processor to calculate a percentage of corrupted video frames as a function of a number of supported users for one or more sets of configurations. [C32] The computer-readable storage medium of C29, wherein the instructions that cause the processor to determine the overall quality value comprise instructions that cause the processor to calculate a percentage of corrupted video frames as a function of quantization configuration for one or more numbers of supported users. [C33] The instructions that cause the processor to encode the video frames further comprise instructions that cause the processor to form v trace data, the v trace data representing the encoding complexity of the video frames, and wherein the instructions that cause the processor to form the v trace data cause the processor to: encode at least some portions of the video frames using inter prediction to form vp trace data; encoding at least some portions of the video frames using intra prediction to form vi trace data; combining the vp trace data and the vi trace data to form the v trace data; 20. The computer-readable storage medium of claim 19, comprising instructions for causing: [C34] The computer-readable storage medium of C29, wherein the instructions that cause the processor to encode the video frame further comprise instructions that cause the processor to form v-trace data, the v-trace data representing the complexity of the encoding of the video frame, and wherein the v-trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value. [C35] The instructions that cause the processor to perform the RAN simulation include: receiving one or more slices of the encoded video frame; packetizing and fragmenting the slices to form packets in p-trace data using a Real-time Transport Protocol (RTP) according to IP packet and payload size configuration data; performing a simulation of the forwarding of said packets; determining that one of the slices is lost when a packet for one of the slices is lost during the simulation of the forwarding of the packet; outputting, for each slice, data indicating whether the slice was received or lost during the simulation of the transfer and the latency for transfer of received packets; 20. The computer-readable storage medium of claim 19, comprising instructions for causing: [C36] The instructions that cause the processor to perform the RAN simulation include: receiving one or more slices of the encoded video frame; forming s-trace data from the slices in accordance with configuration data indicative of bit rate control techniques, slice settings, error resilience techniques, and feedback techniques; 20. The computer-readable storage medium of claim 19, comprising instructions for causing: [C37] A computer-readable storage medium as described in C36, wherein the instructions that cause the processor to form the s-trace data include instructions that cause the processor to assign to each of the slices data representing a frame associated with the slice, a size of the slice, a quality of the slice, an area of the frame covered by the slice, and timing information for the slice. [C38] The instructions that cause the processor to calculate the value representing the individual frame quality for each of the video frames include instructions that cause the processor to: for each received coded slice of the video frame, determining a quality of the slice; For unreceived slices of the video frame, determining whether the unreceived slices were lost or of degraded quality; 20. The computer-readable storage medium of claim 19, comprising instructions for causing: [C39] A device for processing media data, said device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using said tracking and sensor information, said scene data comprising one or more video frames; means for encoding the video frames to form encoded video frames; means for performing a radio access network (RAN) simulation of delivering the encoded video frames over a radio access network; means for decoding the encoded video frames delivered in accordance with the RAN simulation to form decoded video frames; means for calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; means for determining an overall quality value from said values representing said individual frame qualities for each of said video frames; A device comprising:
Claims
1. 1. A method for processing media data, said method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames over a radio access network; and wherein performing the RAN simulation includes: receiving one or more slices of the encoded video frame; forming trace data representing the slice of video data after wireless transmission from the slice in accordance with configuration data indicative of bit rate control techniques, slice settings, error resilience techniques, and feedback techniques. decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; determining an overall quality value from the values representing the individual frame qualities for each of the video frames; and A method comprising:
2. The method of claim 1 , wherein generating the scene data comprises running a video game engine using the tracking and sensor information, video game data, and video game configuration data.
3. The method of claim 1 , wherein encoding the video frames comprises encoding the video frames without changing a quantization configuration.
4. 2. The method of claim 1 , wherein determining the overall quality value comprises either calculating a percentage of corrupted video frames as a function of a number of supported users for one or more sets of configurations, or calculating a percentage of corrupted video frames as a function of a quantization configuration for one or more numbers of supported users.
5. Encoding the video frames further comprises forming v-trace data, the v-trace data representing the encoding complexity of the video frames, and forming the v-trace data comprises: encoding at least some portions of the video frames using inter prediction to form vp trace data; encoding the at least some portions of the video frames using intra prediction to form vi trace data; combining the vp trace data and the vi trace data to form the v trace data; The method of claim 1 , comprising:
6. 2. The method of claim 1, wherein encoding the video frames further comprises forming v-trace data, the v-trace data representing the complexity of the encoding of the video frames, wherein the v-trace data includes one or more of a display picture number value, a coded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.
7. conducting the RAN simulation receiving one or more slices of the encoded video frame; packetizing and fragmenting the slices to form packets using a Real-time Transport Protocol (RTP) according to IP packet and payload size configuration data; performing a simulation of the forwarding of said packets; determining that one of the slices is lost when a packet for one of the slices is lost during the simulation of the forwarding of the packet; outputting, for each slice, data indicating whether the slice was received or lost during the simulation of the transfer and the latency for transfer of received packets; The method of claim 1 , comprising:
8. forming the trace data comprises assigning to each of the slices data representing a frame associated with the slice, a size of the slice, a quality of the slice, an area of the frame covered by the slice, and timing information for the slice; or the bitrate control technique comprises one of a fixed quality, a fixed bitrate, a feedback-based variable bitrate, or a fixed rate factor; or the slice settings comprise one or more of a number of slices or a maximum slice size; or the error resilience technique comprises one of frame-based resilience, slice-based resilience, normal intra-refresh, feedback-based intra-refresh, or feedback-based prediction; or the feedback technique comprises one of statistical feedback or behavioral feedback; The method of claim 1 , wherein the first and second electrodes are one or more of:
9. calculating the value representing the individual frame quality for each of the video frames; for each received coded slice of the video frame, determining a quality of the slice; For unreceived slices of the video frame, determining whether the unreceived slices were lost or of degraded quality; The method of claim 1 , comprising:
10. The method of claim 1 , wherein determining the overall quality value comprises determining the overall quality value using one or more of an average encoding quality or an average number of inaccurate video frames.
11. 10. The method of claim 1, wherein the radio access network comprises a 5G network.
12. the RAN simulation comprises a first RAN simulation of a plurality of RAN simulations, the overall video quality comprises a first overall video quality, and performing the first RAN simulation comprises performing the first RAN simulation using first configuration parameters, and the method further comprises: determining respective configuration parameters for each of the plurality of RAN simulations; performing each of the RAN simulations using the respective configuration parameters; determining an overall video quality value for each of the RAN simulations; configuring the network to deliver XR data using the configuration parameters corresponding to the best overall video quality value for the RAN simulation; The method of claim 1 further comprising:
13. A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames over a radio access network; and wherein performing the RAN simulation includes: receiving one or more slices of the encoded video frame; forming trace data representing the slice of video data after wireless transmission from the slice in accordance with configuration data indicative of bit rate control techniques, slice settings, error resilience techniques, and feedback techniques. decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; determining an overall quality value from the values representing the individual frame qualities for each of the video frames; and A computer-readable storage medium that causes the
14. 1. A device for processing media data, said device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using said tracking and sensor information, said scene data comprising one or more video frames; means for encoding the video frames to form encoded video frames; means for performing a radio access network (RAN) simulation of delivering the encoded video frames over a radio access network, wherein performing the RAN simulation includes: receiving one or more slices of the encoded video frame; forming trace data representing the slice of video data after wireless transmission from the slice in accordance with configuration data indicative of bit rate control techniques, slice settings, error resilience techniques, and feedback techniques. means for decoding the encoded video frames delivered in accordance with the RAN simulation to form decoded video frames; means for calculating a value representing an individual frame quality for each of the video frames from the generated one or more video frames and the decoded video frames; means for determining an overall quality value from said values representing said individual frame qualities for each of said video frames; A device comprising:
15. 15. The device according to claim 14, wherein the device comprises means for carrying out the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Dynamic image communication management device
JP1999069325A
Method and device for decoding compressed multi-media communication and evaluating quality of multi-media communication
JP2001025014A
Real-time display device for multimedia transmission quality
JP2006262453A
A concept for determining the quality of media data streams with varying quality versus bitrate
JP2016530751A
Metrics and messages to improve experience for 360-degree adaptive streaming
US20200037029A1