IMPLEMENTING AND EVALUATING SEPARATE RENDERING ON 5G NETWORKS

IDP000106449BActive Publication Date: 2026-07-14QUALCOMM INC

Patent Information

Authority / Receiving Office
ID · ID
Patent Type
Patents
Current Assignee / Owner
QUALCOMM INC
Filing Date
2021-05-18
Publication Date
2026-07-14

Smart Images

  • Figure 0_ABST
    Figure 0_ABST
Patent Text Reader

Abstract

An example device includes a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing individual frame quality for each video frame from the one or more generated video frames and the decoded video frames; and determine an overall quality value from the values ​​representing individual frame quality for each video frame.
Need to check novelty before this filing date? Find Prior Art

Description

Description IMPLEMENTING AND EVALUATING SEPARATE RENDERING ON THE NETWORK 5G This application claims priority to U.S. Application No. 17 / 322,468 filed May 17, 2021, and U.S. Provisional Application No. 63 / 026,498 , filed May 18, 2020, the entire contents of each of which are hereby incorporated by reference. U.S. Application No. 17 / 322,468 claims the benefit of U.S. Provisional Application No. 63 / 026,498 , filed May 18, 2020. Invention Engineering Field This disclosure relates to the storage and transport of encoded media data. Background of the Invention Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, mobile or satellite radio telephones, video teleconferencing devices, and the like. Digital video devices employ video compression techniques, as described in standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 (also referred to as High Efficiency Video Coding (HEVC)), and extensions to those standards, to more efficiently transmit and receive digital video information. Once video and other media data are encoded, they can be packaged for transmission or storage. Media data can be assembled into video files that conform to one of several standards, such as the International Organization for Standardization (ISO) base media file format and its extensions. Brief Description of the Invention In general, the present disclosure describes techniques related to evaluating the configuration of various components involved in performing discrete rendering for interactive media, such as an online cloud video game. These techniques may be implemented by a computing device or a system of computing devices. These techniques include an evaluation framework for evaluating the efficacy of these devices and the configuration of these devices. For example, different parts of the system may be evaluated in different ways to determine different configuration settings. Specifically, these techniques include evaluation at the video level, the slice level, and the packet level. Evaluation at the packet level may be used to determine whether the corresponding slice has been received correctly. Evaluation at the slice level may be used to determine the quality of the slice and the corresponding frame including the slice.Evaluation at the video level can be used to determine whether the overall set of configuration parameters for all system components is adequate. In one example, a method of processing media data includes receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; simulating a radio access network (RAN) to transmit the encoded video frames over the radio access network; decoding the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculating a value representing individual frame quality for each video frame from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the value representing individual frame quality for each video frame. In another example, a device for processing media data includes a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing individual frame quality for each video frame of the one or more generated video frames and the decoded video frames;and determine the overall quality value from the values ​​representing the individual frame quality for each video frame.; In another example, a computer-readable storage medium has stored thereon instructions that, when executed, cause a processor to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation for transmitting the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing individual frame quality for each video frame from the one or more generated video frames and the decoded video frames; and determine an overall quality value from the values ​​representing individual frame quality for each video frame. Details of one or more examples are set out in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and in the claims. Short Description of Image Figure 1 is a block diagram illustrating an example of a computing system that can perform this disclosure technique. Figure 2 is a conceptual diagram illustrating a detailed example for simulating and measuring the performance of a separate rendering process performed by the computing system of Figure 1 in accordance with the techniques of the present disclosure. Figures 3-5 are example graphs for various radio access network (RAN) simulations as discussed in connection with Figure 2. Figure 6 is a block diagram illustrating an example of a v-trace system model in accordance with the techniques of the present disclosure. Figure 7 is a block diagram illustrating an example system for performing a RAN simulation according to the techniques of the present disclosure. Figure 8 is a block diagram illustrating an example system for measuring the performance of an XR content delivery configuration and modeling according to the techniques of the present disclosure. Figure 9 is a block diagram illustrating an example pipeline for real-time generation of v-trace data according to the techniques of the present disclosure. Figure 10 is a block diagram illustrating an example of a decoding and encoding system according to the techniques of this disclosure. Figure 11 is a graph illustrating an example comparison between average PSRN and PSNR in decibels (dB) for various encoder parameters. Figure 12 is a flowchart illustrating an example of a method of processing media data according to the techniques of this disclosure. Complete Description of the Invention Figure 1 is a block diagram illustrating an example computing system 100 that can perform the techniques of the present disclosure. In this example, the computing system 100 includes an extended reality (XR) server device 110, a network 130, an XR client device 140, and a display device 152. The XR server device 110 includes an XR scene generation unit 112, an XR display port pre-rendering rasterization unit 114, a 2D media encoding unit 116, an XR media content delivery unit 118, and a 5G system (5GS) delivery unit 120. The network 130 may correspond to any network of computing devices that communicate according to one or more network protocols, such as the Internet. In particular, the network 130 may include a 5G radio access network (RAN) including access devices connected by XR client devices 140 to the access network 130 and XR server devices 110. In other examples, other types of networks, such as other types of RANs, may be used.The XR client device 140 includes a 5GS delivery unit 150, a tracking / XR sensor 146, an XR display port rendering unit 142, a 2D media decoder 144, and an XR media content delivery unit 148. The XR client device 140 also interacts with a display device 152 to present XR media data to a user (not shown). In some instances, the XR scene generation unit 112 may be suitable for interactive media entertainment applications, such as video games, that may be executed by one or more processors implemented in the circuitry of the XR server device 110. The XR viewport pre-rendering rasterization unit 114 may format scene data generated by the XR scene generation unit 112 as pre-rendered two-dimensional (2D) media data (e.g., video data) for the viewport user of the XR client device 140. The 2D media encoding unit 116 may encode the formatted scene data from the XR viewport pre-rendering rasterization unit 114, for example, using a video encoding standard, such as ITU-T H.264 / Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 Versatile Video Coding (VVC), or the like. The XR media content delivery unit 118 represents the content delivery transmitter, in this example.In this example, the XR media content delivery unit 148 represents the content delivery receiver, and the 2D media decoder 144 may perform error handling. In general, the XR client device 140 may determine the user's viewport, for example, the direction in which the user is looking and the user's physical location, which may correspond to the orientation of the XR client device 140 and the geographic position of the XR client device 140. The tracking / XR sensor 146 may determine such location and orientation data, for example, using a camera, accelerometer, magnetometer, gyroscope, or the like. The tracking / XR sensor 146 provides location and orientation data to the XR viewport rendering unit 142 and the 5GS dispatch unit 150. The XR client device 140 provides tracking and sensor information 132 to the XR server device 110 via the network 130. The XR server device 110, in turn, receives tracking and sensor information 132 and provides this information to the XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114.In this manner, the XR scene generation unit 112 may generate scene data for the user's viewport and location, and then pre-render the 2D media data for the user's viewport using the XR viewport pre-rendering rasterization unit 114. Therefore, the XR server device 110 may transmit the encoded and pre-rendered 2D media data 134 to the XR client device 140 over the network 130, for example, using a 5G radio configuration. The XR scene generation unit 112 may receive data representing a type of multimedia application (e.g., a type of video game), an application state, some user actions, or the like. The XR display port pre-rendering rasterization unit 114 may format a raster video signal. The 2D media encoding unit 116 may be configured with a particular encoder / decoder (codec), a bitrate for media encoding, a rate control algorithm and corresponding parameters, data for forming image slices of video data, low latency encoding parameters, error robustness parameters, intra-prediction parameters, or the like. The XR media content delivery unit 118 may be configured with real-time transport protocol (RTP) parameters, rate control parameters, error robustness information, and the like. The XR media content delivery unit 148 may be configured with feedback parameters, error masking algorithms and parameters, post-correction algorithms and parameters, and the like. Raster-based split rendering refers to the case where the XR server device 110 runs an XR engine (e.g., an XR scene generation unit 112) to generate an XR scene based on information derived from an XR device, e.g., an XR client device 140 and tracking and sensor information 132. The XR server device 110 may rasterize the XR viewport and perform XR pre-rendering using the viewport pre-rendering rasterization unit. XR 114. In the example of Figure 1, the rendering port is primarily displayed on the XR server device 110, but the XR client device 140 is capable of performing late-stage pose correction, for example, using asynchronous time warping or other XR pose correction to address pose changes. The XR graphics workload can be divided into a rendering workload on the powerful XR server device 110 (in the cloud or edge) and pose correction (such as asynchronous time warp (ATW)) on the XR client device 140. Low motion-to-photon latency is maintained through on-device Asynchronous Time Warping (ATW) or other pose correction methods performed by the XR client device 140. In some examples, the latency of the rendering of video data by the XR server device 110 and the XR client device 140 receiving such pre-rendered video data may be in the range of 50 milliseconds (ms). The latency for the XR client device 140 to provide location and position information (e.g., pose) may be lower, e.g., 20 ms, but the XR server device 110 may perform an asynchronous time warp to compensate for the latest pose at the XR client device 140. The following call flow is an example that highlights the steps to perform this technique: 1) The XR client device 140 connects to the network 130 and joins the XR application (e.g., run by the XR scene generation unit 112). a) The XR 140 client device transmits static device information and capabilities (supported decoders, broadcast ports). 2) Based on this information, the XR server device 110 prepares the encoder and format. 3) Loop: a) The XR client device 140 collects the XR pose (or predicted XR pose) using the tracking / XR sensor 146. b) The XR client device 140 sends XR pose information, in the form of tracking and sensor information 132, to the XR server device 110. c) The XR server device 110 uses tracking information and sensors 132 to pre-render the XR viewport via the XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114. d) The 2D media encoding unit 116 encodes the XR display port. e) the XR media content delivery unit 118 and the 5GS delivery unit 120 send compressed media to the XR client device 140, along with data representing the rendered XR poses for the display port. f) The XR client device 140 decompresses the video data using the 2D media decoder 144. g) The XR client device 140 uses the XR pose data provided with the video frames and the actual XR pose from the XR tracking / sensor 146 for better prediction and to correct the local pose, for example, using ATW performed by the XR viewport rendering unit 142. According to TR 26.928, clause 4.2.2, the relevant processing and delay components are summarized as follows: • User interaction delay is defined as the time duration between when a user action is initiated and when that action is calculated by the content generation engine. In the context of games, this is the time between when a user interacts with the game and when the game engine processes that player's response. • Content lifetime is defined as the time between when the content is created and when it is served to the user. In the context of games, this is the time between the creation of a video frame by the game engine and the time the frame is finally served to the player. Therefore, the round-trip interaction delay is the sum of the Content Age and the User Interaction Delay. If part of the rendering is performed on the XR server and the service generates a frame buffer as a result of rendering the content state, then for raster-based split rendering in cloud gaming applications, the following processes contribute to the delay: • User Interaction Delay (Pose and other interactions) o capturing user interactions in the game client, o sending user interactions to the game engine, i.e. to the server (aka network delay), o processing of user interactions by the game engine / server, • Content Age o creation of one or more video buffers (e.g., one for each eye) by the game engine / server, o encoding of video buffers into video stream frames, o sending video frames to the game client (aka network delay), o decoding of video frames by the game client, o presentation of video frames to the user (aka framerate delay). Since the XR 140 client device implements ATW, the motion-to-photon latency requirement (maximum 20 ms) is met by the XR 140 client device's internal processing. What determines the network requirements for separate rendering are the pose-to-render-to-photon time and the round-trip interaction delay. According to TR 26.928, clause 4.5, the permissible downlink latency is typically 50-60 ms. The rasterized 3D scene available in the frame buffer (see clause 4.4 of TR 26.928) is provided by the XR 112 scene generation unit and needs to be encoded, distributed, and decoded. According to TR 26.928, clause 4.2.1, the relevant format for the frame buffer is 2k by 2k per eye, potentially even higher. The frame rate is expected to be at least 60 fps, potentially up to 90 fps. The frame buffer format is a regular texture video signal that is then directly rendered. Since the processing is graphics-centric, formats beyond the commonly used 4:2:0 signal and YUV signals can be considered. For practical purposes, NVIDIA Encoding functions can be used. The parameters of such encoders are documented at developer.nvidia.com / nvidia-videocodec-sdk. The techniques disclosed in this section can be used to address specific challenges and achieve specific objectives. For example, these techniques can be used to evaluate basic system design options and their performance, generate traffic models for evaluating radio access network (RAN) options, provide guidelines for optimal parameter settings for coding, content delivery, and RAN configuration, identify capacity for specific applications, and determine potential optimizations. These techniques can simulate these various elements with reasonable settings. In the first example, using split rendering, a full simulation for system 100 can be performed. In this example, 5G New Radio (NR) setup and simulation can be performed for communication over network 130, such as, for example, tracking and sensor information 132 and pre-rendered 2D media data 134. Tracking / XR sensors 146 can track and sense user motion examples. The quality of the video data presented to the display device 152 can then be measured during these various simulations. To perform this first example, simulations can be performed using multiple models performing separate tasks. The source video model can include actions performed by the XR scene generation unit 112, the XR viewport pre-rendering rasterization unit 114, the XR viewport rendering unit 142, the 5GS delivery unit 120, and the display device 152. The content delivery model can include actions performed by the 2D media encoding unit 116, the XR media content delivery unit 118, the 5GS delivery unit 120, the 5GS delivery unit 150, and the 2D media decoder 144. The uplink model can include actions performed by the tracking / XR sensors 146 and the 5GS delivery unit 150. Tracking and sensor information 132 can be generated as part of the uplink traffic model, and RAN simulations can be performed to generate pre-rendered 2D media data examples 134. The various components of the XR server device 110, the XR client device 140, and the display device 152 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions associated with these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it is understood that the instructions for the software or firmware may be stored on computer-readable media and executed by the required hardware. Based on the system design in S4-200771, several challenges to potential simulation and to generating traffic models exist because some aspects can be addressed, for example, using certain techniques of the present disclosure. These challenges include: • Evaluate basic system design options and their performance. • Generating Traffic Models for evaluation of RAN options. • Does not violate any content license terms. • Provides guidelines for proper parameter setting in coding, content delivery and also for RAN configuration. • Identify the capacity for that type of application. • Define potential optimization. • Simulate this in a reasonable setting in a repeatable manner. Simulation and modeling per system 100 can be broken down into separate individual components: • Source video model • Content encoding and delivery • Radio access delivery • Content delivery receiver and decoder • Display process • Uplink traffic Figure 2 is a conceptual diagram illustrating an example of a breakdown of the components of an XR server device 110 to simulate and measure the performance of a separate rendering process performed by the computing system 100 of Figure 1 in accordance with the techniques of the present disclosure. The breakdown example of Figure 2 corresponds to the first example discussed above. In this example, the system 160 includes a game engine 162, a model encoding device 164, a content encoding and delivery model 166, a 5GS simulation unit 168, and a content decoding and delivery model 170.Thus, the game engine 162 may correspond to the XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114 of FIG. 1, the model encoding device 164 may correspond to the 2D media encoding unit 116 of FIG. 1, the content encoding and delivery model 166 may correspond to the XR media content delivery unit 118 and the 5GS delivery unit 120 of FIG. 1, the 5GS simulation unit 168 may correspond to the network 130 of FIG. 1, and the content decoding and delivery model 170 may correspond to the 5GS delivery unit 150, the XR media content delivery unit 148, the 2D media decoder 144, and the XR viewport rendering unit 142 of FIG. 1. The game engine 162 receives data 178, including pose model and trace data, game data, and game / machine configuration data, to simulate an extended reality (XR) game using the received pose model and trace data. The game engine 162 generates video frames 180 from this received data and provides the video frames 180 to the model encoding device 164. The model encoding device 164 encodes the video frames 180 to generate encoded video frames and v-trace data and provides the encoded video frames and v-trace data 182 to the content encoding and delivery model 166. The content encoding and delivery model uses encoded video frames and v-trace data 182 to generate slice trace (s-trace) data 184 and packets representing encapsulated and encoded slices of encoded video frames using example configuration parameters 188. The content encoding and delivery model 166 provides the s-trace data 184 and packets to the 5GS simulation unit 168. The 5GS simulation unit 168 performs a RAN simulation using the received s-trace data 184 and packets and configuration parameters 190 to generate the s'-trace data 186. The s-trace data 184 represents video data slices before radio transmission, while the s'-trace data 186 represents video data slices after radio transmission (which may, for example, be corrupted due to radio transmission). The 5GS simulation unit 168 provides the s'-trace data 186 and sent packets encapsulating encoded video data slices to the content decoding and delivery model 170.The content decoding and delivery model 170 performs an example of the decoding process on packet data. Using the encoded and decoded data from the content encoding and delivery model 166 and the content decoding and delivery model 170, values ​​representing the per-frame quality 172 can be generated, to test the performance of the RAN simulation performed by the 5GS simulation unit 168 and the configuration parameters 190. Using the combination / aggregation of the individual values ​​of per-frame quality 172, an overall quality value 174 can be generated. Goals can be set for the values ​​of per-frame quality 172 and overall quality 174 to adjust various configuration parameters. In one example, the goal may be to have no change in video quality as measured by per-frame quality 172 and overall quality 174. One or a small number of application metrics can be derived from RAN simulations performed by the 5GS simulation unit 168. For each scenario, multiple v-traces (e.g., one to three) can be performed but over several hours. The quantization parameters may not change in configuration parameter 188. Some variations in the values ​​for configuration parameter 188 may be adjusted, which may cause Bitrate changes as a side effect and changes in traffic characteristics. However, no quality changes should be made. There may be a small number of configurations tested, e.g., three to ten. Game engine 162, model encoding device 164, content encoding and delivery model 166, simulation unit 5GS 168, and the content decoding and delivery model 170 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions associated with these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it is understood that the instructions for the software or firmware may be stored on computer-readable media and executed by the required hardware. Based on the concepts discussed above, several parameters are relevant for the overall system design: • Games: o Game type o game state, o multi-user actions, and so on. • User interaction: o 6DOF poses based on head and body movements, o Game interaction by controller • Raster video signal format. Typical parameters are: o 1.5K x 1.5K per eye at 60, 90, 120fps o 2K x 2K at 60, 90, 120fps o YUV 4:2:0 or 4:4:4 Encoder configuration o Codec: H.264 / AVC or H.265 / HEVC Bitrate: Bitrate setting to a specific value (e.g., 50 Mbit / s) o Rate control: CBR, Capped VBR, Feedback based, CRF, QP o Slice settings: 1 per frame, 1 per MB line, X per frame o Intra-frame and error robustness settings: IDR Regular, GDR Pattern, Intra adaptive, Intra feedback-based, feedback-based and ACK-based prediction, feedback-based and NACK-based prediction o Latency settings: P image only, look-ahead unit o Complexity settings for encoder • Content Delivery o Slice to IP mapping: Fragmentation o RTP-based timecode and packet numbering o RTP / RTCP ACK / NACK-based feedback o RTP / RTCP Bitrate-based feedback • 5G System / RAN Configuration: o QoS Settings (5QI): GBR, Latency, Loss Rate o HARQ Transmission, scheduling, and so on. • Content Delivery Recipient Configuration: o Loss Detection: sequence number o Latency Delay / handling o Error Resilience o ATW • Quality Aspects o Video quality (encoded) o Data loss o Depth Video quality can be affected by various factors. One factor is encoding artifacts based on the encoding. These artifacts can be determined, for example, by the peak signal-to-noise ratio (PSNR). Another factor is artifacts due to packet loss and the resulting error propagation. According to this disclosure technique, quality can be modeled for each simulation as a combination of: • The average PSNR delivered from the encoder, and • The percentage of corrupted video area o where corrupted macroblocks are defined as Ξ If it is part of the missing slice for this transmission Ξ If received correctly, but predict from wrong macroblock o Macroblock (or coding tree unit) is correct □ If received correctly and predict from uncorrupted macroblock □ Predict from uncorrupted macroblock way • Spatial prediction is correct • Temporal prediction is correct o Recovering MB is done by Intra refresh or by predicting from correct MB again. In general, references to macroblocks above and throughout this disclosure refer to macroblocks of ITU-T H.264 / AVC. However, it should be understood that the concept of macroblock may be replaced with coding unit (CU) or coding tree unit (CTU) of ITU-T H.265 / HEVC or ITU-T H.266 / VVC, without loss of generality. In the example of Figure 2, a source model can be provided based on inputs describing the game under test, a pose model, and a statistical model of the video signal that can be used to implement content delivery and encoding procedures. Based on these traces, 3GPP SA4 can define content delivery and encoding models taking into account the system design based on TR26.928 and as documented in S4-200771. The system model can include encoding, delivery, decoding, and quality definitions. The system can also provide an interface for RAN simulation. The RAN group can then use these models to simulate different traffic characteristics and evaluate different performance options. The following elements may be used by system 160 and other systems described herein: 1) V-trace: A video trace that provides enough information from the encoding model to understand the complexity and data rate for the encoding model, also considering delivery options such as Bitrate control, intra refresh, feedback-based error resilience and so on. Details are tbd, but we are currently investigating the parameters provided by the encoding model from an x265 encoding run. 2) S-trace: a sequence of slices as output from the encoder model. Each slice has been assigned: a. Related Frame (Timestamp) b. Slice Size c. Slice Quality (maybe some more information from the v-trace, e.g. complexity) d. Number of Macroblocks (area covered) e. Time slices, for example, deadlines or at least origin timestamps. 3) S'Trace: a. All information from the trace-s b. Slice loss indicator c. Slice delay information Figure 3 is an example graph for various RAN simulations as discussed in connection with Figure 2. In this example, the graph represents the percentage of corrupted video as a function of the number of supported users. The graph of Figure 3 represents an example graph that might result from the above example where the quantization parameters are unchanged, to prevent changes in video quality. To allow for any number of supported users, various configuration parameters 190 can be modified. As the number of supported users increases, the amount of corrupted video data is expected to increase as well. However, one set of configuration data can be identified that allows for a large number of supported users and a relatively small percentage of corrupted video, such as configuration 3 as shown in Figure 3. Referring again to Figure 2, in another example, the video quality may be allowed to change, as measured by the per-frame quality 172 and the overall quality 174. In this example, one or a small number of application metrics may be derived from the RAN simulation performed by the 5GS simulation unit 168. For each scenario, multiple v-traces (e.g., one to three) may be performed but over a period of several hours. Several different values ​​for the configuration parameters 188 may be adjusted, including the quantization parameters, which may cause Bitrate changes as a side effect and changes in traffic characteristics. However, no quality changes should be made. There may be a small number of configurations tested, e.g., three to ten. Figures 4 and 5 are example graphs for various RAN simulations as discussed in relation to Figure 2. These example graphs correspond to the example above where the quantization parameters are allowed to change. In this example, the graph plots transmit quantization data / video quality against the percentage of corrupted video for a fixed number of users in various scenarios. As shown in Figure 5, a number of users and configuration parameters can be selected that achieve high video quality, low bitrate, and low percentage of data loss, for example, at 0.1% corrupted video. To perform the various test scenarios discussed above, representative source data can be generated to perform v-traces. S-traces can also be performed using content, delivery, and quality coding, including, for example, various system designs, delivery parameters, mapping to radio, modeling, and quality definitions. RAN delivery simulations may be based on the aforementioned S-traces. Figure 6 is a block diagram illustrating an example of a v-trace system model 200 according to the techniques of the present disclosure. The v-trace system model 200 represents a model for evaluating the complexity of source video encoding. In this example, the v-trace system model 200 includes a game engine 202, a predictive model encoding unit 204, an intra-model encoding unit 206, and a trace combination unit 208. The v-trace system model 200 may generally correspond to a portion of the system 160 of Figure 2. For example, the game engine 202 may generally correspond to the game engine 162 of Figure 2, and the predictive model encoding unit 204, the intra-model encoding unit 206, and the trace combination unit 208 may correspond to the model encoding device 164. The game engine 202 receives input data 178 including model pose and trace data, game data, and game configuration data. The game engine 202 generates video frames 210 and provides video frames 210 to each of the predictive model encoding unit 204 and the intra model encoding unit 206. The video frames 210 may, for example, be extended reality (XR) split data at 60 fps, with one frame per eye to achieve the XR effect. In this example, the v-trace system model 200 includes two tracks: one for intra-prediction and one for inter-prediction. In video coding, intra-prediction is performed by predicting a block of video data using a previously decoded neighboring block from the same frame, and inter-prediction is performed by predicting a block of video data using a reference block from a previously decoded frame. The predictive model encoding unit 204 may generate vp-trace data 212, while the intra model encoding unit 206 may generate VI trace data 214. The trace combination unit 208 receives the vp-trace data 212 and the VI trace data 214 and generates V trace data 216. The V-trace data 216 may be formatted according to FFMPEG's -pass log file format after passing through the encoding. Such formats are described in, for example, ffmpeg.org / wiki / Encode / H.264 and slhck.info / video / 2017 / 03 / 01 / rate-control.html. The trace combination unit 208 may also use the -pass[:stream_specifier] n (output,perstream) and -passlogfile configuration parameters. Pseudocode for an example algorithm for generating v-trace data 216 in this format is shown below: void ff_write_pass1_stats ( MpegEncContext *s) { snprintf (s-> avctx -> stats_out, 256, in:% d out:%d type:%dq:%d itex :%d ptex :%d mv:%d other :%d fcode :%d bcode :%d mc-var:%PRId64 var:%PRId64 icount :%d skipcount :%d hbits :%d;\n, s-> current_picture_ptr ->f-> display_picture_number , s-> current_picture_ptr ->f-> coded_picture_number , s->pict_type, s-> current_ picture.f ->qality, s->i_tex_bits, s->p_tex_bits, s->mv_bits, s->misc_bits, s->f_code, s->b_code, s-> current_picture.mc_mb_var_sum, s-> current_ picture.mb _var_sum, s-> i_count, s-> skip_count, s->header_bits); } As an alternative to ITU-T H.264 / AVC, video data can be encoded using ITU-T H.265 / HEVC or other video coding standards. For the H.265 example, a sample log file is described at x265.readthedocs.io / en / default / cli.html#input-output-fileoptions. The log file can include the following parameters: • Encoding Sequence: The sequence of frames that the encoder encodes. • Type: The type of frame slice. • POC: Number of Frame Orders - The order in which frames are displayed. • QP: Quantization Parameters decided for the frame. • Bits: The number of bits consumed by the frame. • Scenecut: 1 if the frame is a scenecut, 0 otherwise. • RateFactor: Only applies if CRF is enabled. The rate factor depends on the CRF provided by the user. It is used to determine the QP to target a specific quality. • BufferFill: Bits available for the next frame. Including bits carried over from the current frame. • BufferFillFinal: Buffer bits available after removing the frame from the CPB. • Latency in terms of the number of frames between when a frame is rendered and when the frame is delivered. • PSNR: Peak signal to noise ratio for the Y, U, and V planes. • SSIM: A quality metric that indicates the structural similarity between frames. • Ref list: POC reference images in list 0 and 1 for frame. • Some statistics about the encoded bitstream and encoder performance are available when --csv-loglevel is greater than or equal to 2: In other examples, other video codecs, such as ITU-T H.266 / Versatile Video Coding (VVC), may be used. Similar datasets can be generated for analysis using VVC or other similar video codecs. The trace combination unit 208 may generate v-trace data 216 according to certain v-trace generation principles. The particular v-trace generation principles are described at slhck.info / video / 2017 / 02 / 24 / crf-guide.html. The trace combination unit 208 may use H.264 / AVC with x264 in FFMPEG and H.265 / HEVC with x265 in FFMPEG. The trace combination unit 208 may use a nearly lossless constant quality mode, with options including: • Constant QP • Constant rate factor • Recommendations according to slhck.info / video / 2017 / 02 / 24 / crf-guide.html for QP / CRF settings: For x264, typical values ​​are between 18 and 28. The default is 23. 18 should be visually transparent. o For x265, typical values ​​are between 24 and 34. The default is 28. 24 should be visually transparent. o A change of ±6 will result in approximately half / double the file size, although results may vary. o Multiple CRFs can be run in parallel from a single output. • It may be determined whether the discussion at slhck.info / video / 2017 / 02 / 24 / crf-guide.html also applies to XR and video games. The game engine 202, the predictive model encoding unit 204, the intra-model encoding unit 206, and the trace combination unit 208 may be implemented using one or more processors implemented in-circuit, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions associated with these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware may be stored on computer-readable media and executed by the required hardware. Figure 7 is a block diagram illustrating an example system 220 for performing RAN simulation according to the techniques of the present disclosure. System 220 includes a content encoding and delivery model 222, an Internet protocol packet / radio link control (IP / RLC) mapping unit 224, a RAN simulation unit 226, and an IP / RLC to slice mapping unit 228. System 220 may generally correspond to a portion of system 160 of Figure 2. For example, content encoding and delivery model 222 may correspond to content encoding and delivery model 166 of Figure 2, and the slice to IP / RLC mapping unit 224, RAN simulation unit 226, and IP / RLC mapping unit 228 may correspond to the 5GS simulation unit 168 of Figure 2. In this example, the content encoding and delivery model 222 receives v-trace data 216, which includes data representing the encoding of video data frames, from the trace combination unit 208 (FIG. 6). The content encoding and delivery model 222 also receives global configuration data 230. Using the global configuration data 230 (which may correspond to configuration parameter 188 of FIG. 2), the content encoding and delivery model 222 may encode video data frames to form encoded video data slices, and to form s-trace data 232 representing the encoded slices. The content encoding and delivery model 222 may generate slices for a specified time, where each slice corresponds to a sequence of macroblocks (ITU-T H.264 / AVC) or the largest coding unit / coding tree unit (ITU-T H.265 / HEVC). The slice-to-IP / RLC mapping unit 224 may use real-time transport protocol (RTP) to form IP packets and RLC fragments from the s-trace data slices 232. The slice-to-IP / RLC mapping unit 224 may be configured with IP packets and payload size information to form the packets and fragments. In general, if a packet is lost, the entire slice associated with that packet is also lost. The slice-to-IP / RLC mapping unit 224 generates the p-trace data 234, which provides the slices to the IP / RLC mapping unit 224 to the RAN simulation unit 226. The RAN simulation unit 226 may be configured with delay requirements and loss requirements for each slice / packet and determine, through RAN simulation, whether each slice is received or lost at the end of the simulation, as well as the latency value or timestamp value received for each slice. The RAN simulation unit 226 may provide P' trace data 236 including data representing whether slices were received or lost and latency / receive times for the received slices to the IP / RLC to slice mapping unit 228. The IP / RLC to slice mapping unit 228 provides the s' trace data 238 to the content encoding and delivery model 222. The content encoding and delivery model 222 may determine an overall quality value 240 by comparing the decoded video data before the RAN simulation with the decoded video data after the RAN simulation. The RAN simulation unit (226) may be configured according to maximum delay requirements for each downlink slice. For example, there may be a MAC-to-MAC deadline of 10 ms and a requirement for no external loss. For the uplink, there may be tracking, sensor, and pose information, uplink content delivery information, and pose update traffic (e.g., at a periodicity of 1.25 ms or 2 ms). The system may determine how frequency impacts quality, which may depend on the type of XR service. For gaming, pose frequency may have a significant impact. Poses and game actions may also be synchronized. RAN simulation options can include open-loop and closed-loop options. For open-loop, for the entire v-trace, the system can generate s-trace data, p-trace data, one-way RAN simulation, P'-trace data, and s'-trace data, and then perform a quality evaluation, in the form of, for example, an overall quality score of 240. For closed-loop, for each v-trace entry, the system can generate s-trace data, p-trace data, one-way RAN simulation, P'-trace data, and s'-trace data, and then feedback the generated data into the next s-trace, and produce a quality evaluation in the form of an overall quality score of 240. The content encoding and delivery model 222, the slice-to-IP / RLC mapping unit 224, the RAN simulation unit 226, and the IP / RCL-to-slice mapping unit 228 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions associated with these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it is understood that the instructions for the software or firmware may be stored on computer-readable media and executed by the required hardware. Figure 8 is a block diagram illustrating an example system 250 for measuring the performance of an XR content delivery configuration and modeling according to the techniques of the present disclosure. System 250 includes a video content model 252, a content encoding and delivery model 254, a radio access network (RAN) simulation unit 256, a content decoding and delivery model 258, and a content decoding and delivery model 260. These elements of system 250 may correspond to the respective components of system 160 of Figure 2. For example, video content model 252 may correspond to game engine 162 and encoder device model 164 of Figure 2, content encoding and delivery model 254 may correspond to content encoding and delivery model 166 of Figure 2, RAN simulation unit 256 may correspond to 5GS simulation unit 168, and content decoding and delivery model 258 may correspond to content decoding and delivery model 170 of Figure 2. In this example, decoded media data is compared before and after passing through the RAN simulation unit 256. Video data from the decoding and content delivery model 258, which has passed through the RAN simulation unit 256, is assessed to determine a per-frame quality value 262. The decoded media data from the decoding and content delivery model 260 and the value for per-frame quality 262 are used to determine the overall quality data 264. Specifically, the video content model 252 provides v-trace data 268 and encoded video data for transmission to the content encoding and delivery model 254. The content delivery model 254 receives global configuration data 266 and uses the global configuration data 266 to encode and deliver video data in accordance with the v-trace data 268. The global configuration data 266 may include, for example, bitrate control information (e.g., to select one or more of constant quality, constant bitrate, feedback-based variable bitrate, or constant rate factor), slice settings (number of slices, maximum slice size), error robustness (frame-based, slice-based, regular intra-refresh, feedback-based intra-refresh, or feedback-based prediction), and feedback data (disabled, statistical, or operational). The maximum slice size setting may depend on various statistics, including the number of slices, which may vary. The content encoding and delivery model 254 outputs s-trace data 270 to the RAN simulation unit 256 and s'-trace data 270 to the content decoding and delivery model 260. The RAN simulation unit 256 performs RAN simulation on the s-trace data 270 and outputs s'-trace data 272 to the content decoding and delivery model 258. The content decoding and delivery model 260 outputs Q-trace data 274, and the content decoding and delivery model 258 outputs values ​​for per-frame quality 262, resulting in q'trace data 276. The overall quality data 264 may be one or more quality measures for the entire video sequence (e.g., an aggregation of quality for each video frame). Various content encoding considerations may affect the encoding of video data corresponding to v-trace data 268 to form s-trace data 270. In one example, the goal is to have consistent quality, use ITU-T H.265 / HEVC encoding, form a fixed number of slices, one slice is encoded using intra prediction in opportunistic mode, and there is no feedback. In this example, the bitrate control may be using a constant rate factor of 28, a number of slices of 10, an intra-periodic refresh value of 10, intra-periodic refresh and reference image cancellation, and feedback turned off.To generate the slice size, the 254 content encoding and delivery model can generate 10 slices for each video frame, using a constant rate factor of 28 (where ±6 results in approximately half / double the frame size, 12% decrease / increase for each + / -), one intra slice is assumed, based on Trace-vi (i.e., 10% of the adjusted frame size I), and 9 inter slices are generated based on Trace-vp (where two aspects can be considered: frame size (10%) and statistical variation). The RAN simulation unit 256 may maintain state data for each macroblock (ITU-T H.264 / AVC) or coding tree unit (ITU-T H.265 / HEVC), such as: corrupted (missing or predicted areas of corruption, either temporally or spatially), or correct. The RAN simulation unit 256 may provide feedback to the encoding and content delivery model 254, including data representing the number of video data blocks lost or corrupted. The encoding and content delivery model 254, in turn, may integrate the feedback for subsequent coding. For example, whether statistical or operational, the encoding and content delivery model 254 may react to the loss of slices. In response, the encoding and content delivery model 254 may adjust the Bitrate (e.g., obtaining the encoding Bitrate and adjusting the quantization parameters up or down, where + / -1 may result in a 12% impact).The content encoding and delivery model 254 may add intra-prediction slices if a slice is missing (significantly more intra-prediction data may be added in the case of a reported loss. The intra-prediction slices may cover a large area, depending on the motion activity vector. In some cases, the content encoding and delivery model 254 may predict from recognized regions only, which may lead to a statistical increase for the frame size for the missing slice, as the most recent slice may not be used for prediction. The s-trace data 270 may be formatted to include data for a timestamp representing the associated frame, slice size, slice quality (which may include more information than the v-trace, e.g., complexity), macroblock number (for ITU-T H.264 / AVC) or coding tree unit (for ITU-T H.265 / HEVC), and / or slice time (e.g., slice reception deadline). The s-trace data 270' and 272 may be formatted to include all of the information from the formats for s-trace data 270 discussed above, and in addition, data representing slice loss and / or slice delay. There may be a constant deadline for each slice, which can be set to a higher number. For example, for 60 fps video data staggered by each eye, i.e., 120 fps video data, the desired latency might be set to 7 ms. The drop deadline can be set differently, and may be useful, but it doesn't complicate the RAN simulation. Content decoding considerations for generating Q-trace data 274 from s-trace data 270 may be used to take the results of a RAN simulation performed by a RAN simulation unit 256 and to map these simulation results to estimated values ​​for per-frame quality 262. In determining the values ​​for per-frame quality 262, slice considerations may include, for example, the quality of the encoded slices received, the number of missing slices, which slices are of such low quality that they are considered missing, and pending slices, which may be considered missing. Other considerations may include error propagation (e.g., slice prediction from degraded quality slices or missing slices), frame / slice complexity may determine error propagation (how quickly erroneous data propagates to frames), new intra-reset quality, and the percentage of correct / erroneous data for each frame and for different configurations. The metric for determining Q-trace data 274 may be the percentage of erroneous data, which may impact lost / error propagation. The format of Q-trace data 274 may include data representing encoding quality, post-transmission quality (percentage of degraded data), and / or frame data rate. System 250 may maintain a status for each macroblock (ITU-T H.264 / AVC) or coding tree unit (CTU) (ITU-T H.265 / HEVC) as either corrupted or correct. If a macroblock / CTU is corrupted, system 250 may determine whether the macroblock / CTU is part of a missing slice for this transmission or if the macroblock / CTU was correctly received by prediction from a mistranslated / missing region of another slice / frame. If the macroblock / CTU is correct, system 250 determines that the macroblock / CTU was correctly received and predicted from an uncorrupted region of another slice / frame. Predicting from an uncorrupted region of another slice / frame means that the spatial prediction is correct, the temporal prediction is correct, or the macroblock / CTU was correctly restored using an intra refresh and prediction from a correct region of another slice / frame after the intra refresh. System 250 may determine an overall quality value 264 by averaging the coding quality (e.g., averaging over quantization parameters, and possibly converting to peak signal-to-noise ratio (PSNR)), averaging the erroneous video data (e.g., averaging the number of erroneous video frames and / or slices), and / or, if necessary, multiplication / combination / aggregation. 3GPP SA4 defines a constant rate factor (CRF)-based coding model with a specified quality factor (e.g., using the FFMPEG 28 standard). Different configurations for error robustness may be implemented. The criterion for the simulation may be the percentage of erroneous video areas such that at most one macroblock / CTU for every X (e.g., 60) seconds is erroneous. For example, if 4096x4096 is used at 60 fps, then this results in an average of 10e6 erroneous areas. The models discussed above can be used for pixel-based split rendering and / or cloud gaming. Chat applications, video streaming applications (e.g., Twitch.tv), and others can also be used. The video content model 252, the content encoding and delivery model 254, the RAN simulation unit 256, the content decoding and delivery model 258, and the content decoding and delivery model 260 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions associated with these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it is understood that the instructions for the software or firmware may be stored on computer-readable media and executed by the required hardware. System 250 is an example of content modeling based on the discussion in S4-200771. This modeling includes a v-model input, a global configuration for the encoder, statistical or dynamic feedback from the content delivery receiver, a decoding model, and a quality model. Content encoding with content encoding and delivery model 254 can be modeled as follows: • For frame i from v-trace (based on timestamp) o Read frame i from v-trace (timestamp) o Read latest dynamic information from dynamic state info o Perform model encoding (based on input parameters) o For slice s=l, 2, ..., S □ Drop slice s with associated parameters □ New available slice creates IP packet □ For IP packet p=l, 2, ..., P • Drop IP packet with associated parameters and with timestamp to s-trace 1, but also parameters such as slice number The global configuration data configuration parameters 266 may include: • Input: Global configuration o Bitrate Control: Constant Quality, Bitrate Constant, Feedback-Based Variable Bitrate, Constant Rate Factor 28 o Slice settings (number of slices 10, maximum slice size) o Error Resilience, Frame-Based, Slice-Based, Periodic Intra Refresh 10, Feedback-Based Intra Refresh, Feedback-Based Prediction (NVIDIA: Periodic Intra Refresh, Picture Cancellation Reference) • Input: Dynamic per slice information o off o statistics for Bitrate or loss with some delay o operational for Bitrate or loss with some delay Aspects of model coding may include: • Impact of QP settings and intra ratio and slice settings • Feedback o Bitrate Adjustment: Encoder gets encoding Bitrate and adjusts QP (see + / -1 ~12%) o Add Intra in case of lost slices: significantly more is added in case of reported loss. Intra covers a large area (depending on motion vector activity) o Predict only from ACK: statistical improvement for frame size for lost slices, as not the most recent one can be used • Slice settings Emulation of decoding by the decoding and content delivery model 258 may be based on delayed and / or lost slices. Late and lost slices may be considered unavailable and cause errors (e.g., in decoding subsequent frames / slices that refer to the lost or late slices for interpretation). The RAN simulation unit 256 can emulate RAN simulation based on existing 5Qis. Quality evaluation can be based on two aspects, including coding quality and quality degradation due to missing slices. The following simulation can be used to identify damaged macroblocks (or coding tree units or coding units, in ITU-T H.265 / HEVC): • Maintain state for each macroblock (or CU / CTU) o Corrupted o Correct • Macroblock is corrupted o If it is part of the missing slice for this transmission o If received correctly, but predicting from the wrong macroblock • Macroblock is correct o If received correctly and predicting from a non-corrupted macroblock • Predicting from a non-corrupted macroblock o Spatial prediction is correct o Temporal prediction is correct o Recovering the MB is done by Intra refresh and predicting from the correct MB again. Depending on the configuration and quality settings of the transmitted video, different results can be obtained. The quality threshold, for example, might be a maximum of 0.1% of the video's area being corrupted. Additionally, the quality of the original content might be a threshold. Figure 9 is a block diagram illustrating an example pipeline 280 for real-time v-trace data generation according to the techniques of the present disclosure. The pipeline 280 includes a game engine 282, an encoder 284, a decoder 286, a raw data store 288, a predictive encoder 290, an intra encoder 292, and a 3GPP to v-trace unit 294. The components of the pipeline 280 may correspond to the respective components of the system 160 of FIG. 2. For example, the game engine 282 may correspond to the game engine 162 of FIG. 2, the encoder 284 may correspond to the model encoder device 164 of FIG. 2, and the decoder 286, raw data 288, the predictive encoder 290, the intra encoder 292, and the 3GPP to v-trace unit 294 may correspond to the content encoding and delivery model 166 of FIG. 2. Game engine 282 receives input data 296 including model pose and trace data, game data, and game configuration data. Game engine 282 generates video frames 298 from the input data 296. Encoder 284 (which may be an NVIDIA NVENC encoder) encodes video frames 298 using configuration data 302 to generate video bitstream 300. Encoder 284 may operate at 600 GBytes per clock, and may run for several hours. Decoder 286 may be an FFMPEG decoder, and decodes video bitstream 300 to generate decoded video data that is stored as raw data 288 in a computer-readable medium, such as a hard disk, flash drive, or the like. Predictive encoder 290 and intra encoder 292 receive configuration data 304 and encode raw data 288 to form vp-Trace and vi-Trace data, respectively. 3GPP unit to v-trace 294 generates v-trace data from vp-Trace and vi-Trace data. As shown in Figure 9, the encoder 284 can generate the bitstream 300, as discussed above, and the 3GPP v-trace unit 294 can generate the v-trace data. The v-trace generation can be initiated using high-quality output from the game engine 282 and encoded by the encoder 284 using ITU-T H.264(AVC) for the game model and repeatable trace. The decoder 286 can decode this information into 2K x 2K at 120 fps source frames. This sequence can be sent through an encoder model including the predictive encoder 290 and the intra encoder 292, to identify the impact of different parameter settings. The following x.265 parameters can be used initially: • --input viki.yuv (4:2:0) • --profile, -P main10 • --input-res 2048x2048 • --fps 120fps, 60fps, 30fps • -- psnr • -- ssim • --frame 500 (initial testing) • -- bframe 0 • -- crf 22, 25, 28, 31, 34 • --csv logfile.csv • --csv-log-level 2 • --log-level 4 • -- numa -pools 8 • -- keyint, -I 1 and -1 • --slices 1, 128 (2048), 8 (for 2048) • --output rvrviki.h265 • -- rc -lookahead 0 / 1 Based on these 180 encoding processes, a good encoder modeling is expected such that only a subset of processes are required for longer game sequences. Parameters for generating V-trace data in this example may include the duration of the P-trace (e.g., several hours), the variety of games run by the game engine 282, pose traces, settings for the encoder 284 (with the goal, for example, of simple video data that FFMPEG can decode high-quality), and FFMPEG configuration data. Four FFMPEG decoder configurations may be tested, and in some examples, piped in parallel: • HEVC in CRF 28, I-PPP • HEVC in CRF 28, I-IIIIIII • AVC with CRF 23, I-PPP • AVC with CRF 23, I-IIIIII x265.readthedocs.io / en / default / cli.html# describes possible configuration parameters. These parameters can be used to configure rvrplugin.ini with a resolution of 2048x2048, 5ps, and a bitstream generated using the server-side ITU-T H.264 / AVC encoder, which can run at 60 or 120 fps. This data can be re-encoded with crf-24 and updated PSNR and SSIM values. Initial command-line arguments can include the following: • ABR(Rata-rata bit rate): x265.exe --input viki.yuv --input-res 1440x1440 --fps 120 -- psnr -- ssim -frames 500 -- bframes 0 --bitrate 600000 --csv logfile.csv --csv-log-level 2 --log-level 4 -- numa -pools 8 -keyint 1 --output rvrviki.h265 • CRF(Bitrate variabel kualitas konstan): x265.exe --input viki.yuv --input-res 1440x1440 --fps 120 -- psnr -ssim --frames 500 -- bframes 0 -- crf 24 --csv logfilecrf .csv --csv-log-level 2 --log-level 4 -- numa -pools 8 -keyint 1 --output rvrviki.h265 The game engine 282, encoder 284, decoder 286, predictive encoder 290, intra encoder 292, and 3GPP v-trace unit 294 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions associated with these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware may be stored on computer-readable media and executed by the required hardware. Figure 10 is a block diagram illustrating an example rendering and encoding system 310 according to the techniques of the present disclosure. In this example, system 310 includes renderer 312 and encoder 314. Renderer 312 receives content changes 318 and submits information 316 to generate rendered video data, for example, at a frame rate of 60 or 120 fps. Pose information 316 may be provided at a very high frequency with a delay of 2.5 ms. Every 1 / 60 Hz, renderer 312 may retrieve the latest pose from pose information 316, render video data from content changes 318, and send the rendered image to encoder 314 for encoding. Encoder 314 may queue an encode call and then render the image. The 314 encoder can encode video data using ITU-T H.265 / HEVC, for example, x265.exe. The 314 encoder can operate with the following command-line options: • --input viki.yuv (4:2:0) • --profile, -P main10 • --input-res 1440x1440 (2048x2048) • --fps 120fps, 60fps, 30 • -- psnr • -- ssim • --frames 500 (initial testing) • -- bframe 0 • -- crf 22, 25, 28, 31, 34 • --csv logfile.csv • --csv-log-level 2 • --log-level 4 • -- numa -pools 8 • -- keyint, -I 1 and -1 • --slices 1, 90 (for 1440), 128 (2048), 6 (for 1440), 8 (for 2048) • --output rvrviki.h265 • --analysis-save-reuse-level 10 • -- rc -lookahead 0 / 1 The 314 encoder does not need to use “-- frame-dup disabled, “-- constrained-intra, and “-- no-deblock. Renderer 312 and encoder 314 may be implemented using one or more processors implemented in circuitry, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functionality associated with these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware may be stored on computer-readable media and executed by the required hardware. Figure 11 is a graph illustrating an example of a comparison between average PSRN and PSNR in decibels (dB) for various encoder parameters. In this example, there is approximately a 25% improvement from 0 to 1, approximately a 50% improvement from 0 to 2, approximately a 75% improvement from 3, and approximately a 100% improvement from 0 to 4. The improvement may depend on the type of image. For example, if a video sequence produces many intra-prediction frames regardless of configuration, the improvement may be less, and very static content may cause less improvement. This can be used to model invalid reference operations. Error propagation can also be modeled. As a basic principle, an error within a slice can destroy a slice in the current image (e.g., X% is corrupted). As the next reference frame is given, the error can propagate both temporally and spatially, until an intra-frame or intra-slice is received for that region, the region may be corrupted, and the error can propagate spatially to the next frame. This may depend on the number of motion vectors, e.g., in the next frame, if more than X% is corrupted, unless intra-frame prediction is applied. In subsequent frames, the size of the corrupted area may increase. Intra-frame prediction can eliminate error propagation, but with multiple reference frames, the corrupted area may be restored. Figure 12 is a flowchart illustrating an example method of processing media data according to the techniques of the present disclosure. The method of Figure 12 is described with respect to the system 100 of Figure 1, and in particular, the XR server device 110, for the purpose of example. However, it should be understood that other devices may be configured to perform this or similar methods. Initially, the XR server device 110 may receive tracking and sensor information 132 from the XR client device 140 (350). For example, the XR client device 140 may determine the orientation the user is seeking using the tracking / XR sensor 146. The XR client device 140 may send tracking and sensor information 132 representing the orientation the user is seeking. The XR server device 110 may receive the tracking and sensor information 132 and provide this information to the XR scene generation unit 112. The XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114 may then generate video frames from the scene data (352). In some examples, a video game engine may generate scene data using the tracking and sensor information 132, video game data, and video game configuration data. The 2D media encoding unit 116 may then encode video frames generated from the scene data 354. In some examples, as shown and described in connection with Figure 6, such encoding may include intra-predictive encoding of all frames in one process and inter- or intra-predictive encoding of frames in another process. The intra-predictive encoding may produce vi-trace data, while the inter-predictive encoding may produce vp-trace data. The XR server device 110 may combine the vi-trace and vp-trace data to form v-trace data, which may represent the complexity of encoding video frames. In some examples, the 2D media encoding unit 116 may encode all frames of a particular v-trace using a common, invariant quantization configuration. The XR media content delivery unit 118 may then package the 356 encoded video frame slices. In general, the 2D media encoding unit 116 may be configured with a specific maximum transmission unit (MTU) size for a radio access network (RAN) packet and generate slices having an amount of data less than or equal to the MTU size. Thus, each slice may be capable of being transmitted in a single packet, in some examples. The XR server device 110 may then perform a RAN simulation of the RAN to transfer packet 358, for example, using network configuration 130. Examples of performing a RAN simulation in this manner are described above in connection with Figures 2-4 and 7-9. For example, when performing a RAN simulation, the XR server device 110 may determine that a slice is lost when packets for the slice are lost or corrupted. The XR server device 110 may also calculate the latency to transfer the packet. The XR server device 110 may also generate trace data and trace data as discussed above, as part of the RAN simulation. The XR server device 110 may then assemble the simulated received packets into encoded video frames (360) and decode the video frames (362). The XR server device 110 may then calculate the quality of the individual frames (364). The quality of the individual frames may represent the difference between the video frames before or after encoding and the decoded video frames, for example, as discussed in connection with Figures 2 and 8. The XR server device 110 may then determine the overall quality for the system from the quality of the individual frames 366. For example, the XR server device 110 may determine the overall quality as the average encoding quality and / or the average number of erroneous video frames (e.g., frames with missing or corrupted data). The XR server device 110 may perform this method for various different types of configurations, for example, to determine the number of users that can be supported for a particular configuration. For example, the XR server device 110 may calculate the percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations. Additionally or alternatively, the XR server device 110 may calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more numbers of supported users. In this manner, the method of Figure 12 represents an example of a media data processing method including receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation to transmit the encoded video frames over the radio access network; decoding the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculating a value representing individual frame quality for each video frame from the one or more generated video frames and the decoded video frames; and determining an overall quality value from the value representing individual frame quality for each video frame. The following clauses represent specific examples of this disclosure technique: Clause 1: A method of transmitting media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information; rendering an XR display port to form video data from the scene data; encoding the video data; and sending the encoded video data to the XR client device over a 5G network. Clause 2: A method of capturing media data, the method comprising: sending tracking and sensor information to an extended reality (XR) server device; receiving, from the XR server device, encoded video data according to the tracking and sensor information; decoding the encoded video data; and rendering an XR display port using the decoded video data. Clause 3: Method consisting of a combination of the methods of clauses 1 and 2. Clause 4: A method of processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video data to form v-trace data; simulating a radio access network (RAN) for transmitting the encoded video data over a 5G network; decoding the transmitted encoded video data to form decoded video data; calculating a value representing quality for each video frame from the generated one or more video frames and the decoded video data; and determining an overall quality value from the value representing quality for each video frame. Clause 5: The method of clause 4, wherein generating scene data comprises running a video game engine using tracking and sensor information, video game data, and video game configuration data. Clause 6: The method of each of clauses 4 and 5, wherein the video data encoding comprises encoding the video data without changing the quantization configuration. Clause 7: The method of any one of clauses 4-6, wherein determining the overall quality value comprises generating a graph representing the percentage of corrupted video frames as a function of the number of supported users for one or more sets of configurations. Clause 8: The method of any one of clauses 4 and 5, wherein determining the overall quality value comprises generating a graph representing the percentage of corrupted video frames as a function of the quantization configuration for one or more supported numbers of users. Clause 9: The method of any one of clauses 4-8, wherein the encoding of video data to form v-trace data comprises: encoding at least some portions of the video data using inter-prediction to form vp-trace data; encoding at least some portions of the video data using intra-prediction to form vi-trace data; and combining the vp-trace data and the vi-trace data to form v-trace data. Clause 10: The method of any one of clauses 4—9, wherein the v-trace data is formatted according to the passlog file format of FFMPEG. Clause 11: The method of any one of clauses 4-10, wherein the v-trace data includes one or more of a display image number value, an encoded image number value, an image type value, a quality value, an intra-texture-bit value, an inter-texture-bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value. Clause 12: The method of any one of clauses 4-11, wherein performing the RAN simulation comprises: receiving one or more encoded video data slices corresponding to v-trace data; packaging and fragmenting the slices to form packets in the p-trace data using real-time transport protocol (RTP) according to the IP packet and payload size configuration data; simulating a packet transfer; determining that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and outputting data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer. Clause 13: The method of any one of clauses 4-12, wherein performing the RAN simulation comprises: receiving one or more encoded video data slices corresponding to v-trace data; and forming s-trace data from the slices according to configuration data indicating the Bitrate control technique, slice arrangement, error robustness technique, and feedback technique. Clause 14: The method of clause 13, wherein the Bitrate control technique comprises one of constant quality, constant Bitrate, feedback-based variable Bitrate, or constant rate factor. Clause 15: The method of any one of clauses 13 and 14, wherein the slice arrangement comprises one or more of a plurality of slices or a maximum slice size. Clause 16: The method of any of clauses 13-15, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, regular intra-refresh, feedback-based intra-refresh, or feedback-based prediction. Clause 17: A method of any of clauses 13-16, wherein the feedback technique comprises one of no feedback, statistical feedback, or operational feedback. Clause 18: The method of any one of clauses 4-17, wherein calculating a value representing the quality for each video frame comprises: determining, for each received encoded slice, the quality of the slice; and determining, for a slice not received, whether the not received slice is missing or of degraded quality. Clause 19: The method of any one of clauses 4-18, wherein determining the overall quality value comprises determining the overall quality value using one or more average or average encoding qualities of the erroneous video data. Clause 20: Devices for processing media data, devices comprising one or more means for carrying out the methods of any of clauses 1-19. Clause 21: Apparatus of clause 20, wherein the apparatus comprises at least one of: an integrated circuit; a microprocessor; and a wireless communications device. Clause 22: a computer-readable storage medium that has stored instructions therein, which when executed, cause the processor to carry out the methods of any of clauses 1-19. Clause 23: Apparatus for sending media data, the apparatus comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information; means for rendering an XR viewport to form video data from the scene data; means for encoding the video data; and means for sending the encoded video data to the XR client device over a 5G network. Clause 24: Apparatus for capturing media data, the apparatus comprising: means for sending tracking and sensor information to an extended reality (XR) server device; means for receiving, from the XR server device, encoded video data corresponding to the tracking and sensor information; means for decoding the encoded video data; and means for rendering an XR display port using the decoded video data. Clause 25: Apparatus for processing media data, the apparatus comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; means for encoding the video data into v-trace data; means for simulating a radio access network (RAN) for transmitting the encoded video data over a 5G network; means for decoding the transmitted encoded video data to form decoded video data; means for calculating a value representing quality for each video frame from the generated one or more video frames and the decoded video data; and means for determining an overall quality value from the value representing quality for each video frame. Clause 26: A method of processing media data, the method comprising: receiving tracking and sensor information from an augmented reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; simulating a radio access network (RAN) to transmit the encoded video frames over the radio access network; decoding the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculating a value representing individual frame quality for each video frame from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the values ​​representing individual frame quality for each video frame. Clause 27: The method of clause 26, wherein generating scene data comprises running a video game engine using tracking and sensor information, video game data, and video game configuration data. Clause 28: The method of clause 26, wherein the video frame encoding comprises encoding the video frame without changing the quantization configuration. Clause 29: The method of clause 26, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets. Clause 30: The method of clause 26, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the quantization configuration for one or more supported numbers of users. Clause 31: The method of clause 26, wherein the encoding of a video frame further comprises forming v-trace data, the v-trace data representing the complexity of encoding the video frame, comprising: encoding at least some portions of the video frame using inter-prediction to form vp-trace data; encoding at least some portions of the video frame using intra-prediction to form vi-trace data; and combining the vp-trace data and the vi-trace data to form v-trace data. Clause 32: The method of clause 26, wherein the encoding of the video frame further comprises generating v-trace data, the v-trace data representing the complexity of the encoding of the video frame, and wherein the v-trace data comprises one or more of displaying a picture number value, a picture number encoded value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value. Clause 33: The method of clause 26, wherein performing a RAN simulation comprises: receiving one or more slices of an encoded video frame; packaging and fragmenting the slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; performing a packet transfer simulation; determining that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and outputting data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer. Clause 34: The method of clause 26, wherein performing the RAN simulation comprises: receiving one or more slices of an encoded video frame; and forming s-trace data from the slices in accordance with configuration data indicating the Bitrate control technique, slice arrangement, error robustness technique, and feedback technique. Clause 35: The method of clause 34, wherein forming the s-trace data comprises determining, for each slice, data representing the frame associated with the slice, the slice size, the slice quality, the area of ​​the frame covered by the slice, and timing information for the slice. Clause 36: The method of clause 34, wherein the Bitrate control technique comprises one of constant quality, constant Bitrate, feedback-based variable Bitrate, or constant rate factor. Clause 37: The method of clause 34, wherein the slice arrangement comprises one or more of a plurality of slices or a maximum slice size. Clause 38: The method of clause 34, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, regular intra-refresh, feedback-based intra-refresh, or feedback-based prediction. Clause 39: The method of clause 34, where the feedback technique comprises either statistical feedback or operational feedback. Clause 40: The method of clause 26, wherein calculating a value representing the individual frame quality for each video frame comprises: determining, for each received encoded video frame slice, the slice quality; and determining, for a video frame slice that is not received, whether the unreceived slice is missing or of degraded quality. Clause 41: The method of clause 26, wherein determining the overall quality value comprises determining the overall quality value using one or more average encoding qualities or an average of the erroneous video frames. Clause 42: The method of clause 26, where the radio access network comprises a 5G network. Clause 43: The method of clause 26, where the RAN simulation comprises the first RAN simulation of a number of simulations RAN, the overall video quality comprises a first overall video quality, and performing the first RAN simulation comprises performing the first RAN simulation using the first RAN simulation configuration parameters, the method further comprising: determining respective configuration parameters for each of the plurality of RAN simulations; performing each RAN simulation using the respective configuration parameters; determining an overall video quality value for each RAN simulation; and configuring the network to transmit XR data using the configuration parameters corresponding to the best overall video quality value for the RAN simulation. Clause 44: Device for processing media data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing the quality of an individual frame for each video frame of the one or more generated video frames and the decoded video frames;and determine the overall quality value from the values ​​representing the individual frame quality for each video frame.; Clause 45: The device of clause 44, wherein one or more processors are configured to execute a video game engine using tracking and sensor information, video game data, and video game configuration data to generate scene data. Clause 46: The device of clause 44, wherein to determine an overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames as a function of a number of supported users for the one or more configuration sets. Clause 47: The device of clause 44, wherein to determine an overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames as a function of a quantization configuration for one or more supported numbers of users. Clause 48: The device of clause 44, wherein the one or more processors are further configured to form v-trace data, the v-trace data representing the complexity of encoding a video frame, and wherein to form the v-trace data, the one or more processors are configured to: encode at least some portions of the video frame using inter-prediction to form vp-trace data; encode at least some portions of the video frame using intra-prediction to form vi-trace data; and combine the vp-trace data and the vi-trace data to form v-trace data. Clause 49: The device of clause 44, wherein the one or more processors are further configured to generate v-trace data, the v-trace data representing the complexity of encoding a video frame, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value. Clause 50: The device of clause 44, wherein to perform the RAN simulation, the one or more processors are configured to: receive one or more encoded video frame slices; packaging and fragmenting slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; performing a packet transfer simulation; determining that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and outputting data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer. Clause 51: The device of clause 44, wherein to perform the RAN simulation, the one or more processors are configured to: receive one or more encoded video frame slices; and form s-trace data from the slices in accordance with configuration data indicating the Bitrate control technique, slice arrangement, error robustness technique, and feedback technique. Clause 52: The device of clause 51, wherein to form the s-trace data, the one or more processors are configured to assign, to each slice, data representing the frame associated with the slice, the slice size, the slice quality, the area of ​​the frame covered by the slice, and timing information for the slice. Clause 53: The apparatus of clause 44, wherein to calculate a value representing the individual frame quality for each video frame, the one or more processors are configured to: determine, for each received encoded video frame slice, the slice quality; and determine, for a video frame slice that is not received, whether the unreceived slice is missing or of degraded quality. Clause 54: a computer-readable storage medium that has stored instructions that, when executed, cause a processor to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames; encode the video frames to form encoded video frames; simulate a radio access network (RAN) to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing the quality of individual frames for each video frame from the one or more generated video frames and the decoded video frames; and determine an overall quality value from the values ​​representing the quality of individual frames for each video frame. Clause 55: The computer-readable storage medium of clause 54, wherein the instructions that cause the processor to generate scene data comprise instructions that cause the processor to execute a video game engine using tracking and sensor information, video game data, and video game configuration data. Clause 56: The computer-readable storage medium of clause 54, wherein the instructions that cause the processor to determine the overall quality value comprise instructions that cause the processor to calculate the percentage of corrupted video frames as a function of a number of supported users for one or more configuration sets. Clause 57: The computer-readable storage medium of clause 54, wherein the instructions causing the processor to determine the overall quality value comprise instructions causing the processor to calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more supported numbers of users. Clause 58: The computer-readable storage medium of clause 54, wherein the instructions that cause the processor to encode a video frame further comprise instructions that cause the processor to form v-trace data, the v-trace data representing the complexity of encoding the video frame, and wherein the instructions that cause the processor to form the v-trace data include instructions that cause the processor to: encode at least some portions of the video frame using inter-prediction to form vp-trace data; encode at least some portions of the video frame using intra-prediction to form vi-trace data; and combine the vp-trace data and the vi-trace data to form v-trace data. Clause 59: The computer-readable storage medium of clause 54, wherein the instructions that cause the processor to encode a video frame further comprise instructions that cause the processor to form v-trace data, the v-trace data representing the complexity of encoding the video frame, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value. Clause 60: The computer-readable storage medium of clause 54, wherein the instructions that cause the processor to perform a RAN simulation comprise instructions that cause the processor to: receive one or more encoded video frame slices; package and fragment the slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; perform a packet transfer simulation; determine that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and output data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer. Clause 61: The computer-readable storage medium of clause 54, wherein the instructions that cause the processor to perform the RAN simulation comprise instructions that cause the processor to: receive one or more encoded video frame slices; and form s-trace data from the slices in accordance with configuration data indicating the Bitrate control technique, slice settings, error resilience technique, and feedback technique. Clause 62: The computer-readable storage medium of clause 61, wherein the instructions that cause the processor to form the s-trace data comprise instructions that cause the processor to submit, to each slice, data representing the frame associated with the slice, the size of the slice, the quality of the slice, the area of ​​the frame covered by the slice, and timing information for the slice. Clause 63: The computer-readable storage medium of clause 54, wherein the instructions that cause the processor to compute a value representing the individual frame quality for each video frame comprise instructions that cause the processor to: determine, for each received encoded slice of the video frame, the slice quality; and determine, for a slice of the video frame that is not received, whether the unreceived slice is lost or of degraded quality. Clause 64: Apparatus for processing media data, the apparatus comprising: means for receiving tracking and sensor information from an augmented reality (XR) client device; means for generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; means for encoding the video frames to form encoded video frames; means for simulating a radio access network (RAN) for transmitting the encoded video frames over the radio access network; means for decoding the encoded video frames transmitted according to the RAN simulation to form decoded video frames; means for calculating a value representing individual frame quality for each video frame from the one or more generated video frames and the decoded video frames; and means for determining an overall quality value from the values ​​representing individual frame quality for each video frame. Clause 65: A method of processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; simulating a radio access network (RAN) to transmit the encoded video frames over the radio access network; decoding the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculating a value representing individual frame quality for each video frame from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the values ​​representing individual frame quality for each video frame. Clause 66: The method of clause 65, wherein generating scene data comprises running a video game engine using tracking and sensor information, video game data, and video game configuration data. Clause 67: The method of any one of clauses 65-66, wherein the video frame encoding comprises encoding the video frame without changing the quantization configuration. Clause 68: The method of any one of clauses 65-67, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets. Clause 69: The method of any one of clauses 65-68, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the quantization configuration for one or more supported numbers of users. Clause 70: The method of any one of clauses 65-69, wherein the encoding of a video frame further comprises forming v-trace data, the v-trace data representing the complexity of encoding the video frame, comprising: encoding at least some portions of the video frame using interprediction to form vp-trace data; encoding at least some portions of the video frame using intraprediction to form vi-trace data; and combining the vp-trace data and the vi-trace data to form v-trace data. Clause 71: The method of any one of clauses 65-70, wherein the video frame encoding further comprises generating v-trace data, the v-trace data representing the complexity of the video frame encoding, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value. Clause 72: The method of any of clauses 65-71, wherein performing a RAN simulation comprises: receiving one or more encoded video frame slices; packaging and fragmenting the slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; performing a packet transfer simulation; determining that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and outputting data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer. Clause 73: The method of any one of clauses 65-72, wherein performing the RAN simulation comprises: receiving one or more encoded video frame slices; and forming s-trace data from the slices in accordance with configuration data indicating the Bitrate control technique, slice arrangement, error robustness technique, and feedback technique. Clause 74: The method of clause 73, wherein the formation of the s-trace data comprises determining, for each slice, data representing the frame associated with the slice, the slice size, the slice quality, the area of ​​the frame covered by the slice, and timing information for the slice. Clause 75: The method of any one of clauses 73 and 74, wherein the Bitrate control technique comprises one of constant quality, constant Bitrate, feedback-based variable Bitrate, or constant rate factor. Clause 76: The method of any one of clauses 73—75, wherein the slice arrangement comprises one or more of a plurality of slices or a maximum slice size. Clause 77: The method of any of clauses 73-76, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, regular intra-refresh, feedback-based intra-refresh, or feedback-based prediction. Clause 78: A method of any of clauses 73-77, wherein the feedback technique comprises either statistical feedback or operational feedback. Clause 79: The method of any one of clauses 65—78, wherein calculating a value representing the individual frame quality for each video frame comprises: determining, for each received encoded video frame slice, the slice quality; and determining, for a video frame slice that is not received, whether the unreceived slice is missing or of degraded quality. Clause 80: The method of any one of clauses 65-79, wherein determining the overall quality value comprises determining the overall quality value using one or more average encoding qualities or averages of the erroneous video frames. Clause 81: The method of any one of clauses 65-80, wherein the radio access network comprises a 5G network. Clause 82: The method of any one of clauses 65-81, wherein the RAN simulation comprises a first RAN simulation of a plurality of RAN simulations, the overall video quality comprises a first overall video quality, and performing the first RAN simulation comprises performing the first RAN simulation using first configuration parameters, the method further comprising: determining respective configuration parameters for each of the plurality of RAN simulations; performing each RAN simulation using the respective configuration parameters; determining an overall video quality value for each RAN simulation; and configuring the network to transmit XR data using configuration parameters corresponding to the best overall video quality value for the RAN simulation. Clause 83: Device for processing media data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; simulate the radio access network (RAN) to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing the quality of the individual frames for each video frame of the generated one or more video frames and the decoded video frames;and determine the overall quality value from the values ​​representing the individual frame quality for each video frame.; Clause 84: The device of clause 83, wherein one or more processors are configured to execute a video game engine using tracking and sensor information, video game data, and video game configuration data to generate scene data. Clause 85: The device of any one of clauses 83 and 84, wherein to determine an overall quality value, one or more processors are configured to calculate a percentage of corrupted video frames as a function of a number of supported users for one or more sets of configurations. Clause 86: The device of any one of clauses 83-85, wherein to determine an overall quality value, one or more processors are configured to calculate a percentage of corrupted video frames as a function of a quantization configuration for one or more user-supported numbers. Clause 87: The device of any one of clauses 83-86, wherein the one or more processors are further configured to form v-trace data, the v-trace data representing the complexity of encoding a video frame, and wherein to form the v-trace data, the one or more processors are configured to: encode at least some portions of the video frame using inter-prediction to form vp-trace data; encode at least some portions of the video frame using intra-prediction to form vi-trace data; and combine the vp-trace data and the vi-trace data to form v-trace data. Clause 88: The device of any of clauses 83-87, wherein the one or more processors are further configured to generate v-trace data, the v-trace data representing the complexity of encoding a video frame, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value. Clause 89: The apparatus of any one of clauses 83—88, wherein to perform a RAN simulation, the one or more processors are configured to: receive one or more encoded video frame slices; package and fragment the slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; perform a packet transfer simulation; determine that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and output data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer. Clause 90: A device of any one of clauses 83—89, wherein to perform the RAN simulation, one or more processors are configured to: receive one or more encoded video frame slices; and form trace data from the slices in accordance with configuration data indicating Bitrate control techniques, slice settings, error robustness techniques, and feedback techniques. Clause 91: The device of clause 90, wherein to form the s-trace data, the one or more processors are configured to assign, to each slice, data representing the frame associated with the slice, the slice size, the slice quality, the area of ​​the frame covered by the slice, and timing information for the slice. Clause 92: The apparatus of any one of clauses 83-91, wherein, to calculate a value representing the individual frame quality for each video frame, the one or more processors are configured to: determine, for each received encoded video frame slice, the slice quality; and determine, for a video frame slice that is not received, whether the unreceived slice is missing or of degraded quality. Clause 93: a computer-readable storage medium that has stored instructions that, when executed, cause a processor to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames; encode the video frames to form encoded video frames; simulate a radio access network (RAN) to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing the quality of individual frames for each video frame from the one or more generated video frames and the decoded video frames; and determine an overall quality value from the values ​​representing the quality of individual frames for each video frame. Clause 94: The computer-readable storage medium of clause 93, wherein the instructions that cause the processor to generate scene data comprise instructions that cause the processor to execute a video game engine using tracking and sensor information, video game data, and video game configuration data. Clause 95: A computer-readable storage medium of any one of clauses 93 and 94, wherein the instructions that cause the processor to determine the overall quality value comprise instructions that cause the processor to calculate the percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets. Clause 96: The computer-readable storage medium of any one of clauses 93-95, wherein the instructions that cause the processor to determine the overall quality value comprise instructions that cause the processor to calculate the percentage of corrupted video frames as a function of configuration quantization for one or more supported numbers of users. Clause 97: The computer-readable storage medium of any of clauses 93-96, wherein the instructions that cause the processor to encode a video frame further comprise instructions that cause the processor to form v-trace data, the v-trace data representing the complexity of encoding the video frame, and wherein the instructions that cause the processor to form the v-trace data include instructions that cause the processor to: encode at least some portions of the video frame using inter-prediction to form vp-trace data; encode at least some portions of the video frame using intra-prediction to form vi-trace data; and combine the vp-trace data and the vi-trace data to form v-trace data. Clause 98: A computer-readable storage medium of any of clauses 93—97, wherein the instructions that cause the processor to encode a video frame further comprise instructions that cause the processor to generate v-trace data, the v-trace data representing the complexity of encoding the video frame, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an in-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a bit value header. Clause 99: A computer-readable storage medium of any of clauses 93—98, wherein the instructions that cause the processor to perform the RAN simulation comprise instructions that cause the processor to: receive one or more encoded video frame slices; package and fragment the slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; perform a packet transfer simulation; determine that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and output data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer. Clause 100: A computer-readable storage medium of any of clauses 93—99, wherein the instructions causing the processor to perform the RAN simulation comprise instructions causing the processor to: receive one or more encoded video frame slices; and form s-trace data from the slices in accordance with configuration data indicating Bitrate control techniques, slice settings, error resilience techniques, and feedback techniques. Clause 101: The computer-readable storage medium of clause 100, wherein the instructions that cause the processor to form the s-trace data comprise instructions that cause the processor to submit, to each slice, data representing the frame associated with the slice, the size of the slice, the quality of the slice, the area of ​​the frame covered by the slice, and timing information for the slice. Clause 102: The computer-readable storage medium of any of clauses 93-101, wherein the instructions that cause the processor to compute a value representing the individual frame quality for each video frame comprise instructions that cause the processor to: determine, for each received slice of an encoded video frame, the slice quality; and determine, for a slice of a video frame that is not received, whether the unreceived slice is lost or of degraded quality. Clause 103: Apparatus for processing media data, the apparatus comprising: means for receiving tracking and sensor information from an augmented reality (XR) client device; means for generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; means for encoding the video frames to form encoded video frames; means for simulating a radio access network (RAN) for transmitting the encoded video frames over the radio access network; means for decoding the encoded video frames transmitted according to the RAN simulation to form decoded video frames; means for calculating a value representing the quality of an individual frame for each video frame from the one or more generated video frames and the decoded video frames; and means for determining an overall quality value from the values ​​representing the quality of an individual frame for each video frame. In one or more examples, the described functionality may be implemented in hardware, software, firmware, or a combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable medium may include a computer-readable storage medium, corresponding to a tangible medium such as a data storage medium, or a communications medium including any medium that facilitates the transfer of a computer program from one location to another, for example, according to a communications protocol. In this way, computer-readable medium may generally correspond to (1) a non-transitory computer-readable storage medium or (2) a communications medium such as a carrier signal or wave.The data storage medium may be any available medium that is accessible to one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media. For example, and without limitation, such computer-readable storage media may consist of RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. Also, any connection is properly called a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of medium.However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but rather refer to tangible storage media that are not temporary. Disks and discs, as used herein, include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where discs typically reproduce data magnetically, while discs reproduce data optically with a laser. The combination of the above should also be included within the scope of computer-readable media. Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor," as used herein, may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or combined in a combined codec. Also, the techniques may be fully implemented in one or more circuits or logic elements. The techniques of the present disclosure may be implemented in a variety of devices or apparatus, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize the functional aspects of a device configured to perform the disclosed techniques, but do not necessarily require realization by distinct hardware units. Instead, as described above, the various units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors as described above, together with appropriate software and / or firmware. Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method of processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; simulating a radio access network (RAN) to transmit the encoded video frames over the radio access network; decoding the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculating a value representing individual frame quality for each video frame from the generated one or more video frames and the decoded video frames; and determining an overall quality value from the value representing individual frame quality for each video frame.

2. The method of claim 1, wherein generating the scene data comprises running a video game engine using tracking and sensor information, video game data, and video game configuration data.

3. The method of claim 1, wherein the video frame encoding comprises encoding the video frame without changing the quantization configuration.

4. The method of claim 1, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets.

5. The method of claim 1, wherein determining the overall quality value comprises calculating the percentage of corrupted video frames as a function of the quantization configuration for one or more supported numbers of users.

6. The method of claim 1, wherein the encoding of the video frame further comprises forming v-trace data, the v-trace data representing the complexity of encoding the video frame, comprising: encoding at least some portions of the video frame using inter-prediction to form vp-trace data; encoding at least some portions of the video frame using intra-prediction to form vi-trace data; and combining the vp-trace data and the vi-trace data to form v-trace data.

7. The method of claim 1, wherein the encoding of the video frame further comprises generating v-trace data, the v-trace data representing the complexity of the encoding of the video frame, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value.

8. The method of claim 1, wherein performing the RAN simulation comprises: receiving one or more encoded video frame slices; packaging and fragmenting the slices to form packets in p-trace data using real-time transport protocol (RTP) according to the IP packet and payload size configuration data; performing the packet transfer simulation; determining that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and outputting data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer.

9. The method of claim 1, wherein performing the RAN simulation comprises: receiving one or more encoded video frame slices; and forming s-trace data from the slices according to configuration data indicating a Bitrate control technique, a slice setting, an error robustness technique, and a feedback technique.

10. The method of claim 9, wherein the formation of the s-trace data comprises providing, for each slice, data representing the frame associated with the slice, the slice size, the slice quality, the area of ​​the frame covered by the slice, and timing information for the slice.

11. The method of claim 9, wherein the Bitrate control technique comprises one of constant quality, constant Bitrate, feedback-based variable Bitrate, or constant rate factor.

12. The method of claim 9, wherein the slice arrangement comprises one or more of a plurality of slices or a maximum slice size.

13. The method of claim 9, wherein the error resilience technique comprises one of frame-based resilience, slice-based resilience, regular intra refresh, feedback-based intra refresh, or feedback-based prediction.

14. The method of claim 9, wherein the feedback technique comprises either statistical feedback or operational feedback.

15. The method of claim 1, wherein calculating a value representing the individual frame quality for each video frame comprises: determining, for each received encoded video frame slice, the slice quality; and determining, for a video frame slice that is not received, whether the unreceived slice is missing or of degraded quality.

16. The method of claim 1, wherein determining the overall quality value comprises determining the overall quality value using one or more average encoding qualities or an average number of erroneous video frames.

17. The method in claim 1, wherein the radio access network comprises a 5G network.

18. The method of claim 1, wherein the RAN simulation comprises a first RAN simulation of a plurality of RAN simulations, the overall video quality comprises a first overall video quality, and performing the first RAN simulation comprises performing the first RAN simulation using first configuration parameters, the method further comprising: determining respective configuration parameters for each of the plurality of RAN simulations; the respective configuration parameters; determining an overall video quality value for each RAN simulation; and configuring the network to transmit XR data using configuration parameters corresponding to the best overall video quality value for the RAN simulation.

19. A device for processing media data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; simulate a radio access network (RAN) to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing the quality of individual frames for each video frame of the generated one or more video frames and the decoded video frames;and determines the overall quality value from the values ​​representing the individual frame quality for each video frame.; 20. The device of claim 19, wherein the one or more processors are configured to execute a video game engine using tracking and sensor information, video game data, and video game configuration data to generate scene data.

21. The device of claim 19, wherein to determine an overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames as a function of a number of supported users for the one or more configuration sets.

22. The device of claim 19, wherein to determine the overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames as a function of the quantization configuration for the one or more supported number of users.

23. The device of claim 19, wherein the one or more processors are further configured to form v-trace data, the v-trace data representing the complexity of encoding a video frame, and wherein to form the v-trace data, the one or more processors are configured to: encode at least some portions of the video frame using inter-prediction to form vp-trace data; encode at least some portions of the video frame using intra-prediction to form vi-trace data; and combine the vp-trace data and the vi-trace data to form v-trace data.

24. The device of claim 19, wherein the one or more processors are further configured to form v-trace data, the v-trace data representing the complexity of encoding a video frame, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value.

25. The device of claim 19, wherein to perform the RAN simulation, the one or more processors are configured to: receive one or more encoded video frame slices; package and fragment the slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; perform the packet transfer simulation; determine that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and output data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer.

26. The device of claim 19, wherein to perform the RAN simulation, the one or more processors are configured to: receive one or more encoded video frame slices; and form s-trace data from the slices in accordance with configuration data indicating a Bitrate control technique, slice settings, error robustness technique, and feedback technique.

27. The device of claim 26, wherein to form the s-trace data, the one or more processors are configured to assign, to each slice, data representing the frame associated with the slice, the slice size, the slice quality, the area of ​​the frame covered by the slice, and timing information for the slice.

28. The device of claim 19, wherein to calculate a value representing an individual frame quality for each video frame, the one or more processors are configured to: determine, for each received encoded video frame slice, the slice quality; and determine, for a video frame slice that is not received, whether the unreceived slice is missing or of degraded quality.

29. A computer-readable storage medium that has stored instructions that, when executed, cause a processor to: receive tracking and sensor information from an extended reality (XR) client device; generate scene data using the tracking and sensor information, the scene data comprising one or more video frames; encode the video frames to form encoded video frames; simulate a radio access network (RAN) to transmit the encoded video frames over the radio access network; decode the transmitted encoded video frames in accordance with the RAN simulation to form decoded video frames; calculate a value representing individual frame quality for each video frame from the one or more generated video frames and the decoded video frames; and determine an overall quality value from the values ​​representing individual frame quality for each video frame.

30. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to generate scene data comprise instructions causing the processor to execute a video game engine using tracking and sensor information, video game data, and video game configuration data.

31. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to determine the overall quality value comprise instructions causing the processor to calculate the percentage of corrupted video frames as a function of a number of supported users for one or more configuration sets.

32. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to determine the overall quality value comprise instructions causing the processor to calculate the percentage of corrupted video frames as a function of the quantization configuration for one or more supported numbers of users.

33. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to encode the video frame further comprise instructions causing the processor to form v-trace data, the v-trace data representing the complexity of encoding the video frame, and wherein the instructions causing the processor to form the v-trace data include instructions causing the processor to: encode at least some portions of the video frame using inter-prediction to form vp-trace data; encode at least some portions of the video frame using intra-prediction to form vi-trace data; and combine the vp-trace data and the vi-trace data to form v-trace data.

34. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to encode the video frame further comprise instructions causing the processor to generate v-trace data, the v-trace data representing the complexity of encoding the video frame, and wherein the v-trace data includes one or more of a display picture number value, an encoded picture number value, a picture type value, a quality value, an intra-texture bit value, an inter-texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, a mc_mb_var_sum value, a mb_var_sum value, an i_count value, a skip count value, or a header bit value.

35. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to perform a RAN simulation comprise instructions causing the processor to: receive one or more encoded video frame slices; package and fragment the slices to form packets in p-trace data using real-time transport protocol (RTP) according to IP packet and payload size configuration data; perform a packet transfer simulation; determine that one of the slices is lost when a packet for one of the slices is lost during the packet transfer simulation; and output data representing, for each slice, whether the slice was received or lost during the transfer simulation and the latency for the received packet transfer.

36. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to perform the RAN simulation comprise instructions causing the processor to: receive one or more encoded video frame slices; and form s-trace data from the slices in accordance with configuration data indicating a Bitrate control technique, a slice setting, an error resilience technique, and a feedback technique.

37. The computer-readable storage medium of claim 36, wherein the instructions causing the processor to form the s-trace data comprise instructions causing the processor to submit, to each slice, data representing the frame associated with the slice, the size of the slice, the quality of the slice, the area of ​​the frame covered by the slice, and timing information for the slice.

38. The computer-readable storage medium of claim 29, wherein the instructions causing the processor to compute a value representing an individual frame quality for each video frame comprise instructions causing the processor to: determine, for each received encoded video frame slice, the slice quality; and determine, for a video frame slice not received, whether the not received slice is missing or of degraded quality.

39. A device for processing media data, the device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; means for encoding the video frames to form encoded video frames; means for simulating a radio access network (RAN) for transmitting the encoded video frames over the radio access network; means for decoding the encoded video frames transmitted according to the RAN simulation to form decoded video frames; means for calculating a value representing individual frame quality for each video frame from the one or more generated video frames and the decoded video frames; and means for determining an overall quality value from the values ​​representing individual frame quality for each video frame.