Perform and evaluate segmentation rendering over 5G networks

By simulating and optimizing the segmented rendering process of the XR system and using the 5G network for RAN simulation, the problems of high latency and insufficient network capacity in XR applications are solved, efficient video data transmission and encoding are achieved, and user experience and system performance are improved.

CN115398927BActive Publication Date: 2025-09-19QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180028078.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-05-17
Filing Date
2021-05-18
Publication Date
2025-09-19
Estimated Expiration
2041-05-18

AI Technical Summary

Technical Problem

Existing video data transmission and encoding technologies suffer from high latency, insufficient network capacity, and poor encoding parameter settings in extended reality (XR) applications, affecting user experience and system performance.

Method used

By simulating and evaluating the segmented rendering process of the XR system, using the 5G network to simulate the radio access network (RAN), combined with video frame quality assessment and encoder parameter optimization, it provides good encoding and content delivery parameter settings, identifies system capacity and defines optimization solutions.

Benefits of technology

Effectively reduce latency, optimize network performance, improve video quality and user experience, and ensure smooth video under high user load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115398927B_ABST
    Figure CN115398927B_ABST
Patent Text Reader

Abstract

An example device includes: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; use the tracking and sensor information to generate scene data, the scene data comprising one or more video frames of video data; encode the video frames to form encoded video frames; perform a radio access network simulation of delivering the encoded video frames via a radio access network (RAN); decode the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determine an overall quality value based on the value representing the individual frame quality for each of the video frames.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of priority to U.S. Application No. 17 / 322,468, filed on May 17, 2021, and U.S. Provisional Application No. 63 / 026,498, filed on May 18, 2020, each of which is hereby incorporated by reference in its entirety. U.S. Application No. 17 / 322,468 claims the benefit of U.S. Provisional Application No. 63 / 026,498, filed on May 18, 2020. Technical Field

[0002] The present disclosure relates to the storage and transmission of encoded media data. Background Art

[0003] Digital video capabilities may be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio telephones, video teleconferencing equipment, etc. Digital video devices implement video compression techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263 or ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Coding (AVC)), ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)), and extensions of such standards, to more efficiently transmit and receive digital video information.

[0004] After the video data and other media data have been encoded, the media data can be packetized for transmission or storage. The media data can be assembled into a video file that conforms to any of a variety of standards, such as the International Organization for Standardization (ISO) base media file format and its extensions. Summary of the Invention

[0005] In summary, the present disclosure describes techniques related to evaluating the configuration of various components involved in performing segmented rendering for interactive media (such as online cloud video games). These techniques can be implemented by a computing device or a system of computing devices. These techniques include an evaluation framework for evaluating the performance of these devices and the configuration of these devices. For example, different parts of the system can be evaluated in different ways to determine different configuration settings. Specifically, these techniques include video-level, slice-level, and group-level evaluations. Evaluation at the group-level can be used to determine whether the corresponding slice has been correctly received. Evaluation at the slice-level can be used to determine the quality of the slice and the corresponding frame including the slice. Evaluation at the video level can be used to determine whether the overall configuration parameter set for all components of the system is appropriate.

[0006] In one example, a method of processing media data includes: receiving tracking and sensor information from an extended reality (XR) client device; using the tracking and sensor information to generate scene data, the scene data including one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determining an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0007] In another example, a device for processing media data includes: a memory configured to store video data; and one or more processors implemented in circuits and configured to: receive tracking and sensor information from an extended reality (XR) client device; use the tracking and sensor information to generate scene data, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decode the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determine an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0008] In another example, a computer-readable storage medium having instructions stored thereon that, when executed, cause a processor to: receive tracking and sensor information from an extended reality (XR) client device; use the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decode the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determine an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0009] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a block diagram illustrating an example computing system that can perform the techniques of this disclosure.

[0011] Figure 2 is a diagram showing the technique according to the present disclosure for simulating and measuring the Figure 1 A conceptual diagram of an example decomposition of the performance of a computing system performing a segmented rendering process.

[0012] Figure 3-Figure 5 Is it about Figure 2 Example graphs for various radio access network (RAN) simulations discussed.

[0013] Figure 6 is a block diagram illustrating an example v-tracking system model in accordance with techniques of this disclosure.

[0014] Figure 7 is a block diagram illustrating an example system for performing RAN simulation in accordance with techniques of this disclosure.

[0015] Figure 8 is a block diagram illustrating an example system for measuring the performance of XR content delivery configurations and modeling, in accordance with techniques of this disclosure.

[0016] Figure 9 is a block diagram illustrating an example pipeline for generating v-tracking data in real time, in accordance with techniques of this disclosure.

[0017] Figure 10 is a block diagram illustrating an example rendering and encoding system in accordance with techniques of this disclosure.

[0018] Figure 11 is a graph showing an example comparison between average PSRN and PSNR (in decibels (dB)) for various encoder parameters.

[0019] Figure 12 is a flowchart illustrating an example method of processing media data according to the techniques of this disclosure. DETAILED DESCRIPTION

[0020] Figure 11 is a block diagram illustrating an example computing system 100 that can perform techniques of this disclosure. In this example, computing system 100 includes an extended reality (XR) server device 110, a network 130, an XR client device 140, and a display device 152. XR server device 110 includes an XR scene generation unit 112, an XR viewport pre-rendering rasterization unit 114, a 2D media encoding unit 116, an XR media content delivery unit 118, and a 5G system (5GS) delivery unit 120. Network 130 can correspond to any network of computing devices that communicate according to one or more network protocols, such as the Internet. Specifically, network 130 can include a 5G radio access network (RAN), which includes access devices to which XR client device 140 connects to access network 130 and XR server device 110. In other examples, other types of networks, such as other types of RANs, can be used. The XR client device 140 includes a 5GS delivery unit 150, a tracking / XR sensor 146, an XR viewport rendering unit 142, a 2D media decoder 144, and an XR media content delivery unit 148. The XR client device 140 also interfaces with a display device 152 to present XR media data to a user (not shown).

[0021] In some examples, the XR scene generation unit 112 may correspond to an interactive media entertainment application (such as a video game), which may be executed by one or more processors implemented in circuitry of the XR server device 110. The XR viewport prerendering rasterization unit 114 may format the scene data generated by the XR scene generation unit 112 into prerendered two-dimensional (2D) media data (e.g., video data) for use in the viewport of a user of the XR client device 140. The 2D media encoding unit 116 may encode the formatted scene data from the XR viewport prerendering rasterization unit 114 using, for example, a video coding standard such as ITU-T H.264 / Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC), etc. In this example, the XR media content delivery unit 118 represents a content delivery sender. In this example, the XR media content delivery unit 148 represents a content delivery recipient, and the 2D media decoder 144 may perform error handling.

[0022] Typically, the XR client device 140 can determine a user's viewport, e.g., the direction the user is looking and the user's physical location, which can correspond to the orientation of the XR client device 140 and the geographic location of the XR client device 140. The tracking / XR sensors 146 can determine such position and orientation data using, for example, a camera, accelerometer, magnetometer, gyroscope, etc. The tracking / XR sensors 146 provide the position and orientation data to the XR viewport rendering unit 142 and the 5GS delivery unit 150. The XR client device 140 provides the tracking and sensor information 132 to the XR server device 110 via the network 130. The XR server device 110, in turn, receives the tracking and sensor information 132 and provides it to the XR scene generation unit 112 and the XR viewport prerendering rasterization unit 114. In this manner, the XR scene generation unit 112 can generate scene data for the user's viewport and position, and then use the XR viewport prerendering rasterization unit 114 to prerender 2D media data for the user's viewport. Thus, the XR server device 110 may deliver encoded, pre-rendered 2D media data 134 to the XR client device 140 via the network 130 (eg, using a 5G radio configuration).

[0023] The XR scene generation unit 112 may receive data indicating the type of multimedia application (e.g., the type of video game), the state of the application, multiple user actions, and the like. The XR viewport pre-rendering rasterization unit 114 may format the rasterized video signal. The 2D media encoding unit 116 may be configured with a specific encoder / decoder (codec), a bit rate for media encoding, a rate control algorithm and corresponding parameters, data for slices forming pictures of the video data, low-latency encoding parameters, error resilience parameters, intra-frame prediction parameters, and the like. The XR media content delivery unit 118 may be configured with real-time transport protocol (RTP) parameters, rate control parameters, error resilience information, and the like. The XR media content delivery unit 148 may be configured with feedback parameters, error concealment algorithms and parameters, post-correction algorithms and parameters, and the like.

[0024] Raster-based segmented rendering refers to a scenario in which the XR server device 110 runs an XR engine (e.g., the XR scene generation unit 112) to generate an XR scene based on information from an XR device (e.g., the XR client device 140) and the tracking and sensor information 132. The XR server device 110 may rasterize the XR viewport and perform XR prerendering using the XR viewport prerendering rasterization unit 114.

[0025] exist Figure 1In the example, the viewport is primarily rendered in the XR server device 110, but the XR client device 140 is capable of up-to-date pose correction, for example, using asynchronous timewarp or other XR pose correction to account for changes in pose. The XR graphics workload can be divided into the rendering workload on the powerful XR server device 110 (in the cloud or at the edge) and pose correction (such as asynchronous timewarp (ATW)) on the XR client device 140. Low motion-to-photon latency is maintained via on-device asynchronous timewarp (ATW) or other pose correction methods performed by the XR client 140.

[0026] In some examples, the latency from the XR server device 110 rendering video data and the XR client device 140 receiving such pre-rendered video data can be in the range of 50 milliseconds (ms). The latency for the XR client device 140 to provide position and pose (e.g., posture) information can be lower (e.g., 20 ms), but the XR server device 110 can perform asynchronous time warping to compensate for the latest pose in the XR client device 140.

[0027] The following call flow is an example highlight of the steps involved in executing these techniques:

[0028] The XR client device 140 connects to the network 130 and joins an XR application (eg, executed by the XR scene generation unit 112 ).

[0029] The XR client device 140 sends static device information and capabilities (supported decoders, viewport).

[0030] Based on this information, the XR server device 110 sets the encoder and format.

[0031] cycle:

[0032] The XR client device 140 uses the tracking / XR sensors 146 to collect XR poses (or predicted XR poses).

[0033] The XR client device 140 sends XR pose information (in the form of tracking and sensor information 132 ) to the XR server device 110 .

[0034] The XR server device 110 uses the tracking and sensor information 132 to pre-render the XR viewport via the XR scene generation unit 112 and the XR viewport pre-render rasterization unit 114 .

[0035] The 2D media encoding unit 116 encodes the XR viewport.

[0036] The XR media content delivery unit 118 and the 5GS delivery unit 120 send the compressed media to the XR client device 140 along with data representing the XR pose for its rendering viewport.

[0037] The XR client device 140 uses the 2D media decoder 144 to decompress the video data.

[0038] The XR client device 140 uses the XR pose data provided with the video frames and the actual XR pose from the tracking / XR sensors 146 to improve the prediction, and correct the local pose using, for example, ATW performed by the XR viewport rendering unit 142 .

[0039] According to TR 26.928 clause 4.2.2, the relevant processing and delay components are summarized as follows:

[0040] User interaction latency is defined as the duration between the moment a user action is initiated and the time such action is considered by the content creation engine. In the context of games, this is the time between the moment a user interacts with the game and the moment the game engine processes such player response.

[0041] The content age is defined as the duration between the moment the content is created and the time it is presented to the user. In the context of games, this is the time between when the game engine creates a video frame and when that frame is finally presented to the player.

[0042] Therefore, the round-trip latency is Content Stage and User interaction delay If part of the rendering is done on the XR server and the server produces a frame buffer as the rendering result of the content state, then for raster-based split rendering in cloud gaming applications, the following process causes this delay:

[0043] User interaction latency (gestures and other interactions)

[0044] Capturing user interactions in the game client,

[0045] Delivering user interactions to the game engine, which is then delivered to the server (also known as network latency),

[0046] User interactions are handled by the game engine / server,

[0047] Content Stage

[0048] One or more video buffers are created by the game engine / server (e.g. one for each eye),

[0049] Encode the video buffer into video stream frames,

[0050] Delivering video frames to the game client (also known as network latency),

[0051] The game client decodes the video frames.

[0052] Presenting video frames to the user (also known as frame rate latency).

[0053] When the XR client device 140 applies ATW, the internal processing of the XR client device 140 meets the motion-to-photon latency requirement (up to 20 ms). The network requirements for split rendering are determined by the pose-to-render-to-photon time and the round-trip interaction latency. According to TR 26.928, Section 4.5, the allowed downlink latency is typically 50-60 ms.

[0054] The rasterized 3D scene available in the frame buffer (see clause 4.4 of TR 26.928) is provided by the XR scene generation unit 112 and needs to be encoded, distributed, and decoded. According to clause 4.2.1 of TR 26.928, the relevant format for the frame buffer is 2k by 2k per eye, potentially even higher. The expected frame rate is at least 60 fps, potentially up to 90 fps. The format of the frame buffer is a regular texture video signal, which is then rendered directly. Since the processing is graphics-centric, formats other than the commonly used 4:2:0 signal and YUV signal can be considered.

[0055] For practical purposes, the NVIDIA encoder function can be used. The parameters of this encoder are documented at developer.nvidia.com / nvidia-video-codec-sdk.

[0056] The techniques of this disclosure can be used to address certain challenges and achieve certain tasks. For example, these techniques can be used to evaluate basic system design options and their performance, generate business models for evaluating radio access network (RAN) options, provide guidance on good parameter settings for coding, content delivery, and RAN configuration, identify capacity for this type of application, and define potential optimizations. These techniques can simulate these various elements using reasonable settings.

[0057] In a first example, using segmented rendering, a full simulation of system 100 can be performed. In this example, a 5G New Radio (NR) setup and simulation can be performed for communications via network 130, such as tracking and sensor information 132 and pre-rendered 2D media data 134. Tracking / XR sensors 146 can track and sense example user movements. The quality of the video data presented at display device 152 can then be measured during these various simulations.

[0058] To perform this first example, simulations can be performed using various models that perform separate tasks. The source video model can include actions performed by the XR scene generation unit 112, the XR viewport pre-rendering rasterization unit 114, the XR viewport rendering unit 142, the 5GS delivery unit 120, and the display device 152. The content delivery model can include actions performed by the 2D media encoding unit 116, the XR media content delivery unit 118, the 5GS delivery unit 120, the 5GS delivery unit 150, and the 2D media decoder 144. The uplink model can include actions performed by the tracking / XR sensor 146 and the 5GS delivery unit 150. Tracking and sensor information 132 can be generated as part of the business model uplink, and a RAN simulation can be performed to generate example pre-rendered 2D media data 134.

[0059] The various components of the XR server device 110, the XR client device 140, and the display device 152 may be implemented using one or more processors implemented in circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions attributed to these various components may be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware may be stored on a computer-readable medium and executed by the necessary hardware.

[0060] Based on the system design in S4-200771, there are some challenges for potential simulation and for the generation of business models, as several aspects can be addressed using certain techniques of the present disclosure, for example. These challenges include:

[0061] Evaluate basic system design options and their performance.

[0062] Generate business models for evaluating RAN options.

[0063] Does not violate any content licensing conditions.

[0064] Guidance is provided for good parameter settings regarding encoding, content delivery and also for RAN configuration.

[0065] Identify the capacity used for this type of application.

[0066] Define potential optimizations.

[0067] Simulate this in a reasonable setting in a reproducible manner.

[0068] The simulation and modeling of each system 100 can be broken down into its individual components:

[0069] Source video model

[0070] Encoding and content delivery

[0071] Radio Access Delivery

[0072] Content delivery receivers and decoders

[0073] Display process

[0074] Uplink service

[0075] Figure 2 is a diagram showing the technique according to the present disclosure for simulating and measuring the Figure 1 A conceptual diagram of an example breakdown of components of computing system 100 performing a split rendering process for performance of XR server device 110 . Figure 2 The example decomposition of corresponds to the first example discussed above. In this example, the system 160 includes a game engine 162, a model encoding device 164, a content encoding and delivery model 166, a 5GS simulation unit 168, and a content delivery and decoding model 170. Thus, the game engine 162 may correspond to Figure 1 The XR scene generation unit 112 and the XR viewport pre-rendering rasterization unit 114, the model encoding device 164 may correspond to Figure 1 2D media encoding unit 116, the content encoding and delivery model 166 may correspond to Figure 1 The XR media content delivery unit 118 and the 5GS delivery unit 120, the 5GS simulation unit 168 may correspond to Figure 1 The network 130 and the content delivery and decoding model 170 may correspond to Figure 1 5GS delivery unit 150, XR media content delivery unit 148, 2D media decoder 144 and XR viewport rendering unit 142.

[0076] The game engine 162 receives data 178 (including gesture models and tracking data, game data, and game / engine configuration data) for simulating an extended reality (XR) game using the received gesture models and tracking data. The game engine 162 generates video frames 180 based on the received data and provides the video frames 180 to the model encoding device 164. The model encoding device 164 encodes the video frames 180 to generate encoded video frames and v-tracking data, and provides the encoded video frames and v-tracking data 182 to the content encoding and delivery model 166.

[0077] The content coding and delivery model uses the encoded video frames and v-track data 182 to generate slice tracking (s-track) data 184 and packets representing encapsulated, encoded slices of the encoded video frames using example configuration parameters 188. The content coding and delivery model 166 provides the s-track data 184 and packets to the 5GS simulation unit 168. The 5GS simulation unit 168 uses the received s-track data 184 and packets, along with configuration parameters 190, to perform a RAN simulation to generate s'-track data 186. The s-track data 184 represents slices of the video data before over-the-air transmission, while the s'-track data 186 represents slices of the video data after over-the-air transmission (e.g., which may be corrupted due to over-the-air transmission). The 5GS simulation unit 168 provides the s'-track data 186 and packets of the delivered encoded slices of the encapsulated video data to the content delivery and decoding model 170. The content delivery and decoding model 170 performs an example decoding process on the packetized data.

[0078] Using the encoded and decoded data from the content encoding and delivery model 166 and the content delivery and decoding model 170, a value representing per-frame quality 172 may be generated to test the performance of the RAN simulation performed by the 5GS simulation unit 168 and configuration parameters 190. Using the combination / aggregation of the individual values ​​of per-frame quality 172, an overall quality value 174 may be generated.

[0079] Targets can be set for per-frame quality 172 and overall quality 174 values ​​for adjusting various configuration parameters. In one example, the goal can be to maintain no change in video quality as measured by per-frame quality 172 and overall quality 174. One or a small number of application metrics can be derived from the RAN simulation performed by the 5GS simulation unit 168. For each scenario, a small number (e.g., one to three) of v-traces can be performed, but over a duration of several hours. The quantization parameter can be constant in configuration parameters 188. Several different values ​​for configuration parameters 188 can be adjusted, which can cause bitrate to change as a side effect and result in changes to service characteristics. However, no quality change should occur. A relatively small number of configurations can be tested, e.g., three to ten.

[0080] The game engine 162, model encoding device 164, content encoding and delivery model 166, 5GS simulation unit 168, and content delivery and decoding model 170 can be implemented using one or more processors implemented in circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions attributed to these various components can be implemented using hardware, software, or firmware. When implemented using software or firmware, it should be understood that the instructions for the software or firmware can be stored on a computer-readable medium and executed by the necessary hardware.

[0081] Based on the concepts discussed above, several parameters are relevant to the overall system design:

[0082] game:

[0083] Game Type

[0084] Game state,

[0085] Multi-user actions, etc.

[0086] User Interaction:

[0087] 6DOF pose based on head and body movement,

[0088] Game interaction via controller

[0089] The format of the rasterized video signal. Typical parameters are:

[0090] 1.5K x 1.5K per eye at 60, 90, 120 fps,

[0091] 2K x 2K at 60, 90, 120 fps

[0092] YUV 4:2:0 or 4:4:4

[0093] Encoder configuration

[0094] Codec: H.264 / AVC or H.265 / HEVC

[0095] Bitrate: The bitrate is set to a specific value (e.g., 50 Mbit / s)

[0096] Rate control: CBR, capped VBR, feedback-based, CRF, QP

[0097] Slice settings: 1 per frame, 1 per MB line, X per frame

[0098] Intra-frame settings and error recovery: conventional IDR, GDR mode, adaptive intra, feedback-based intra, feedback-based prediction and ACK-based, feedback-based prediction and NACK-based

[0099] Delay setting: P picture only, front view unit

[0100] Complexity settings for encoders

[0101] Content Delivery

[0102] Slice to IP Mapping: Segmentation

[0103] RTP-based timecode and packet numbering

[0104] Feedback ACK / NACK based on RTP / RTCP

[0105] RTP / RTCP-based feedback on bitrate

[0106] 5G system / RAN configuration:

[0107] QoS settings (5QI): GBR, latency, and loss rate

[0108] HARQ transmission, scheduling, etc.

[0109] Content delivery receiver configuration:

[0110] Loss Detection: Serial Number

[0111] Delay / latency processing

[0112] Error recovery

[0113] ATW

[0114] Quality

[0115] Video quality (encoded)

[0116] Lost data

[0117] Immersion

[0118] Video quality can be affected by various factors. One factor is decoding artifacts based on encoding. For example, such artifacts can be determined by the Peak Signal-to-Noise Ratio (PSNR). Another factor is artifacts caused by lost packets and the resulting error propagation.

[0119] According to the techniques of this disclosure, quality can be modeled for each simulation as a combination of:

[0120] The average PSNR from the encoding, and

[0121] Percentage of damaged video area

[0122] Among them, the damaged macroblock is defined as

[0123] If it is part of a slice that is missing for this transmission

[0124] If it is received correctly, but it is predicted from the wrong macroblock

[0125] The macroblock (or decoding tree unit) is correct

[0126] If it is received correctly and it predicts for an undamaged macroblock

[0127] Prediction based on undamaged macroblocks means

[0128] The spatial predictions are correct

[0129] The time prediction is correct

[0130] Recovering the MB is done by intra refresh or by predicting again from the correct MB.

[0131] Generally speaking, references to macroblocks above and throughout this disclosure refer to macroblocks of ITU-T H.264 / AVC. However, it should be understood that the concept of macroblocks can be replaced with coding units (CUs) or coding tree units (CTUs) of ITU-T H.265 / HEVC or ITU-T H.266 / VVC without loss of generality.

[0132] exist Figure 2 In the example, a source model can be provided based on the inputs describing the test game, a gesture model, and a statistical model of the video signal that can be used to apply the content delivery and encoding processes. Based on these traces, 3GPP SA4 can define a content coding and delivery model that takes into account the system design based on TR 26.928 and as documented in S4-200771. The system model can include encoding, delivery, decoding, and also quality definitions. The system can also provide an interface for RAN simulation. The RAN team can then use this model to simulate different traffic characteristics and evaluate different performance options.

[0133] System 160 and other systems described herein may use the following elements:

[0134] V-tracking: Video tracking that provides enough information from the encoding model to understand the complexity and data rate for the encoding model, and also takes into account delivery options such as bitrate control, intra refresh, feedback-based error recovery, etc. Details are TBD, but we are currently investigating parameters provided by the model encoded from x265 encoding runs.

[0135] S-track: A sequence of slices that are the output of the model encoder. Each slice has been assigned the following:

[0136] Associated frame (timestamp)

[0137] Slice size

[0138] The quality of the slice (possibly more information from V-tracking, such as complexity)

[0139] Number of macroblocks (area covered)

[0140] The timing of the slice, e.g., the cutoff time or at least the origin timestamp.

[0141] S'Trace:

[0142] All information from S-Tracking

[0143] Slice Lost Indicator

[0144] Slice delay information

[0145] Figure 3 It is about Figure 2 Example graphs for various RAN simulations discussed. In this example, the graphs show the percentage of corrupted video as a function of the number of users supported. Figure 3 The graph of represents an example graph that can be generated from the above example, where the quantization parameter is constant to prevent changes in video quality. Various configuration parameters 190 can be modified to allow for each number of supported users. As the number of supported users increases, the amount of corrupted video data is expected to increase. However, a configuration data set can be identified that allows for a relatively large number of supported users and a relatively small percentage of corrupted video, such as in Figure 3 "config 3" shown in .

[0146] Reference again Figure 2 In another example, video quality (as measured by per-frame quality 172 and overall quality 174) can be allowed to vary. In this example, one or a small number of application metrics can be derived from the RAN simulation performed by the 5GS simulation unit 168. For each scenario, a small number (e.g., one to three) of v-traces can be performed, but over a duration of several hours. Various values ​​for configuration parameters 188 (including quantization parameters) can be adjusted, which can result in bitrate changes as a side effect and changes to service characteristics. However, no quality changes should occur. A relatively small number of configurations can be tested, e.g., three to ten.

[0147] Figure 4 and Figure 5 It is about Figure 2 Example graphs for various RAN simulations discussed. These example graphs correspond to the above examples where the quantization parameter is allowed to vary. In these examples, the graphs plot the delivered quantizer / video quality data for a fixed number of users in various scenarios versus the percentage of corrupted video. Figure 5 As shown, the number of users and configuration parameters can be selected to achieve high video quality, low bit rate, and a low percentage of lost data (e.g., at 0.1% corrupted video).

[0148] To perform the various test scenarios discussed above, representative source data can be generated to perform v-trace. S-trace can also be performed using content encoding, delivery, and quality, including, for example, various system designs, delivery parameters, mapping to radio, modeling, and quality definitions. RAN delivery simulation can be based on the S-trace mentioned above.

[0149] Figure 6 is a block diagram illustrating an example v-tracking system model 200 according to techniques of the present disclosure. The v-tracking system model 200 represents a model for evaluating the coding complexity of a source video. In this example, the v-tracking system model 200 includes a game engine 202, a prediction model encoding unit 204, an intra-frame model encoding unit 206, and a tracking combination unit 208. The v-tracking system model 200 may generally correspond to Figure 2 For example, the game engine 202 may generally correspond to Figure 2 The game engine 162 of FIG. 1 may include the prediction model encoding unit 204, the intra model encoding unit 206, and the tracking combination unit 208, and may correspond to the model encoding device 164. The game engine 202 receives input data 178, which includes pose and model tracking data, game data, and game configuration data.

[0150] Game engine 202 generates video frame 210 and provides video frame 210 to each of prediction model encoding unit 204 and intra model encoding unit 206. Video frame 210 may be, for example, extended reality (XR) segmentation data at 60 fps, with one frame per eye to achieve XR effects. In this example, v-tracking system model 200 includes two tracks: one for intra prediction and one for inter prediction. In video coding, intra prediction is performed by predicting a block of video data using adjacent, previously decoded blocks of the same frame, and inter prediction is performed by predicting a block of video data using reference blocks of previously decoded frames.

[0151] The prediction model encoding unit 204 can generate VP-track data 212, and the intra-frame model encoding unit 206 can generate VI-track data 214. The track combination unit 208 receives the VP-track data 212 and the VI-track data 214 and generates V-track data 216. The V-track data 216 can be formatted according to FFMPEG's passlogfile format after the encoding pass. Such a format is described in, for example, ffmpeg.org / wiki / Encode / H.264 and slhck.info / video / 2017 / 03 / 01 / rate-control.html. The track combination unit 208 can also use the configuration parameters "-pass[:stream_specifier] n (output,per-stream)" and "-passlogfile". Pseudo code for an example algorithm for generating v-track data 216 in this format is shown below:

[0152] void ff_write_pass1_stats (MpegEncContext *s) {

[0153] snprintf(s->avctx->stats_out, 256,

[0154] "in:%d out:%d type:%dq:%d itex:%d ptex:%d mv:%d misc:%d"

[0155] "fcode:%d bcode:%d mc-var:%"PRId64" var:%"PRId64" icount:%dskipcount:%d hbits:%d;\n",

[0156] s->current_picture_ptr->f->display_picture_number,

[0157] s->current_picture_ptr->f->coded_picture_number,

[0158] s->pict_type,

[0159] s->current_picture.f->quality,

[0160] s->i_tex_bits,

[0161] s->p_tex_bits,

[0162] s->mv_bits,

[0163] s->misc_bits,

[0164] s->f_code,

[0165] s->b_code,

[0166] s->current_picture.mc_mb_var_sum,

[0167] s->current_picture.mb_var_sum,

[0168] s->i_count, s->skip_count,

[0169] s->header_bits);

[0170] }

[0171] As an alternative to ITU-T H.264 / AVC, ITU-T H.265 / HEVC or other video coding standards can be used to decode video data. For an example of H.265, an example log file is described at x265.readthedocs.io / en / default / cli.html#input-output-file-options. Such a log file may include the following parameters:

[0172] Coding order: The order in which frames are encoded by the encoder.

[0173] Type: The slice type of the frame.

[0174] POC: Picture Order Count - the order in which frames are displayed.

[0175] QP: Quantization parameter determined for each frame.

[0176] Bits: The number of bits consumed by the frame.

[0177] Scenecut: This is 1 if the frame is a scene cut, 0 otherwise.

[0178] RateFactor: Applicable only when CRF is enabled. The rate factor depends on the CRF given by the user. This is used to determine the QP in order to target a specific quality.

[0179] BufferFill: The bits available for the next frame. This includes bits carried over from the current frame.

[0180] BufferFillFinal: Buffer bits available after the frame is removed from the CPB.

[0181] Latency, expressed in terms of the number of frames between when the frame is input and when it is output.

[0182] PSNR: Peak signal-to-noise ratio for the Y, U, and V planes.

[0183] SSIM: A quality metric representing the structural similarity between frames.

[0184] Reference List: The POC of the reference pictures in Lists 0 and 1 for the frame.

[0185] When --csv log level is greater than or equal to 2, several statistics about the encoded bitstream and encoder performance are available.

[0186] In still other examples, other video codecs may be used, such as ITU-T H.266 / Versatile Video Coding (VVC). Similar datasets may be generated for analysis using VVC or other such video codecs.

[0187] The track combination unit 208 can generate v-track data 216 according to certain v-track generation principles. Certain v-track generation principles are described at slhck.info / video / 2017 / 02 / 24 / crf-guide.html. The track combination unit 208 can use H.264 / AVC using x264 in FFMPEG and H.265 / HEVC using x265 in FFMPEG. The track combination unit 208 can use a near-lossless constant quality mode with options including:

[0188] Constant QP

[0189] Constant rate factor

[0190] According to slhck.info / video / 2017 / 02 / 24 / crf-guide.html, the recommendations for QP / CRF settings are:

[0191] For x264, typical values ​​are between 18 and 28. The default value is 23. 18 should be visually transparent.

[0192] For x265, typical values ​​are between 24 and 34. The default is 28. 24 should be visually transparent.

[0193] A change of ±6 should result in half / double the file size, although results may vary.

[0194] Multiple CRFs can be run in parallel from one output.

[0195] It may be possible to determine whether the discussion at slhck.info / video / 2017 / 02 / 24 / crf-guide.html also applies to XR and video games.

[0196] The game engine 202, prediction model encoding unit 204, intra-frame model encoding unit 206, and tracking combination unit 208 can be implemented using one or more processors implemented in circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions attributed to these various components can be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware can be stored on a computer-readable medium and executed by the necessary hardware.

[0197] Figure 7 is a block diagram illustrating an example system 220 for performing RAN simulation according to techniques of this disclosure. The system 220 includes a content encoding and delivery model 222, a slice to Internet Protocol packet / radio link control (IP / RLC) mapping unit 224, a RAN simulation unit 226, and an IP / RLC to slice mapping unit 228. The system 220 may generally correspond to Figure 2 For example, the content encoding and delivery model 222 may correspond to Figure 2 The content coding and delivery model 166 of the embodiment of the present invention, and the slice to IP / RLC mapping unit 224, the RAN simulation unit 226 and the IP / RLC mapping unit 228 may correspond to Figure 2 168 5GS analog units.

[0198] In this example, the content encoding and delivery model 222 is derived from the tracking combination unit 208 ( Figure 6 ) receives v-tracking data 216, which includes data representing the encoding of frames of video data. The content encoding and delivery model 222 also receives global configuration data 230. Using the global configuration data 230 (which may correspond to Figure 2), the content coding and delivery model 222 may encode frames of video data to form coded slices of the video data, and form s-track data 232 representing the coded slices. The content coding and delivery model 222 may generate slices for a particular time, where each slice corresponds to a sequence of macroblocks (ITU-T H.264 / AVC) or a largest coding unit / coding tree unit (ITU-T H.265 / HEVC).

[0199] The slice-to-IP / RLC mapping unit 224 can use the Real-time Transport Protocol (RTP) to form IP packets and RLC segments from the slices of the s-trace data 232. The slice-to-IP / RLC mapping unit 224 can be configured with IP packet and payload size information for forming packets and segments. Typically, if a packet is lost, the entire slice corresponding to the packet is also lost. The slice-to-IP / RLC mapping unit 224 generates p-trace data 234 and provides the p-trace data 234 to the RAN simulation unit 226. The RAN simulation unit 226 can be configured with delay and loss requirements for each slice / packet and, through RAN simulation, determine whether each slice is received or lost, as well as the delay value or received timestamp value for each slice at the end of the simulation.

[0200] The RAN simulation unit 226 can provide P'-tracking data 236, including data indicating whether a slice was received or lost, and the latency / reception time for the received slice, to the IP / RLC to slice mapping unit 228. The IP / RLC to slice mapping unit 228 provides the s'-tracking data 238 to the content coding and delivery model 222. The content coding and delivery model 222 can determine an overall quality value 240 by comparing the video data decoded before the RAN simulation with the video data decoded after the RAN simulation.

[0201] The RAN simulation unit 226 can be configured based on the maximum latency requirements for each slice downlink. For example, there may be a 10 ms MAC-to-MAC deadline and a zero-loss requirement from the outside. For the uplink, there may be tracking, sensor and gesture information, content delivery uplink information, and gesture update traffic (e.g., with a period of 1.25 ms or 2 ms). The system can determine how gesture frequency affects quality, which may depend on the type of XR service. For gaming, gesture frequency may have a significant impact. Gestures can also be synchronized with game actions.

[0202] The RAN simulation options may include open-loop and closed-loop options. For open-loop, the system may generate s-trace data, p-trace data, one-way RAN simulation, p'-trace data, and s'-trace data for the entire v-trace, and then perform a quality assessment (e.g., in the form of an overall quality value 240). For closed-loop, the system may generate s-trace data, p-trace data, one-way RAN simulation, p'-trace data, and s'-trace data for each v-trace item, and then feed the resulting data into the next s-trace, and generate a quality assessment in the form of an overall quality value 240.

[0203] The content coding and delivery model 222, the slice-to-IP / RLC mapping unit 224, the RAN simulation unit 226, and the IP / RLC-to-slice mapping unit 228 can be implemented using one or more processors implemented in circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions attributed to these various components can be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware can be stored on a computer-readable medium and executed by the necessary hardware.

[0204] Figure 8 is a block diagram illustrating an example system 250 for measuring the performance of XR content delivery configurations and modeling according to techniques of this disclosure. System 250 includes a video content model 252, a content encoding and delivery model 254, a radio access network (RAN) simulation unit 256, a content delivery and decoding model 258, and a content delivery and decoding model 260. These elements of system 250 may correspond to Figure 2 For example, the video content model 252 may correspond to the corresponding component of the system 160. Figure 2 The game engine 162 and the model encoding device 164, the content encoding and delivery model 254 may correspond to Figure 2 The content encoding and delivery model 166 of RAN simulation unit 256 may correspond to the 5GS simulation unit 168, and the content delivery and decoding model 258 may correspond to Figure 2 Content delivery and decoding model 170.

[0205] In this example, the decoded media data is compared before passing through the RAN simulation unit 256 and after having passed through the RAN simulation unit 256. The video data from the content delivery and decoding model 258 (which has passed through the RAN simulation unit 256) is evaluated to determine a value for per-frame quality 262. The decoded media data and the value for per-frame quality 262 from the content delivery and decoding model 260 are used to determine overall quality data 264.

[0206] Specifically, the video content model 252 provides v-tracking data 268 and the encoded video data to be sent to the content encoding and delivery model 254. The content delivery model 254 receives global configuration data 266 and uses the global configuration data 266 to encode and deliver the video data corresponding to the v-tracking data 268. The global configuration data 266 may include, for example, bitrate control information (e.g., for selecting one or more of constant quality, constant bitrate, feedback-based variable bitrate, or constant rate factor), slice settings (number of slices, maximum slice size), error recovery (frame-based, slice-based, regular intra-frame refresh, feedback-based intra-frame refresh, or feedback-based prediction), and feedback data (off, statistics, or action). The maximum slice size setting may depend on various statistics that may vary, including the number of slices.

[0207] The content encoding and delivery model 254 outputs s-track data 270 to the RAN simulation unit 256 and outputs s'-track data 270' to the content delivery and decoding model 260. The RAN simulation unit 256 performs RAN simulation on the s-track data 270 and outputs s'-track data 272 to the content delivery and decoding model 258. The content delivery and decoding model 260 outputs q-track data 274, and the content delivery and decoding model 258 outputs a value for each frame quality 262 to produce q-track data 276. The overall quality data 264 can be one or more quality metrics for the entire video sequence (e.g., an aggregation of the quality of individual video frames).

[0208] Various content coding considerations may influence the encoding of video data corresponding to v-track data 268 to form s-track data 270. In one example, the goal is consistent quality, using ITU-T H.265 / HEVC encoding, a fixed number of slices (one slice is encoded in opportunistic mode using intra prediction), and no feedback. In this example, bitrate control may use a constant rate factor of 28, a number of slices of 10, a periodic intra refresh value of 10, periodic intra refresh and reference picture deactivation, and feedback disabled. To create the slice sizes, the content coding and delivery model 254 may generate 10 slices for each video frame, using a constant rate factor of 28 (where ±6 results in approximately half / double the frame size, for each + / - 12% decrease / increase), assuming one intra slice based on the VI-track (i.e., 10% of the adjusted I-frame size), and generating nine inter slices based on the VP-track (where two aspects may be considered: frame size (10%) and statistical variation).

[0209] The RAN simulation unit 256 can maintain status data for each macroblock (ITU-T H.264 / AVC) or coding tree unit (ITU-T H.265 / HEVC), such as whether it is damaged (lost or predicted based on damaged areas, whether temporal or spatial), or correct. The RAN simulation unit 256 can provide feedback to the content coding and delivery model 254, including data indicating the number of lost or damaged blocks of video data. The content coding and delivery model 254 can then incorporate the feedback for subsequent encoding. For example, the content coding and delivery model 254 can react to slice loss, either statistically or operationally. In response, the content coding and delivery model 254 can adjust the bitrate (e.g., obtain the encoding bitrate and adjust the quantization parameter upward or downward, where + / - 1 can have a 12% effect). If a slice is lost, the content coding and delivery model 254 can add an intra-predicted slice (significantly more intra-prediction data can be added in the case of a reported loss). Depending on motion vector activity, the intra-predicted slice can cover a large area. In some cases, the content coding and delivery model 254 may make predictions based only on confirmed regions, which may result in a statistical increase in frame size for lost slices because the latest slices may not be used for prediction.

[0210] The s-track data 270 may be formatted to include data for: a timestamp representing the associated frame, the size of the slice, the quality of the slice (which may include more information from the v-track, such as complexity), the number of macroblocks (for ITU-T H.264 / AVC) or coding tree units (for ITU-T H.265 / HEVC), and / or the timing of the slice (e.g., a deadline for slice reception). The s'-track data 270' and 272 may be formatted to include all of the information from the format for the s-track data 270 discussed above, as well as additional data indicating slice loss and / or slice delay.

[0211] There can be a constant cutoff time for each slice, which can be set to a high number. For example, for 60 fps video data staggered for each eye, i.e. 120 fps video data, there can be a desired delay set to 7 ms. The discard cutoff time can be set in different ways and can be useful, but should not complicate the RAN simulation.

[0212] Content decoding considerations used to generate Q-trace data 274 from s'-trace data 270' can be used to take results from the RAN simulation performed by the RAN simulation unit 256 and map these simulation results to an estimate for per-frame quality 262. When determining the value for per-frame quality 262, slice considerations can include, for example, the quality of the received encoded slices, the number of lost slices (which slices are degraded in quality to the point of being considered lost), and delayed slices (which can be considered lost). Other considerations can include error propagation (e.g., prediction of slices from degraded or lost slices), frame / slice complexity, which can determine error propagation (how quickly incorrect data propagates within a frame), new intra-frame reset quality, and the percentage of correct / erroneous data for each frame and for different configurations.

[0213] A metric used to determine the q-tracking data 274 may be the percentage of incorrect data, which may affect loss / error propagation. The q-tracking data 274 format may include data indicating the encoding quality, the resulting quality after delivery (percentage of degraded data), and / or the data rate of the frame.

[0214] System 250 can maintain a state for each macroblock (ITU-T H.264 / AVC) or coding tree unit (CTU) (ITU-T H.265 / HEVC) for either "impaired" or "correct." If the macroblock / CTU is impaired, system 250 can determine whether the macroblock / CTU is part of a slice that was lost for that transmission, or whether the macroblock / CTU was correctly received by predicting from an incorrectly decoded / lost area of ​​another slice / frame. If the macroblock / CTU is correct, system 250 determines that the macroblock / CTU was correctly received and was predicted from an unimpaired area of ​​another slice / frame. Predicting from an unimpaired area of ​​another slice / frame means that the spatial prediction is correct, the temporal prediction is correct, or that the macroblock / CTU was correctly recovered using intra refresh and prediction from a correct area of ​​another slice / frame after the intra refresh.

[0215] System 250 can determine an overall quality value 264 by averaging the encoding quality (e.g., averaging over quantization parameters and possibly converting to peak signal-to-noise ratio (PSNR)), averaging incorrect video data (e.g., averaging the number of incorrect video frames and / or slices), and / or (if necessary) multiplying / combining / aggregating. 3GPP SA4 defines a coding model based on a constant rate factor (CRF) with a specific quality factor (e.g., using the FFMPEG default of 28). Different configurations for error resilience can be applied. A criterion for simulation can be the percentage of incorrect video areas, such that at most one macroblock / CTU is erroneous for every X (e.g., 60) seconds. For example, if 4096x4096 is used at 60 fps, this results in an average of 10e-6 damaged areas.

[0216] The models discussed above can be applied to pixel-based segmented rendering and / or cloud gaming. Conversational applications, video streaming applications (e.g., Twitch.tv), and other applications can also be used.

[0217] Video content model 252, content encoding and delivery model 254, RAN simulation unit 256, content delivery and decoding model 258, and content delivery and decoding model 260 can be implemented using one or more processors implemented in circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions attributed to these various components can be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware can be stored on a computer-readable medium and executed by the necessary hardware.

[0218] System 250 represents an example content modeling based on the discussion in S4-200771. The modeling includes v-model inputs, global configuration for encoders, statistical or dynamic feedback from content delivery receivers, decoding models, and quality models.

[0219] The content encoding performed by the content encoding and delivery model 254 can be modeled as follows:

[0220] For frame i from V-tracking (based on sample time)

[0221] Read frame i (timestamp) from V-track

[0222] Read the latest dynamic information from the dynamic status information

[0223] Encode the model (based on input parameters)

[0224] For slices s=1, 2, …, S

[0225] Discards the slice s with associated parameters

[0226] New slices become available to create IP packets

[0227] For IP packets p=1, 2, …, P

[0228] Drop the IP packet with associated parameters and with timestamp to S-Trace 1 (and with parameters such as slice number)

[0229] Configuration parameters of the global configuration data 266 may include:

[0230] Input: Global Configuration

[0231] Bitrate control: constant quality, constant bitrate, variable bitrate based on feedback, constant rate factor 28

[0232] Slice settings (number of slices 10, maximum slice size)

[0233] Error resilience, frame-based, slice-based, periodic intra refresh10, feedback-based intra refresh, feedback-based prediction (NVIDIA: periodic intra refresh, reference picture invalidation)

[0234] Input: Dynamic per-slice information

[0235] closure

[0236] Statistics of bit rate or loss at some delay

[0237] Operating at a loss of bit rate or with some delay

[0238] Aspects of model encoding can include:

[0239] Effects of QP settings and intra-frame ratio and slice settings

[0240] feedback

[0241] Bitrate adjustment: The encoder obtains the encoding bitrate and adjusts the QP (see + / -1 è 12%)

[0242] Added intras in case of lost slices: Significantly more intras in case of reported loss. Intras cover large areas (depending on motion vector activity)

[0243] Prediction based on ACKs only: Statistical increase in frame size for lost slices, since the latest slice cannot be used

[0244] Slice Settings

[0245] The decoding simulation performed by the content delivery and decoding model 258 can be based on the delay and / or loss of slices. Delayed and lost slices can be considered unusable and cause errors (e.g., when decoding subsequent frames / slices that reference the lost or delayed slice for inter-frame prediction).

[0246] The RAN simulation unit 256 can simulate RAN simulation based on the existing 5Qis.

[0247] Quality assessment can be based on two aspects, including coding quality and quality degradation due to lost slices. The following simulation can be used to identify damaged macroblocks (or coding tree units or coding units in ITU-T H.265 / HEVC):

[0248] Maintain state for each macroblock (or CU / CTU)

[0249] damaged

[0250] correct

[0251] Macroblocks are damaged

[0252] If it is part of a slice that is missing for this transmission

[0253] If it is received correctly, but it is predicted from the wrong macroblock

[0254] The macroblocks are correct

[0255] If it is received correctly and it predicts for an undamaged macroblock

[0256] Prediction based on non-damaged macroblocks means

[0257] The spatial predictions are correct

[0258] The time prediction is correct

[0259] The MB is recovered by intra refresh and prediction is performed again based on the correct MB.

[0260] Depending on the configuration and settings of the delivered video quality, different results can be obtained. For example, the quality threshold can be a video area with a maximum of 0.1% impairment. In addition, the quality of the original content can be a threshold.

[0261] Figure 9 is a block diagram illustrating an example pipeline 280 for generating v-tracking data in real time according to techniques of this disclosure. Pipeline 280 includes a game engine 282, an encoder 284, a decoder 286, a raw data storage 288, a predictive encoder 290, an intra encoder 292, and a v-tracking to 3GPP unit 294. The components of pipeline 280 may correspond to Figure 2 For example, the game engine 282 may correspond to the corresponding components of the system 160. Figure 2 The game engine 162, the encoder 284 may correspond to Figure 2 The model encoding device 164, and the decoder 286, the original data 288, the prediction encoder 290, the intra encoder 292 and the v-tracking to 3GPP unit 294 may correspond to Figure 2 Content encoding and delivery model 166.

[0262] The game engine 282 receives input data 296, which includes pose models and tracking data, game data, and game configuration data. The game engine 282 generates video frames 298 based on the input data 296. The encoder 284, which may be an NVIDIA NVENC encoder, uses the configuration data 302 to encode the video frames 298 to generate a video bitstream 300. The encoder 284 can operate at 600 gigabytes per hour and can run for several hours. The decoder 286, which may be an FFMPEG decoder, decodes the video bitstream 300 to generate decoded video data, which is stored as raw data 288 on a computer-readable medium (such as a hard drive, flash drive, etc.). The predictive encoder 290 and the intra encoder 292 receive the configuration data 304 and encode the raw data 288 to form corresponding VP-track and VI-track data. The v-track to 3GPP unit 294 generates v-track data based on the VP-track and VI-track data.

[0263] like Figure 9As shown, as discussed above, the encoder 284 can generate a bitstream 300 and the V-track to 3GPP unit 294 can generate v-track data. V-track generation can be initiated by using the high-quality output of the game engine 282 and encoded by the encoder 284 using ITU-T H.264 (AVC) for model gaming and repeatable tracking. The decoder 286 can decode this information into 2K x 2K source frames at 120 fps. This sequence can be sent through a model encoder including a predictive encoder 290 and an intra encoder 292 to identify the impact of different parameter settings. The following x.265 parameters can be used initially:

[0264] --input viki.yuv (4:2:0)

[0265] --profile, -P main10

[0266] --input-res 2048x2048

[0267] --fps 120fps, 60fps, 30 fps

[0268] --psnr

[0269] --ssim

[0270] --frames 500 (initial testing)

[0271] --bframes 0

[0272] --crf 22, 25, 28, 31, 34

[0273] --csv logfile.csv

[0274] --csv-log-level 2

[0275] --log-level 4

[0276] --numa-pools "8"

[0277] --keyint, -I 1 and -1

[0278] --slices 1, 128 (2048), 8 (for 2048)

[0279] --output rvrviki.h265

[0280] --rc-lookahead 0 / 1

[0281] Based on these 180 encoding runs, good encoder modeling is expected such that for longer game sequences only a subset of runs is necessary.

[0282] The parameters used to generate V-tracking data in this example may include the duration of P-tracking (e.g., a number of hours), the various games being executed by the game engine 282, gesture tracking, settings for the encoder 284 (where the goal is, for example, simple, high-quality FFMPEG decodable video data), and FFMPEG configuration data. Four FFMPEG decoder configurations may be tested and, in some examples, pipelined in parallel:

[0283] HEVC under CRF 28 and I-PPP

[0284] HEVC under CRF 28, I-IIIIII

[0285] AVC with CRF 23, I-PPP

[0286] AVC with CRF 23, I-IIIIII

[0287] The possible configuration parameters are described at x265.readthedocs.io / en / default / cli.html#. These parameters can be used to configure rvrplugin.ini with the following: 2048x2048 resolution, 5 fps, and a bitstream generated on the server using the ITU-T H.264 / AVC encoder (which can run at 60 or 120 fps). This data can be re-encoded using CRF-24 and updated PSNR and SSIM values. Initial command line parameters can include the following:

[0288] ABR (average bitrate): x265.exe --input viki.yuv --input-res 1440x1440 --fps120 --psnr --ssim --frames 500 --bframes 0 --bitrate 600000 --csv logfile.csv --csv-log-level 2 --log-level 4 --numa-pools "8" --keyint 1 --outputrvrviki.h265

[0289] CRF (Constant Quality Variable Bitrate): x265.exe --input viki.yuv --input-res1440x1440 --fps 120 --psnr --ssim --frames 500 --bframes 0 --crf 24 --csvlogfilecrf.csv --csv-log-level 2 --log-level 4 --numa-pools "8" --keyint 1 --output rvrviki.h265

[0290] The game engine 282, encoder 284, decoder 286, predictive encoder 290, intra encoder 292, and v-tracking to 3GPP unit 294 can be implemented using one or more processors implemented in circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions attributed to these various components can be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware can be stored on a computer-readable medium and executed by the necessary hardware.

[0291] Figure 10 3 is a block diagram illustrating an example rendering and encoding system 310 according to techniques of the present disclosure. In this example, system 310 includes a renderer 312 and an encoder 314. Renderer 312 receives content changes 318 and pose information 316 to generate rendered video data (e.g., at a frame rate of 60 or 120 fps). Pose information 316 can be provided at a very high frequency, with a 2.5 ms latency. Every 1 / 60 Hz, renderer 312 can extract the latest pose from pose information 316, render video data based on content changes 318, and send the rendered image to encoder 314 for encoding. Encoder 314 can queue encoding calls and then render the image.

[0292] The encoder 314 may encode the video data using ITU-T H.265 / HEVC (e.g., "x265.exe"). The encoder 314 may operate with the following command line options:

[0293] --input viki.yuv (4:2:0)

[0294] --profile, -P main10

[0295] --input-res 1440x1440 (2048x2048)

[0296] --fps 120fps, 60fps, 30

[0297] --psnr

[0298] --ssim

[0299] --frames 500 (initial testing)

[0300] --bframes 0

[0301] --crf 22, 25, 28, 31, 34

[0302] --csv logfile.csv

[0303] --csv-log-level 2

[0304] --log-level 4

[0305] --numa-pools "8"

[0306] --keyint, -I 1 and -1

[0307] --slices 1, 90 (for 1440), 128 (2048), 6 (for 1440), 8 (for 2048)

[0308] --output rvrviki.h265

[0309] --analysis-save-reuse-level 10

[0310] --rc-lookahead 0 / 1

[0311] The encoder 314 does not need to use "-frame dup disabled", "--consrained-intra", and "--no-deblock".

[0312] The renderer 312 and encoder 314 can be implemented using one or more processors implemented in circuits, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. The functions attributed to these various components can be implemented in hardware, software, or firmware. When implemented in software or firmware, it should be understood that the instructions for the software or firmware can be stored on a computer-readable medium and executed by the necessary hardware.

[0313] Figure 11 is a graph showing an example comparison between average PSRN and PSNR (in decibels (dB)) for various encoder parameters. In this example, there is an increase of approximately 25% from 0 to 1, approximately 50% from 0 to 2, approximately 75% from 0 to 3, and approximately 100% from 0 to 4. This increase can depend on the picture type. For example, if the video sequence generates many intra-predicted frames, the increase may be smaller regardless of the configuration, and very static content may result in a smaller increase. This can be used to model ineffective reference operations.

[0314] Error propagation can also be modeled. As a basic principle, an error in a slice may destroy a slice in the current picture (e.g., X% is damaged). Since the next frame references this frame, the error may propagate both temporally and spatially until an intra frame is received, or an intra slice for an area that may be damaged, and the error may propagate spatially to the next frame. Unless intra is applied, this may depend on the number of motion vectors (e.g., in the next frame) if more than X% is damaged. In the next frame, the size of the damaged area may increase. Intra prediction can remove error propagation, but using multiple reference frames can restore the damaged area.

[0315] Figure 12 is a flowchart illustrating an example method of processing media data according to the techniques of this disclosure. Figure 1 The system 100 (and specifically, the XR server device 110 ) explains Figure 12 However, it should be understood that other devices may be configured to perform this method or similar methods.

[0316] First, the XR server device 110 may receive tracking and sensor information 132 from the XR client device 140 (350). For example, the XR client device 140 may determine the orientation in which a user is looking using the tracking / XR sensor 146. The XR client device 140 may send tracking and sensor information 132 indicating the orientation in which the user is looking. The XR server device 110 may receive the tracking and sensor information 132 and provide the information to the XR scene generation unit 112. The XR scene generation unit 112 and the XR viewport prerendering rasterization unit 114 may then generate a video frame of scene data (352). In some examples, a video game engine may use the tracking and sensor information 132, the video game data, and the video game configuration data to generate the scene data.

[0317] The 2D media encoding unit 116 may then encode the generated video frames of scene data (354). In some examples, such as Figure 6 shown in and about Figure 6 As described, such encoding may include intra-prediction encoding of all frames in one run and inter- or intra-prediction encoding of the frames in another run. The intra-prediction encoding may generate vi-track data, while the inter-prediction encoding may generate vp-track data. The XR server device 110 may combine the vi-track data and the vp-track data to form v-track data, which may represent the complexity of the encoding of the video frame. In some examples, the 2D media encoding unit 116 may encode all frames of a particular v-track using a common, unchanging quantization configuration.

[0318] The XR media content delivery unit 118 may then packetize the slices of the encoded video frame (356). Typically, the 2D media encoding unit 116 may be configured with a specific maximum transmission unit (MTU) size for packets for a radio access network (RAN) and generate slices with a data volume less than or equal to the MTU size. Thus, in some examples, each slice can be transmitted within a single packet.

[0319] XR server device 110 may then perform RAN simulation of the RAN, for example, using the configuration of network 130 to transmit packets (358). Figure 2-Figure 4 and Figure 7-Figure 9 An example of performing RAN simulation in this manner is described. For example, when performing RAN simulation, the XR server device 110 can determine that a slice is lost when a packet for a slice is lost or damaged. The XR server device 110 can also calculate the latency for transmitting the packet. The XR server device 110 can also generate S-trace data and P-trace data as discussed above as part of the RAN simulation.

[0320] The XR server device 110 may then assemble the simulated received packets into encoded video frames (360) and decode the video frames (362). The XR server device 110 may then calculate the individual frame quality (364). The individual frame quality may represent the difference between the video frame before or after encoding and the decoded video frame, for example, as described with respect to Figure 2 and Figure 8 discussed.

[0321] The XR server device 110 may then determine an overall quality for the system based on the individual frame qualities (366). For example, the XR server device 110 may determine the overall quality as an average encoding quality and / or an average number of incorrect video frames (e.g., frames for which data is lost or corrupted). The XR server device 110 may perform this method for various different types of configurations, for example, to determine the number of users that can be supported for a given configuration. For example, the XR server device 110 may calculate a percentage of corrupted video frames based on the number of supported users for one or more sets of configurations. Additionally or alternatively, the XR server device 110 may calculate a percentage of corrupted video frames based on a quantized configuration for one or more numbers of supported users.

[0322] In this way, Figure 12 The method represents an example of a method of processing media data, which includes: receiving tracking and sensor information from an extended reality (XR) client device; using the tracking and sensor information to generate scene data, the scene data including one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation that delivers the encoded video frames via a radio access network; decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames based on the one or more generated video frames and the decoded video frames; and determining an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0323] The following clauses represent some examples of the techniques of this disclosure:

[0324] Clause 1: A method of sending media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; generating scene data using the tracking and sensor information; rendering an XR viewport based on the scene data to form video data; encoding the video data; and sending the encoded video data to the XR client device via a 5G network.

[0325] Clause 2: A method of obtaining media data, the method comprising: sending tracking and sensor information to an extended reality (XR) server device; receiving encoded video data corresponding to the tracking and sensor information from the XR server device; decoding the encoded video data; and rendering an XR viewport using the decoded video data.

[0326] Clause 3: A method comprising a combination of the methods according to clauses 1 and 2.

[0327] Clause 4: A method of processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; using the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; encoding the video data to form v-tracking data; performing a radio access network (RAN) simulation of delivering the encoded video data via a 5G network; decoding the delivered encoded video data to form decoded video data; calculating a value representing a quality for each of the video frames based on the generated one or more video frames and the decoded video data; and determining an overall quality value based on the value representing the quality for each of the video frames.

[0328] Clause 5: The method of clause 4, wherein generating the scene data comprises executing a video game engine using the tracking and sensor information, video game data, and video game configuration data.

[0329] Clause 6: The method of any of clauses 4 and 5, wherein encoding the video data comprises encoding the video data without changing a quantization scheme.

[0330] Clause 7: The method of any of clauses 4-6, wherein determining the overall quality value comprises generating a graph representing a percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets.

[0331] Clause 8: The method of any of clauses 4 and 5, wherein determining the overall quality value comprises generating a graph representing a percentage of corrupted video frames as a function of a quantization configuration for one or more supported numbers of users.

[0332] Clause 9: A method according to any one of clauses 4-8, wherein encoding the video data to form the v-tracking data comprises: encoding at least some portions of the video data using inter-frame prediction to form vp-tracking data; encoding at least some portions of the video data using intra-frame prediction to form vi-tracking data; and combining the vp-tracking data with the vi-tracking data to form the v-tracking data.

[0333] Clause 10: The method of any of clauses 4-9, wherein the v-track data is formatted according to FFMPEG's passlogfile format.

[0334] Clause 11: A method according to any of clauses 4-10, wherein the v-tracking data includes one or more of the following: a display picture number value, a decoded picture number value, a picture type value, a quality value, an intra-frame texture bit value, an inter-frame texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.

[0335] Clause 12: A method according to any of clauses 4-11, wherein performing the RAN simulation includes: receiving one or more slices of encoded video data corresponding to the v-trace data; packetizing and segmenting the slices according to IP packet and payload size configuration data using the Real-time Transport Protocol (RTP) to form packets in the p-trace data; simulating the transmission of the packets; determining that one of the slices is lost when a packet for the one of the slices is lost during the simulation of the transmission of the packets; and outputting data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and the latency for transmitting the received packet.

[0336] Clause 13: A method according to any one of clauses 4-12, wherein performing the RAN simulation includes: receiving one or more slices of encoded video data corresponding to the v-trace data; and forming s-trace data from the slices according to configuration data, the configuration data indicating a bit rate control technique, a slice setting, an error recovery technique and a feedback technique.

[0337] Clause 14: The method of clause 13, wherein the bitrate control technique comprises one of: constant quality, constant bitrate, variable bitrate based on feedback, or constant rate factor.

[0338] Clause 15: The method of any one of clauses 13 and 14, wherein the slice setting comprises one or more of a number of slices or a maximum slice size.

[0339] Clause 16: The method of any of clauses 13-15, wherein the error resilience technique comprises one of: frame-based recovery, slice-based recovery, regular intra refresh, feedback-based intra refresh, or feedback-based prediction.

[0340] Clause 17: The method of any of clauses 13-16, wherein the feedback technique comprises one of: no feedback, statistical feedback, or operational feedback.

[0341] Clause 18: A method according to any one of clauses 4-17, wherein calculating the value representing the quality for each of the video frames includes: determining the quality of each received encoded slice; and determining whether the unreceived slice is lost or degraded in quality for each unreceived slice.

[0342] Clause 19: The method of any of clauses 4-18, wherein determining the overall quality value comprises determining the overall quality value using one or more of an average encoding quality or an average incorrect video data.

[0343] Clause 20: An apparatus for processing media data, the apparatus comprising one or more means for performing the method according to any of clauses 1-19.

[0344] Clause 21: The apparatus of clause 20, wherein the means comprises at least one of: an integrated circuit; a microprocessor; and a wireless communication device.

[0345] Clause 22: A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to perform the method of any of clauses 1-19.

[0346] Clause 23: A device for sending media data, the device comprising: means for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information; means for rendering an XR viewport based on the scene data to form video data; means for encoding the video data; and means for sending the encoded video data to the XR client device via a 5G network.

[0347] Clause 24: An apparatus for obtaining media data, the apparatus comprising: means for sending tracking and sensor information to an extended reality (XR) server device; means for receiving encoded video data corresponding to the tracking and sensor information from the XR server device; means for decoding the encoded video data; and means for rendering an XR viewport using the decoded video data.

[0348] Clause 25: A device for processing media data, the device comprising: a unit for receiving tracking and sensor information from an extended reality (XR) client device; a unit for using the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; a unit for encoding the video data to form v-tracking data; a unit for performing a radio access network (RAN) simulation of delivering the encoded video data via a 5G network; a unit for decoding the delivered encoded video data to form decoded video data; a unit for calculating a value representing a quality for each of the video frames based on the generated one or more video frames and the decoded video data; and a unit for determining an overall quality value based on the value representing the quality for each of the video frames.

[0349] Clause 26: A method of processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; using the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determining an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0350] Clause 27: The method of clause 26, wherein generating the scene data comprises executing a video game engine using the tracking and sensor information, video game data, and video game configuration data.

[0351] Clause 28: The method of clause 26, wherein encoding the video frame comprises encoding the video frame without changing a quantization configuration.

[0352] Clause 29: The method of clause 26, wherein determining the overall quality value comprises calculating a percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets.

[0353] Clause 30: The method of clause 26, wherein determining the overall quality value comprises calculating a percentage of corrupted video frames according to a quantization configuration for one or more supported numbers of users.

[0354] Clause 31: A method according to clause 26, wherein encoding the video frame further comprises forming v-tracking data, wherein the v-tracking data represents the complexity of the encoding of the video frame, comprising: encoding at least some portions of the video frame using inter-frame prediction to form vp-tracking data; encoding at least some portions of the video frame using intra-frame prediction to form vi-tracking data; and combining the vp-tracking data with the vi-tracking data to form the v-tracking data.

[0355] Clause 32: A method according to clause 26, wherein encoding the video frame further comprises forming v-track data, wherein the v-track data represents the complexity of the encoding of the video frame, and wherein the v-track data comprises one or more of the following: a display picture number value, a decoded picture number value, a picture type value, a quality value, an intra-frame texture bit value, an inter-frame texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.

[0356] Clause 33: A method according to clause 26, wherein performing the RAN simulation includes: receiving one or more slices of the encoded video frame; packetizing and segmenting the slices according to IP packet and payload size configuration data using the Real-time Transport Protocol (RTP) to form packets in the p-track data; performing a simulation of transmission of the packets; determining that one of the slices is lost when a packet for the one of the slices is lost during the simulation of the transmission of the packets; and outputting data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and the latency for transmitting the received packet.

[0357] Clause 34: A method according to clause 26, wherein performing the RAN simulation includes: receiving one or more slices of the encoded video frame; and forming s-track data from the slices based on configuration data, wherein the configuration data indicates a bit rate control technique, a slice setting, an error recovery technique, and a feedback technique.

[0358] Clause 35: A method according to clause 34, wherein forming the s-tracking data includes: assigning data representing the following to each of the slices: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice.

[0359] Clause 36: The method of clause 34, wherein the bitrate control technique comprises one of: constant quality, constant bitrate, variable bitrate based on feedback, or constant rate factor.

[0360] Clause 37: The method of clause 34, wherein the slice settings include one or more of: number of slices or maximum slice size.

[0361] Clause 38: The method of clause 34, wherein the error resilience technique comprises one of: frame-based recovery, slice-based recovery, regular intra refresh, feedback-based intra refresh, or feedback-based prediction.

[0362] Clause 39: The method of clause 34, wherein the feedback technique comprises one of: statistical feedback or operational feedback.

[0363] Clause 40: A method according to clause 26, wherein calculating the value representing the individual frame quality for each of the video frames includes: determining the quality of each received encoded slice in the video frame; and determining whether the unreceived slice in the video frame is lost or degraded in quality.

[0364] Clause 41: The method of clause 26, wherein determining the overall quality value comprises determining the overall quality value using one or more of average encoding quality or average incorrect video frames.

[0365] Clause 42: The method of clause 26, wherein the wireless access network comprises a 5G network.

[0366] Clause 43: A method according to clause 26, wherein the RAN simulation includes a first RAN simulation among a plurality of RAN simulations, the overall video quality includes a first overall video quality, and performing the first RAN simulation includes performing the first RAN simulation using first configuration parameters, the method further comprising: determining corresponding configuration parameters for each RAN simulation among the plurality of RAN simulations; performing each RAN simulation among the RAN simulations using the corresponding configuration parameters; determining an overall video quality value for each RAN simulation among the RAN simulations; and configuring a network to deliver XR data using the configuration parameters corresponding to the best overall video quality value for the RAN simulation.

[0367] Item 44: A device for processing media data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; use the tracking and sensor information to generate scene data, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decode the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determine an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0368] Clause 45: The device of clause 44, wherein the one or more processors are configured to execute a video game engine using the tracking and sensor information, video game data, and video game configuration data to generate the scene data.

[0369] Clause 46: The apparatus of clause 44, wherein to determine the overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets.

[0370] Clause 47: The apparatus of clause 44, wherein to determine the overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames according to a quantization configuration for one or more supported numbers of users.

[0371] Clause 48: An apparatus according to clause 44, wherein the one or more processors are further configured to form v-tracking data, the v-tracking data representing the complexity of the encoding of the video frame, and wherein, to form the v-tracking data, the one or more processors are configured to: encode at least some portions of the video frame using inter-frame prediction to form vp-tracking data; encode at least some portions of the video frame using intra-frame prediction to form vi-tracking data; and combine the vp-tracking data with the vi-tracking data to form the v-tracking data.

[0372] Clause 49: An apparatus according to clause 44, wherein the one or more processors are further configured to form v-tracking data, the v-tracking data representing the complexity of the encoding of the video frame, and wherein the v-tracking data includes one or more of the following: a display picture number value, a decoded picture number value, a picture type value, a quality value, an intra-frame texture bit value, an inter-frame texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.

[0373] Clause 50: An apparatus according to clause 44, wherein, to perform the RAN simulation, the one or more processors are configured to: receive one or more slices of the encoded video frame; packetize and segment the slices using a real-time transport protocol (RTP) according to IP packet and payload size configuration data to form packets in p-tracking data; perform a simulation of transmission of the packets; determine that one of the slices is lost when a packet for the one of the slices is lost during the simulation of the transmission of the packets; and output data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and a latency for transmitting the received packet.

[0374] Clause 51: An apparatus according to clause 44, wherein, in order to perform the RAN simulation, the one or more processors are configured to: receive one or more slices of the encoded video frame; and form s-track data from the slices according to configuration data, wherein the configuration data indicates a bit rate control technique, a slice setting, an error recovery technique, and a feedback technique.

[0375] Clause 52: An apparatus according to clause 51, wherein, to form the s-tracking data, the one or more processors are configured to: assign to each of the slices data representing: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice.

[0376] Clause 53: An apparatus according to clause 44, wherein, in order to calculate the value representing the individual frame quality for each of the video frames, the one or more processors are configured to: determine the quality of each received encoded slice in the video frame; and determine whether the unreceived slice in the video frame is lost or degraded in quality.

[0377] Item 54: A computer-readable storage medium having instructions stored thereon that, when executed, cause a processor to: receive tracking and sensor information from an extended reality (XR) client device; use the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decode the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determine an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0378] Clause 55: The computer-readable storage medium of clause 54, wherein the instructions to cause the processor to generate the scene data include instructions to cause the processor to execute a video game engine using the tracking and sensor information, video game data, and video game configuration data.

[0379] Clause 56: The computer-readable storage medium of clause 54, wherein the instructions causing the processor to determine the overall quality value comprise instructions causing the processor to calculate a percentage of corrupted video frames as a function of the number of supported users for one or more configuration sets.

[0380] Clause 57: The computer-readable storage medium of clause 54, wherein the instructions causing the processor to determine the overall quality value comprise instructions causing the processor to calculate a percentage of corrupted video frames based on a quantization configuration for one or more supported user numbers.

[0381] Clause 58: A computer-readable storage medium according to clause 54, wherein the instructions that cause the processor to encode the video frame further include instructions that cause the processor to perform the following operations: form v-tracking data, wherein the v-tracking data represents the complexity of the encoding of the video frame, and wherein the instructions that cause the processor to form the v-tracking data include instructions that cause the processor to perform the following operations: encode at least some portions of the video frame using inter-frame prediction to form vp-tracking data; encode at least some portions of the video frame using intra-frame prediction to form vi-tracking data; and combine the vp-tracking data with the vi-tracking data to form the v-tracking data.

[0382] Clause 59: A computer-readable storage medium according to clause 54, wherein the instructions that cause the processor to encode the video frame also include instructions that cause the processor to perform the following operations: form v-tracking data, the v-tracking data representing the complexity of the encoding of the video frame, and wherein the v-tracking data includes one or more of the following: a display picture number value, a decoded picture number value, a picture type value, a quality value, an intra-frame texture bit value, an inter-frame texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.

[0383] Clause 60: A computer-readable storage medium according to clause 54, wherein the instructions that cause the processor to perform the RAN simulation include instructions that cause the processor to perform the following operations: receive one or more slices of the encoded video frame; use the real-time transport protocol (RTP) to packetize and segment the slices according to IP packet and payload size configuration data to form packets in p-tracking data; perform a simulation of transmission of the packets; when a packet for one of the slices is lost during the simulation of the transmission of the packets, determine that the one of the slices is lost; and output data representing the following for each slice: whether the slice was received or lost during the simulation of the transmission, and the latency for transmitting the received packet.

[0384] Clause 61: A computer-readable storage medium according to clause 54, wherein the instructions causing the processor to perform the RAN simulation include instructions causing the processor to perform the following operations: receiving one or more slices of the encoded video frame; and forming s-track data from the slices based on configuration data, wherein the configuration data indicates a bit rate control technique, a slice setting, an error recovery technique, and a feedback technique.

[0385] Clause 62: A computer-readable storage medium according to clause 61, wherein the instructions causing the processor to form the s-tracking data include instructions causing the processor to perform the following operations: assigning data representing the following items to each of the slices: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice.

[0386] Clause 63: Computer-readable storage medium according to clause 54, wherein the instructions causing the processor to calculate the value representing the individual frame quality for each of the video frames include instructions causing the processor to perform the following operations: determining the quality of each received encoded slice in the video frame; and determining whether the unreceived slice in the video frame is lost or degraded in quality.

[0387] Clause 64: A device for processing media data, the device comprising: a unit for receiving tracking and sensor information from an extended reality (XR) client device; a unit for using the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; a unit for encoding the video frames to form encoded video frames; a unit for performing a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; a unit for decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; a unit for calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and a unit for determining an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0388] Clause 65: A method of processing media data, the method comprising: receiving tracking and sensor information from an extended reality (XR) client device; using the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; encoding the video frames to form encoded video frames; performing a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determining an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0389] Clause 66: The method of clause 65, wherein generating the scene data comprises executing a video game engine using the tracking and sensor information, video game data, and video game configuration data.

[0390] Clause 67: The method of any of clauses 65-66, wherein encoding the video frame comprises encoding the video frame without changing a quantization configuration.

[0391] Clause 68: The method of any of clauses 65-67, wherein determining the overall quality value comprises calculating, for one or more configuration sets, a percentage of corrupted video frames as a function of the number of supported users.

[0392] Clause 69: The method of any of clauses 65-68, wherein determining the overall quality value comprises calculating a percentage of corrupted video frames according to a quantization configuration for one or more supported numbers of users.

[0393] Clause 70: A method according to any one of clauses 65-69, wherein encoding the video frame further comprises forming v-tracking data, wherein the v-tracking data represents the complexity of the encoding of the video frame, comprising: encoding at least some portions of the video frame using inter-frame prediction to form vp-tracking data; encoding at least some portions of the video frame using intra-frame prediction to form vi-tracking data; and combining the vp-tracking data with the vi-tracking data to form the v-tracking data.

[0394] Clause 71: A method according to any one of clauses 65-70, wherein encoding the video frame further comprises forming v-tracking data, wherein the v-tracking data represents the complexity of the encoding of the video frame, and wherein the v-tracking data comprises one or more of the following: a display picture number value, a decoded picture number value, a picture type value, a quality value, an intra-frame texture bit value, an inter-frame texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.

[0395] Clause 72: A method according to any one of clauses 65-71, wherein performing the RAN simulation comprises: receiving one or more slices of the encoded video frame; packetizing and segmenting the slices according to IP packet and payload size configuration data using the Real-time Transport Protocol (RTP) to form packets in the p-tracking data; performing a simulation of the transmission of the packets; determining that one of the slices is lost when a packet for the one of the slices is lost during the simulation of the transmission of the packets; and outputting data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and the latency for transmitting the received packet.

[0396] Clause 73: A method according to any one of clauses 65-72, wherein performing the RAN simulation comprises: receiving one or more slices of the encoded video frame; and forming s-track data from the slices based on configuration data, wherein the configuration data indicates a bit rate control technique, a slice setting, an error recovery technique, and a feedback technique.

[0397] Clause 74: A method according to clause 73, wherein forming the s-tracking data includes: assigning data representing the following to each of the slices: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice.

[0398] Clause 75: The method of any of clauses 73 and 74, wherein the bitrate control technique comprises one of: constant quality, constant bitrate, variable bitrate based on feedback, or constant rate factor.

[0399] Clause 76: The method of any of clauses 73-75, wherein the slice settings comprise one or more of: number of slices or maximum slice size.

[0400] Clause 77: The method of any of clauses 73-76, wherein the error resilience technique comprises one of: frame-based recovery, slice-based recovery, regular intra-frame refresh, feedback-based intra-frame refresh, or feedback-based prediction.

[0401] Clause 78: The method of any of clauses 73-77, wherein the feedback technique comprises one of: statistical feedback or operational feedback.

[0402] Clause 79: A method according to any one of clauses 65-78, wherein calculating the value representing the individual frame quality for each of the video frames comprises: determining the quality of each received encoded slice in the video frame; and determining whether the unreceived slice in the video frame is lost or degraded in quality.

[0403] Clause 80: The method of any of clauses 65-79, wherein determining the overall quality value comprises determining the overall quality value using one or more of average encoding quality or average incorrect video frames.

[0404] Clause 81: A method as described in any of clauses 65-80, wherein the wireless access network comprises a 5G network.

[0405] Clause 82: A method according to any one of clauses 65-81, wherein the RAN simulation comprises a first RAN simulation among a plurality of RAN simulations, the overall video quality comprises a first overall video quality, and performing the first RAN simulation comprises performing the first RAN simulation using first configuration parameters, the method further comprising: determining corresponding configuration parameters for each RAN simulation among the plurality of RAN simulations; performing each RAN simulation among the RAN simulations using the corresponding configuration parameters; determining an overall video quality value for each RAN simulation among the RAN simulations; and configuring a network to deliver XR data using the configuration parameters corresponding to the best overall video quality value for the RAN simulation.

[0406] Item 83: A device for processing media data, the device comprising: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive tracking and sensor information from an extended reality (XR) client device; use the tracking and sensor information to generate scene data, the scene data comprising one or more video frames of the video data; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decode the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determine an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0407] Clause 84: The apparatus of clause 83, wherein the one or more processors are configured to execute a video game engine using the tracking and sensor information, video game data, and video game configuration data to generate the scene data.

[0408] Clause 85: The apparatus of any of clauses 83 and 84, wherein to determine the overall quality value, the one or more processors are configured to calculate, for one or more configuration sets, a percentage of corrupted video frames as a function of the number of supported users.

[0409] Clause 86: The apparatus of any of clauses 83-85, wherein, to determine the overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames according to a quantization configuration for one or more supported numbers of users.

[0410] Clause 87: An apparatus according to any one of clauses 83-86, wherein the one or more processors are further configured to form v-tracking data, the v-tracking data representing the complexity of the encoding of the video frame, and wherein, to form the v-tracking data, the one or more processors are configured to: encode at least some portions of the video frame using inter-frame prediction to form vp-tracking data; encode at least some portions of the video frame using intra-frame prediction to form vi-tracking data; and combine the vp-tracking data with the vi-tracking data to form the v-tracking data.

[0411] Clause 88: An apparatus according to any one of clauses 83-87, wherein the one or more processors are further configured to form v-tracking data, wherein the v-tracking data represents the complexity of the encoding of the video frame, and wherein the v-tracking data includes one or more of the following: a display picture number value, a decoded picture number value, a picture type value, a quality value, an intra-frame texture bit value, an inter-frame texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.

[0412] Clause 89: An apparatus according to any of clauses 83-88, wherein, in order to perform the RAN simulation, the one or more processors are configured to: receive one or more slices of the encoded video frame; packetize and segment the slices using the Real-time Transport Protocol (RTP) according to IP packet and payload size configuration data to form packets in the p-tracking data; perform a simulation of the transmission of the packets; determine that one of the slices is lost when a packet for the one of the slices is lost during the simulation of the transmission of the packets; and output data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and the latency for transmitting the received packet.

[0413] Clause 90: An apparatus according to any one of clauses 83-89, wherein, in order to perform the RAN simulation, the one or more processors are configured to: receive one or more slices of the encoded video frame; and form s-track data from the slices based on configuration data, the configuration data indicating a bit rate control technique, a slice setting, an error recovery technique and a feedback technique.

[0414] Clause 91: An apparatus according to clause 90, wherein, to form the s-tracking data, the one or more processors are configured to: assign to each of the slices data representing: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice.

[0415] Clause 92: An apparatus according to any one of clauses 83-91, wherein, in order to calculate the value representing the individual frame quality for each of the video frames, the one or more processors are configured to: determine the quality of each received encoded slice in the video frame; and determine whether the unreceived slice in the video frame is lost or degraded in quality.

[0416] Item 93: A computer-readable storage medium having instructions stored thereon that, when executed, cause a processor to: receive tracking and sensor information from an extended reality (XR) client device; use the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; encode the video frames to form encoded video frames; perform a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; decode the encoded video frames delivered according to the RAN simulation to form decoded video frames; calculate a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and determine an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0417] Clause 94: The computer-readable storage medium of clause 93, wherein the instructions to cause the processor to generate the scene data include instructions to cause the processor to execute a video game engine using the tracking and sensor information, video game data, and video game configuration data.

[0418] Clause 95: A computer-readable storage medium according to any one of clauses 93 and 94, wherein the instructions causing the processor to determine the overall quality value include instructions causing the processor to perform the following operations: calculate the percentage of damaged video frames based on the number of supported users for one or more configuration sets.

[0419] Clause 96: A computer-readable storage medium according to any of clauses 93-95, wherein the instructions causing the processor to determine the overall quality value include instructions causing the processor to perform the following operations: calculate the percentage of damaged video frames based on a quantization configuration for one or more supported user numbers.

[0420] Clause 97: A computer-readable storage medium according to any one of clauses 93-96, wherein the instructions that cause the processor to encode the video frame further include instructions that cause the processor to form v-tracking data, wherein the v-tracking data represents the complexity of the encoding of the video frame, and wherein the instructions that cause the processor to form the v-tracking data include instructions that cause the processor to encode at least some portions of the video frame using inter-frame prediction to form vp-tracking data; encode at least some portions of the video frame using intra-frame prediction to form vi-tracking data; and combine the vp-tracking data with the vi-tracking data to form the v-tracking data.

[0421] Clause 98: A computer-readable storage medium according to any one of clauses 93-97, wherein the instructions that cause the processor to encode the video frame also include instructions that cause the processor to perform the following operations: form v-tracking data, the v-tracking data representing the complexity of the encoding of the video frame, and wherein the v-tracking data includes one or more of the following: a display picture number value, a decoded picture number value, a picture type value, a quality value, an intra-frame texture bit value, an inter-frame texture bit value, a motion vector bit value, a miscellaneous bit value, an f-code value, a b-code value, an mc_mb_var_sum value, an mb_var_sum value, an i_count value, a skip count value, or a header bit value.

[0422] Clause 99: A computer-readable storage medium according to any one of clauses 93-98, wherein the instructions causing the processor to perform the RAN simulation include instructions causing the processor to perform the following operations: receiving one or more slices of the encoded video frame; packetizing and segmenting the slices according to IP packet and payload size configuration data using the Real-time Transport Protocol (RTP) to form packets in the p-tracking data; performing a simulation of the transmission of the packets; when a packet for one of the slices is lost during the simulation of the transmission of the packets, determining that the one of the slices is lost; and outputting data representing the following for each slice: whether the slice was received or lost during the simulation of the transmission, and the latency for transmitting the received packet.

[0423] Clause 100: A computer-readable storage medium according to any one of clauses 93-99, wherein the instructions causing the processor to perform the RAN simulation include instructions causing the processor to perform the following operations: receiving one or more slices of the encoded video frame; and forming s-tracking data from the slices based on configuration data, wherein the configuration data indicates a bit rate control technique, a slice setting, an error recovery technique, and a feedback technique.

[0424] Clause 101: A computer-readable storage medium according to clause 100, wherein the instructions causing the processor to form the s-tracking data include instructions causing the processor to perform the following operations: assigning data representing the following items to each of the slices: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice.

[0425] Clause 102: Computer-readable storage medium according to any one of clauses 93-101, wherein the instructions causing the processor to calculate the value representing the individual frame quality for each of the video frames include instructions causing the processor to perform the following operations: determining the quality of each received encoded slice in the video frame; and determining whether the unreceived slice in the video frame is lost or degraded in quality.

[0426] Clause 103: A device for processing media data, the device comprising: a unit for receiving tracking and sensor information from an extended reality (XR) client device; a unit for using the tracking and sensor information to generate scene data, the scene data comprising one or more video frames; a unit for encoding the video frames to form encoded video frames; a unit for performing a radio access network (RAN) simulation of delivering the encoded video frames via a radio access network; a unit for decoding the encoded video frames delivered according to the RAN simulation to form decoded video frames; a unit for calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frames; and a unit for determining an overall quality value based on the value representing the individual frame quality for each of the video frames.

[0427] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted through a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to tangible media such as data storage media, or communication media, which includes, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include computer-readable media.

[0428] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies (such as infrared, radio, and microwaves), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwaves) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transitory tangible storage media. As used herein, disks and optical disks include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs utilize lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0429] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the term "processor" as used herein may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Furthermore, the techniques may be implemented entirely in one or more circuits or logic elements.

[0430] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but do not necessarily require implementation by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) in conjunction with appropriate software and / or firmware.

[0431] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for processing media data, the method comprising: Receive tracking and sensor information from the extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frame to form an encoded video frame and video tracking data (v-tracking data), the v-tracking data including information related to complexity and data rate for an encoding model; generating slice tracking data, i.e., s-tracking data, and packets using the coded video frame and the v-tracking data, the packets representing encapsulated coded slices of the coded video frame, wherein forming the s-tracking data comprises: assigning to each of the slices data representing: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice; performing a radio access network RAN ​​simulation of delivering the encoded video frame via the radio access network RAN, wherein performing the RAN simulation comprises: generating s'-tracking data based on the s-tracking data and the packets using configuration data, the configuration data indicating a bitrate control technique, a slice setting, an error recovery technique, and a feedback technique, the s'-tracking data including: associated frames, slice size, quality, number of macroblocks, and timing information, as well as slice loss indicators and slice delay information; decoding the encoded video frames delivered according to the RAN simulation using the s′-trace data to form decoded video frames; Calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frame, wherein calculating the value representing the individual frame quality for each of the video frames comprises: determining, for each received encoded slice in the video frame, a quality of the slice; and determining, for any unreceived slice in the video frame, whether the unreceived slice is lost or degraded; and An overall quality value is determined based on values ​​representing the individual frame quality for each of the video frames.

2. The method according to claim 1, wherein Generating the scenario data includes executing a video game engine using the tracking and sensor information, the video game data, and the video game configuration data.

3. The method according to claim 1, wherein Encoding the video frame includes encoding the video frame without changing a quantization configuration.

4. The method according to claim 1, wherein Determining the overall quality value includes calculating a percentage of corrupted video frames based on the number of supported users for one or more configuration sets.

5. The method according to claim 1, wherein Determining the overall quality value includes calculating a percentage of corrupted video frames based on a quantization configuration for one or more supported user numbers.

6. The method according to claim 1, wherein Forming the v-tracking data includes: encoding at least some portions of the video frames using inter-frame prediction to form vp-tracking data; encoding said at least some portions of said video frames using intra prediction to form vi-track data; and The vp-tracking data is combined with the vi-tracking data to form the v-tracking data.

7. The method according to claim 1, wherein Executing the RAN simulation includes: receiving one or more slices of the encoded video frame; packetizing and segmenting the slices according to the IP packet and payload size configuration data to form IP packets and p-track data using a real-time transport protocol (RTP); performing a simulation of transmission of the IP packet; determining that one of the slices is lost when an IP packet for the one of the slices is lost during the simulation of the transmission of the IP packet; and Data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and a delay for transmitting the received IP packet are output.

8. The method according to claim 1, wherein The bitrate control techniques include one of: constant quality, constant bitrate, variable bitrate based on feedback, or constant rate factor.

9. The method according to claim 1, wherein The slice settings include one or more of the following: number of slices or maximum slice size.

10. The method according to claim 1, wherein The error resilience technique includes one of: frame-based recovery, slice-based recovery, regular intra-frame refresh, feedback-based intra-frame refresh, or feedback-based prediction.

11. The method according to claim 1, wherein Determining the overall quality value includes using one or more of an average encoding quality or an average number of incorrect video frames to determine the overall quality value.

12. The method according to claim 1, wherein The wireless access network includes a 5G network.

13. The method according to claim 1, wherein The RAN simulation comprises a first RAN simulation of a plurality of RAN simulations, the overall video quality comprises a first overall video quality, and performing the first RAN simulation comprises performing the first RAN simulation using first configuration parameters, the method further comprising: determining corresponding configuration parameters for each RAN simulation in the plurality of RAN simulations; executing each of the RAN simulations using the corresponding configuration parameters; determining an overall video quality value for each of the RAN simulations; and A network is configured to deliver XR data using the configuration parameters corresponding to the best overall video quality value simulated for the RAN.

14. A device for processing media data, the device comprising: a memory configured to store video data; as well as One or more processors implemented in circuitry and configured to: Receive tracking and sensor information from the extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames of the video data; encoding the video frame to form an encoded video frame and video tracking data (v-tracking data), the v-tracking data including information related to complexity and data rate for an encoding model; generating slice tracking data, i.e., s-tracking data, and packets using the coded video frame and the v-tracking data, the packets representing encapsulated coded slices of the coded video frame, wherein forming the s-tracking data comprises: assigning to each of the slices data representing: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice; performing a radio access network RAN ​​simulation of delivering the encoded video frame via the radio access network RAN, wherein performing the RAN simulation comprises: generating s'-tracking data based on the s-tracking data and the packets using configuration data, the configuration data indicating a bitrate control technique, a slice setting, an error recovery technique, and a feedback technique, the s'-tracking data including: associated frames, slice size, quality, number of macroblocks, and timing information, as well as slice loss indicators and slice delay information; decoding the encoded video frames delivered according to the RAN simulation using the s′-trace data to form decoded video frames; Calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frame, wherein to calculate the value representing the individual frame quality for each of the video frames, the one or more processors are configured to: determining, for each received encoded slice in the video frame, a quality of the slice; and determining, for any unreceived slice in the video frame, whether the unreceived slice is lost or degraded; and An overall quality value is determined based on values ​​representing the individual frame quality for each of the video frames.

15. The apparatus according to claim 14, wherein The one or more processors are configured to execute a video game engine using the tracking and sensor information, video game data, and video game configuration data to generate the scene data.

16. The apparatus according to claim 14, wherein To determine the overall quality value, the one or more processors are configured to calculate, for one or more configuration sets, a percentage of corrupted video frames as a function of the number of supported users.

17. The apparatus according to claim 14, wherein To determine the overall quality value, the one or more processors are configured to calculate a percentage of corrupted video frames for one or more supported user numbers according to a quantization configuration.

18. The apparatus according to claim 14, wherein To form the v-tracking data, the one or more processors are configured to: encoding at least some portions of the video frames using inter-frame prediction to form vp-tracking data; encoding said at least some portions of said video frames using intra prediction to form vi-track data; and The vp-tracking data is combined with the vi-tracking data to form the v-tracking data.

19. The apparatus according to claim 14, wherein To perform the RAN simulation, the one or more processors are configured to: receiving one or more slices of the encoded video frame; packetizing and segmenting the slices according to the IP packet and payload size configuration data to form IP packets and p-track data using a real-time transport protocol (RTP); performing a simulation of transmission of the IP packet; determining that one of the slices is lost when an IP packet for the one of the slices is lost during the simulation of the transmission of the IP packet; as well as Data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and a delay for transmitting the received IP packet are output.

20. A computer-readable storage medium having stored thereon instructions that, when executed, cause a processor to: Receive tracking and sensor information from the extended reality (XR) client device; generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; encoding the video frame to form an encoded video frame and video tracking data (v-tracking data), the v-tracking data including information related to complexity and data rate for an encoding model; Slice tracking data, s-tracking data, and packets representing encapsulated coded slices of the coded video frame are generated using the coded video frame and the v-tracking data, wherein Forming the s-tracking data includes assigning to each of the slices data representing: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice; performing a radio access network RAN ​​simulation of delivering the encoded video frame via the radio access network RAN, wherein performing the RAN simulation comprises: generating s'-tracking data based on the s-tracking data and the packets using configuration data, the configuration data indicating a bitrate control technique, a slice setting, an error recovery technique, and a feedback technique, the s'-tracking data including: associated frames, slice size, quality, number of macroblocks, and timing information, as well as slice loss indicators and slice delay information; decoding the encoded video frames delivered according to the RAN simulation using the s′-trace data to form decoded video frames; calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frame, wherein the instructions causing the processor to calculate the value representing the individual frame quality for each of the video frames include instructions causing the processor to: determining, for each received encoded slice in the video frame, a quality of the slice; and determining, for any unreceived slice in the video frame, whether the unreceived slice is lost or degraded; and An overall quality value is determined based on values ​​representing the individual frame quality for each of the video frames.

21. The computer-readable storage medium of claim 20, wherein: The instructions that cause the processor to generate the scene data include instructions that cause the processor to execute a video game engine using the tracking and sensor information, video game data, and video game configuration data.

22. The computer-readable storage medium of claim 20, wherein: The instructions causing the processor to determine the overall quality value include instructions causing the processor to calculate, for one or more configuration sets, a percentage of corrupted video frames as a function of a number of supported users.

23. The computer-readable storage medium of claim 20, wherein: The instructions causing the processor to determine the overall quality value include instructions causing the processor to calculate a percentage of corrupted video frames based on a quantization configuration for one or more supported user numbers.

24. The computer-readable storage medium of claim 20, wherein: The instructions causing the processor to form the v-tracking data include instructions causing the processor to: encoding at least some portions of the video frames using inter-frame prediction to form vp-tracking data; encoding said at least some portions of said video frames using intra prediction to form vi-track data; and The vp-tracking data is combined with the vi-tracking data to form the v-tracking data.

25. The computer-readable storage medium of claim 20, wherein: The instructions causing the processor to perform the RAN simulation include instructions causing the processor to: receiving one or more slices of the encoded video frame; packetizing and segmenting the slices according to the IP packet and payload size configuration data to form IP packets and p-track data using a real-time transport protocol (RTP); performing a simulation of transmission of the IP packet; determining that one of the slices is lost when an IP packet for the one of the slices is lost during the simulation of the transmission of the IP packet; as well as Data representing, for each slice, whether the slice was received or lost during the simulation of the transmission and a delay for transmitting the received IP packet are output.

26. A device for processing media data, the device comprising: A unit for receiving tracking and sensor information from an extended reality (XR) client device; means for generating scene data using the tracking and sensor information, the scene data comprising one or more video frames; means for encoding the video frame to form an encoded video frame and video tracking data, i.e., v-tracking data, the v-tracking data comprising information related to complexity and data rate for a coding model; means for generating slice tracking data, i.e., s-tracking data, and packets representing encapsulated coded slices of the coded video frame using the coded video frame and the v-tracking data, wherein forming the s-tracking data comprises: assigning to each of the slices data representing: a frame associated with the slice, a size of the slice, a quality of the slice, an area of ​​the frame covered by the slice, and timing information for the slice; means for performing a radio access network RAN ​​simulation of delivering the coded video frame via the RAN, wherein the means for performing the RAN simulation comprises: means for generating s'-tracking data from the s-tracking data and the packets using configuration data indicating a bitrate control technique, a slice setting, an error recovery technique, and a feedback technique, the s'-tracking data including: associated frames, slice size, quality, number of macroblocks, and timing information, as well as slice loss indicator and slice delay information; means for decoding the encoded video frames delivered according to the RAN simulation using the s'-trace data to form decoded video frames; means for calculating a value representing an individual frame quality for each of the video frames based on the generated one or more video frames and the decoded video frame, wherein the means for calculating a value representing the individual frame quality for each of the video frames comprises: means for determining, for each received coded slice in the video frame, a quality of the slice; and means for determining, for any unreceived slice in the video frame, whether the unreceived slice is lost or degraded; and Means for determining an overall quality value based on values ​​representing the individual frame quality for each of the video frames.