Intra refresh and error tracking based on feedback information
By adaptively inserting intra-frame predictive decoding frames and slices based on feedback information in video decoding technology, the problem of decreased user experience caused by lost or damaged video information is solved, achieving low motion-to-photon latency and efficient video transmission, which is suitable for extended reality systems.
Patent Information
- Application Number
- CN202080054035.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-01
- Filing Date
- 2020-07-23
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2040-07-23
AI Technical Summary
Existing video decoding technologies struggle to effectively adapt to lost or corrupted video information, leading to issues such as decreased user experience and motion sickness. This is especially true in extended reality systems, where low motion-to-photon latency is difficult to guarantee.
By adaptively inserting intra-frame predictive decoding frames or slices based on feedback information, the insertion of I-frames and I-slices in the video bitstream is dynamically adjusted. By utilizing error concealment techniques and common reference clock synchronization, video data transmission and processing are optimized.
It improves the reliability of video data transmission and user experience, reduces the risk of motion sickness, optimizes video quality and transmission efficiency, and adapts to the synchronization needs of multi-user environments.
Smart Images

Figure CN114208170B_ABST
Abstract
Description
[0001] Claiming priority
[0002] This patent application claims priority to U.S. nonprovisional application No. 16 / 529,710, filed August 1, 2019, entitled “DYNAMIC VIDEO INSERTIONBASED ON FEEDBACK INFORMATION,” which has been assigned to the assignee of this application and is expressly incorporated herein by reference. Technical Field
[0003] This application relates to media-related technologies. For example, aspects of this application relate to systems, methods, and computer-readable media for performing dynamic video insertion based on feedback regarding lost or damaged video information. Background Technology
[0004] Many devices and systems allow video data to be processed and output for consumption. Digital video data comprises vast amounts of data to meet the needs of both consumers and video providers. For example, consumers of video data expect the highest quality video, characterized by high fidelity, high resolution, and high frame rates. As a result, the large amounts of video data required to meet these needs place a burden on the communication networks and equipment used to process and store this video data.
[0005] Various video decoding technologies can be used to compress video data. Video decoding can be performed according to one or more video decoding standards. Examples of video decoding standards include Universal Video Codec (VCC), High Efficiency Video Codec (HEVC), Advanced Video Codec (AVC), Moving Picture Experts Group (MPEG) Related Decoding (e.g., MPEG-2 Part 2 Decoding), VP9, and Alliance for Open Media (AOMedia) Video 1 (AV1). Video decoding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of redundancy present in video images or sequences. A key goal of video decoding technology is to compress video data to a lower bitrate while avoiding or minimizing video quality degradation.
[0006] Video can be used in many different media environments. One example of such a media environment is extended reality (XR) systems, which include augmented reality (AR), virtual reality (VR), mixed reality (MR), and so on. Each of these forms of XR allows users to experience or interact with virtual content, sometimes in combination with real content. XR systems need to provide low motion-to-photon latency, which is the delay from when user movement occurs to when the corresponding content is displayed. Low motion-to-photon latency is important to provide a good user experience and prevent motion sickness or other adverse effects on client devices (e.g., head-mounted displays or HMDs).
[0007] As more and more video services become available (including XR technology), there is a need for decoding technologies with better decoding efficiency, as well as other video processing and management technologies. Summary of the Invention
[0008] Systems and techniques are described for adaptively controlling encoding devices (e.g., video encoders in a boundless extended reality (XR) architecture and / or other suitable systems) based on feedback information indicating that video data has missing (or corrupted) or corrupted packets. The feedback information may indicate that a video frame, video slice, a portion thereof, or other video information is a missing or corrupted packet. In one example, the feedback information may indicate that at least a portion of a video slice of a video frame is missing or corrupted. The missing or corrupted portion of the video slice may include certain packets of the video slice.
[0009] Feedback information can be provided from the client device to the encoding device. The encoding device can use the feedback information to determine when to adaptively insert intra-predictive decoded frames (also known as intra-decoded frames or pictures) or intra-predictive decoded slices (also known as intra-decoded slices) into the encoded video bitstream. The client device can rely on concealment (e.g., asynchronous time warp (ATW) concealment) until it receives error-free intra-decoded frames or slices.
[0010] In some implementations, intra-decoded frames (I-frames) can be dynamically inserted into the encoded video bitstream based on feedback information. For example, in systems with a strictly constant bit rate (CBR), I-frames can be dynamically inserted into the encoded video bitstream. In such implementations, using feedback from the client device indicating that packets are lost or corrupted, the encoding device can relax (or even eliminate) the periodic insertion of I-frames into the coding structure.
[0011] In some implementations, intra-frame decoded slices (I-slices) with intra-frame refresh periods can be dynamically inserted into the encoded video bitstream. The intra-frame refresh period spreads the intra-frame decoded blocks of the I-frames over several frames. The slice size within the intra-frame refresh period can be increased to ensure that the intra-frame blocks of the intra-frame refresh slices cover completely missing slices with any possible propagating motion. Using feedback information, the periodic insertion of intra-frame refresh periods can be relaxed (or even eliminated in some cases).
[0012] In some implementations, individual I-slices can be dynamically inserted into the encoded video bitstream based on feedback information. For example, if the encoding device allows, it can insert I-slices into error-affected portions of the bitstream. The size of the I-slice can be increased to ensure that intra-frame blocks of the I-slice cover completely missing slices with any possible propagating motion. In some examples, the encoding device can decide to insert the required I-slices across multiple frames to avoid introducing instantaneous degradation in quality.
[0013] In some examples, systems and techniques are also provided for synchronizing encoding devices with a common reference clock (e.g., set by a wireless access point or other device or entity) and with other encoding devices, which can be helpful in multi-user environments. Regardless of the encoding configuration, synchronization with a common clock helps reduce bit rate fluctuations on the wireless link. In some cases, encoding devices can utilize the dynamic insertion of I-frames and / or I-slices to synchronize with the common clock. For example, when an encoding device receives feedback indicating one or more missing packets and needs to force an I-frame or I-slice, any non-urgent (e.g., non-feedback-based) I-frame and / or I-slice insertions from other encoding devices can be delayed so that an encoding device can insert an I-frame or I-slice as soon as possible.
[0014] In an illustrative example of adaptive insertion of intra-frame predictive decoded frames based on feedback information, a method for processing video data is provided. The method includes: determining, by a computing device, that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted. The method further includes: sending feedback information to an encoding device. The feedback information indicates that at least the portion of the video slice is missing or corrupted. The method further includes: receiving an updated video bitstream from the encoding device in response to the feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the missing or corrupted video slice. The size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the missing or corrupted slice and the propagation error in the video frame caused by the missing or corrupted slice.
[0015] In another example of adaptive insertion of intra-frame predictive decoding frames based on feedback information, an apparatus for processing video data is provided, the apparatus including a memory and a processor implemented in circuitry and coupled to the memory. In some examples, more than one processor may be coupled to the memory. The processor is configured to: determine that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted. The processor is also configured to: send feedback information to an encoding device. The feedback information indicates that at least the portion of the video slice is missing or corrupted. The processor is further configured to: receive an updated video bitstream from the encoding device in response to the feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the missing or corrupted video slice. The size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the missing or corrupted slice and the propagation error in the video frame caused by the missing or corrupted slice.
[0016] In another example of adaptive insertion of intra-frame predictive decoded frames based on feedback information, a non-transitory computer-readable medium of a computing device is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: determine that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted; send feedback information to an encoding device indicating that at least the portion of the video slice is missing or corrupted; and, in response to the feedback information, receive an updated video bitstream from the encoding device, the updated video bitstream comprising at least one intra-frame decoded video slice having a size larger than the missing or corrupted video slice, wherein the size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the missing or corrupted slice and the propagation error in the video frame caused by the missing or corrupted slice.
[0017] In another example of adaptive insertion of intra-frame predictive decoding frames based on feedback information, an apparatus for processing video data is provided. The apparatus includes: a unit for determining that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted. The apparatus further includes: a unit for sending feedback information to an encoding device. The feedback information indicates that at least the portion of the video slice is missing or corrupted. The apparatus further includes: a unit for receiving an updated video bitstream from the encoding device in response to the feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the missing or corrupted video slice. The size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the missing or corrupted slice and the propagation error in the video frame caused by the missing or corrupted slice.
[0018] In some aspects, the propagation error in the video frame caused by the lost or damaged slices is based on the motion search range.
[0019] In some aspects, the lost or damaged slice spans from the first row to the second row in the video frame, and the size of the at least one intra-frame decoded video slice is defined as including the first row minus the motion search range to the second row plus the motion search range.
[0020] In some aspects, the methods, apparatus, and computer-readable media described above may include: performing error hiding on one or more video frames in response to determining that at least a portion of the video slice is missing or corrupted, until an error-free intra-decoded video slice is received in the updated video bitstream.
[0021] In some aspects, the at least one intra-decoded video slice includes an intra-decoded frame. In some aspects, the at least one intra-decoded video slice is included as part of an intra-refresh period. The intra-refresh period includes at least one video frame, each of the at least one video frame including one or more intra-decoded video slices. In some examples, the number of the at least one video frame in the intra-refresh period is based on at least one of the following: the number of slices in the video frame including the video slice, the position of the video slice in the video frame, and when the intra-refresh period is inserted into the updated video bitstream based on the feedback information.
[0022] In one example, when the position of the video slice is the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle comprises at least two frames. In such an example, the above-described method, apparatus, and computer-readable medium may include performing error hiding on the first frame of the at least two frames but not on the second frame of the at least two frames, the second frame being after the first frame in the video bitstream. In another example, when the position of the video slice is not the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle comprises an intra-frame decoded frame. In another example, when the position of the video slice is the last slice in the video frame, the at least one video frame in the intra-frame refresh cycle comprises at least two frames. In such an example, the above-described method, apparatus, and computer-readable medium may include performing error hiding on the first and second frames of the at least two frames based on the video slice being the last slice in the video frame.
[0023] In some implementations, the computing device and / or the apparatus includes an extended reality display device configured to provide motion information to the encoding device to generate the video bitstream for display by the extended reality display device.
[0024] In another illustrative example of adaptive insertion of intra-frame predictive decoded frames based on feedback information, a method for processing video data is provided. The method includes: receiving feedback information from a computing device at an encoding device. The feedback information indicates that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted. The method further includes: generating an updated video bitstream in response to the feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the missing or corrupted video slice. The size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the missing or corrupted slice and the propagation error in the video frame caused by the missing or corrupted slice.
[0025] In another example of adaptive insertion of intra-frame predictive decoded frames based on feedback information, an apparatus for processing video data is provided, the apparatus including a memory and a processor implemented in circuitry and coupled to the memory. In some examples, more than one processor may be coupled to the memory. The processor is configured to receive feedback information from a computing device. The feedback information indicates that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted. The processor is also configured to generate an updated video bitstream in response to the feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the missing or corrupted video slice. The size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the missing or corrupted slice and the propagation error in the video frame caused by the missing or corrupted slice.
[0026] In another example of adaptive insertion of intra-frame predictive decoded frames based on feedback information, a non-transitory computer-readable medium of an encoding device is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: receive feedback information from a computing device, the feedback information indicating that at least a portion of a video slice of a video frame in a video bitstream is lost or corrupted; and, in response to the feedback information, generate an updated video bitstream comprising at least one intra-frame decoded video slice having a size larger than the lost or corrupted video slice, wherein the size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the lost or corrupted slice and the propagation error in the video frame caused by the lost or corrupted slice.
[0027] In another example of adaptive insertion of intra-frame predictive decoding frames based on feedback information, an apparatus for processing video data is provided. The apparatus includes: a unit for receiving feedback information from a computing device. The feedback information indicates that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted. The apparatus further includes: a unit for generating an updated video bitstream in response to the feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the missing or corrupted video slice. The size of the at least one intra-frame decoded video slice is determined to cover (or compensate for) the missing or corrupted slice and the propagation error in the video frame caused by the missing or corrupted slice.
[0028] In some aspects, the propagation error in the video frame caused by the lost or damaged slices is based on the motion search range.
[0029] In some aspects, the missing or corrupted slice spans from the first row to the second row in the video frame. In such an aspect, the methods, apparatus, and computer-readable media described above may include determining the size of the at least one intra-frame decoded video slice to include the first row minus the motion search range to the second row plus the motion search range.
[0030] In some aspects, in response to at least a portion of the video slice being lost or corrupted, error concealment is performed on one or more video frames until an error-free intra-decoded video slice is received in the updated video bitstream.
[0031] In some aspects, the at least one intra-decoded video slice includes an intra-decoded frame. In some aspects, the at least one intra-decoded video slice is included as part of an intra-refresh period. The intra-refresh period includes at least one video frame, each of the at least one video frame including one or more intra-decoded video slices. In some examples, the above-described methods, apparatus, and computer-readable media may include: determining the number of the at least one video frame in the intra-refresh period based on at least one of the following: the number of slices in the video frame including the video slice, the position of the video slice in the video frame, and when to insert the intra-refresh period into the updated video bitstream based on the feedback information.
[0032] In one example, when the position of the video slice is the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle is determined to include at least two frames. In such an example, error hiding can be performed on the first frame of the at least two frames without performing error hiding on the second frame of the at least two frames, which is after the first frame in the video bitstream. In another example, when the position of the video slice is not the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle includes an intra-frame decoded frame. In another example, when the position of the video slice is the last slice in the video frame, the at least one video frame in the intra-frame refresh cycle is determined to include at least two frames. In such an example, error hiding is performed on the first and second frames of the at least two frames based on the fact that the video slice is the last slice in the video frame.
[0033] In some aspects, the methods, apparatus, and computer-readable media described above may include: storing the updated video bitstream. In some aspects, the methods, apparatus, and computer-readable media described above may include: sending the updated video bitstream to the computing device.
[0034] In some aspects, the methods, apparatus, and computer-readable media described above may include: adding intra-frame decoded video data to the video bitstream according to a reference clock shared with at least one other encoding device, the reference clock defining a staggered scheduling of intra-frame decoded video from the encoding device and the at least one other encoding device; sending a request to adapt the reference clock in response to the feedback information to allow the encoding device to add intra-frame decoded video data to the video bitstream in unscheduled time slots; receiving an indication that the reference clock has been updated to define an updated schedule; and adding the intra-frame decoded video slice to the video bitstream according to the updated reference clock based on the updated schedule.
[0035] In some implementations, the computing device includes an extended reality display device, and the encoding device is part of a server. The encoding device is configured to generate the video bitstream based on motion information received by the encoding device from the extended reality display device for display by the extended reality display device.
[0036] In an illustrative example of synchronizing an encoding device with a common reference clock, a method for processing video data is provided. The method includes: generating a video bitstream by the encoding device; inserting intra-frame decoded video data into the video bitstream according to a reference clock shared with at least one other encoding device (e.g., by the encoding device). The reference clock defines a scheduling for interleaving intra-frame decoded video from the encoding device and the at least one other encoding device. The method further includes: obtaining feedback information from the encoding device indicating that at least a portion of a video slice of the video bitstream is missing or corrupted. The method further includes: sending a request to adapt the reference clock in response to the feedback information to allow the encoding device to insert intra-frame decoded video data into the video bitstream at unscheduled time slots. The method further includes: receiving an indication that the reference clock has been updated to define an updated schedule; and inserting the intra-frame decoded video data into the video bitstream according to the updated reference clock based on the updated schedule.
[0037] In another example of an encoding device synchronized with a common reference clock, an apparatus for processing video data is provided, the apparatus including a memory and a processor implemented in circuitry and coupled to the memory. In some examples, more than one processor may be coupled to the memory. The processor is configured to: generate a video bitstream; insert intra-frame decoded video data into the video bitstream according to a reference clock shared with at least one other encoding device (e.g., by the apparatus, which may include the encoding device). The reference clock defines a scheduling for interleaving intra-frame decoded video from the encoding device and the at least one other encoding device. The processor is also configured to: receive feedback information indicating that at least a portion of a video slice of the video bitstream is missing or corrupted. The processor is further configured to: send a request to adapt the reference clock in response to the feedback information to allow the encoding device to insert intra-frame decoded video data into the video bitstream in unscheduled time slots. The processor is also configured to: receive an indication that the reference clock has been updated to define an updated scheduling. The processor is also configured to insert the intra-frame decoded video data into the video bitstream based on the updated scheduling and according to the updated reference clock.
[0038] In another example of synchronization between an encoding device and a common reference clock, a non-transitory computer-readable medium of a computing device is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: generate a video bitstream, wherein intra-frame decoded video data is inserted into the video bitstream according to a reference clock shared with at least one other encoding device, the reference clock defining a schedule for interleaving intra-frame decoded video from the encoding device and the at least one other encoding device; obtain feedback information indicating that at least a portion of a video slice of the video bitstream is missing or corrupted; in response to the feedback information, send a request to adapt the reference clock to allow the encoding device to insert intra-frame decoded video data into the video bitstream at unscheduled time slots; receive an indication that the reference clock has been updated to define an updated schedule; and, based on the updated schedule, insert the intra-frame decoded video data into the video bitstream according to the updated reference clock.
[0039] In another example of synchronization between an encoding device and a common reference clock, an apparatus for processing video data is provided. The apparatus includes: a unit for generating a video bitstream; and a unit for inserting intra-frame decoded video data into the video bitstream according to a reference clock shared with at least one other encoding device (e.g., by the encoding device). The reference clock defines a schedule for interleaving intra-frame decoded video from the encoding device and the at least one other encoding device. The apparatus further includes: a unit for obtaining feedback information indicating that at least a portion of a video slice of the video bitstream is missing or corrupted. The apparatus further includes: a unit for sending a request to adapt the reference clock in response to the feedback information, allowing the encoding device to insert intra-frame decoded video data into the video bitstream in unscheduled time slots. The apparatus further includes: a unit for receiving an indication that the reference clock has been updated to define an updated schedule; and a unit for inserting the intra-frame decoded video data into the video bitstream according to the updated reference clock based on the updated schedule.
[0040] In some aspects, based on the updated scheduling, the at least one other encoding device delays the scheduling of intra-frame decoded video relative to the time slots of the previously scheduled time slots defined by the reference clock.
[0041] In some aspects, multiple encoding devices are synchronized with the reference clock. In such aspects, each of the multiple encoding devices can be assigned a different time reference and transmit encoded data according to the different time references. In some cases, a first time reference assigned to the encoding device is different from a second time reference assigned to the at least one other encoding device.
[0042] In some aspects, the unscheduled time slots deviate from multiple time slots defined by the reference clock for the encoding device.
[0043] In some respects, the updated reference clock is shared with the at least one other encoding device.
[0044] In some aspects, the intra-frame decoded video data includes one or more intra-frame decoded video frames. In some aspects, the intra-frame decoded video data includes one or more intra-frame decoded video slices. In some aspects, the intra-frame decoded video data includes an intra-frame refresh period, the intra-frame refresh period including at least one video frame. For example, each video frame in the at least one video frame may include one or more intra-frame decoded video slices.
[0045] In some aspects, the feedback information is provided from a computing device. In some implementations, the computing device includes an extended reality display device, and the encoding device is part of a server. The encoding device is configured to generate the video bitstream for display on the extended reality display device based on motion information received by the encoding device from the extended reality display device.
[0046] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to define the scope of the claimed subject matter. The subject matter should be understood by referring to the appropriate portions of the entire specification, any or all of the drawings, and each claim.
[0047] The foregoing, as well as other features and embodiments, will become more apparent upon reference to the following description, claims, and drawings. Attached Figure Description
[0048] The illustrative embodiments of this application are described in detail below with reference to the accompanying drawings:
[0049] Figure 1 This is a block diagram illustrating examples of encoding and decoding devices based on some examples;
[0050] Figure 2 This is a block diagram illustrating an example of a split rendering system based on some examples of Extended Reality (XR);
[0051] Figure 3 This is a diagram illustrating an example of a video decoding structure using Strict Constant Bit Rate (CBR) based on some examples;
[0052] Figure 4 This is a diagram illustrating an example of a video decoding structure using a dynamic intra-frame decoded frame (I-frame) with a strictly constant bit rate (CBR), according to some examples.
[0053] Figure 5 This is a diagram illustrating an example of a video decoding structure with an intra-frame refresh period in an error-free link, based on some examples;
[0054] Figures 6A-6D This is a diagram illustrating an example of a video decoding structure using dynamic I-slices, based on some examples;
[0055] Figures 7A-7F These are diagrams illustrating other examples of video decoding structures using dynamic I-slices, based on some examples;
[0056] Figures 8A-8H These are diagrams illustrating other examples of video decoding structures using dynamic I-slices, based on some examples;
[0057] Figure 9A and Figure 9B This is a diagram illustrating an example of a video decoding structure using dynamically individual I-slices, based on some examples;
[0058] Figure 10 This is a diagram illustrating an example of a system including an encoding device synchronized with a reference clock, based on some examples;
[0059] Figure 11 This is a diagram illustrating an example of a video decoding structure with periodic I-frames, where two encoding devices synchronized with a reference clock are shown, according to some examples.
[0060] Figure 12 This is a diagram illustrating another example of a video decoding structure with dynamic I-frames, where two encoding devices synchronized with a reference clock are shown, based on some examples.
[0061] Figure 13 This is a flowchart illustrating an example of a process for processing video data, based on some examples;
[0062] Figure 14 This is a flowchart illustrating another example of a process for processing video data, based on some examples;
[0063] Figure 15 This is a flowchart illustrating another example of a process for processing video data, based on some examples; and
[0064] Figure 16 This is an example computing device architecture that can implement the various technologies described in this article. Detailed Implementation
[0065] Certain aspects and embodiments of this disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments can be applied independently, and some can be applied in combination. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The accompanying drawings and description are not intended to be limiting.
[0066] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of exemplary embodiments will provide those skilled in the art with feasible descriptions for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0067] This document describes systems and techniques for adaptively controlling encoding devices (such as video encoders or other types of encoding devices) based on feedback information (indicating video data with missing or corrupted video packets) provided from client devices to video encoders. For example, a video encoder can use the feedback information to determine when to adaptively insert intra-decoded frames (also referred to as pictures) and / or intra-decoded slices into the encoded video bitstream. Systems and techniques for synchronizing encoding devices with a common reference clock (e.g., set by a wireless access point or other device) and other encoding devices are also described, which can help schedule video traffic in multi-user environments.
[0068] Video decoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques may include applying different prediction modes (including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques) to reduce or remove redundancy inherent in the video sequence. A video encoder can segment each frame of the original video sequence into rectangular regions, referred to as video blocks or decoding units (described in more detail below). Specific prediction modes can be used to encode these video blocks.
[0069] Video blocks can be divided into one or more smaller blocks in one or more ways. Blocks may include decode tree blocks, prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise specified, the reference to “block” generally refers to such a video block (e.g., a decode tree block, decode block, prediction block, transform block, or other suitable block or sub-block, as will be understood by one of ordinary skill in the art). Furthermore, each of these blocks may also be interchangeably referred to herein as a “unit” (e.g., decode tree unit (CTU), decode unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may refer to a decoded logic unit encoded in a bitstream, while a block may refer to a portion of the video frame buffer to which the process is targeted.
[0070] For inter-frame prediction mode, the video encoder searches for a block similar to the block being encoded in a frame (or picture) located at another time position; this is called a reference frame or reference picture. The video encoder can restrict the search to a specific spatial displacement from the block to be encoded. A best match can be located using two-dimensional (2D) motion vectors that include horizontal and vertical displacement components. For intra-frame prediction mode, the video encoder uses spatial prediction techniques to form the predicted block based on data from previously encoded adjacent blocks within the same picture.
[0071] A video encoder can determine prediction error. For example, a prediction can be determined as the difference between the pixel values in the block being encoded and the predicted block. Prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a Discrete Cosine Transform (DCT) or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form a decoded representation of the video sequence. In some cases, the video encoder can perform entropy decoding on the syntax elements, further reducing the number of bits required for its representation.
[0072] A video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., prediction blocks) for decoding the current frame. For example, a video decoder can add the predicted block to the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis function using quantization coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.
[0073] The techniques described herein can be applied to one or more of a wide variety of block-based video decoding techniques, where video is reconstructed block by block. For example, the techniques described herein can be applied to any existing video codec (e.g., High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), or other suitable existing video codecs), and / or can be efficient decoding tools for any video decoding standard under development and / or future video decoding standards, such as Universal Video Decoding (VVC), Joint Exploratory Model (JEM), VP9, AV1, and / or other video decoding standards under development or to be developed.
[0074] Figure 1This is a block diagram illustrating an example of a system 100 including encoding device 104 and decoding device 112. Encoding device 104 may be part of a source device, and decoding device 112 may be part of a receiving device (also referred to as a client device). The source device and / or receiving device may include electronic devices such as server devices in a server system including one or more server devices (e.g., an extended reality (XR) split rendering system, a video streaming server system, or other suitable server system), head-mounted displays (HMDs), head-up displays (HUDs), smart glasses (e.g., virtual reality (VR) glasses, augmented reality (AR) glasses, or other smart glasses), mobile or landline phones (e.g., smartphones, cellular phones, etc.), desktop computers, laptops or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, Internet Protocol (IP) cameras, or any other suitable electronic devices. In an illustrative example, such as Figure 2 As shown and described in more detail below, the source device may include a server, and the receiving device may include an XR client device (e.g., an HMD or other suitable device) in an XR split rendering system. In some examples, the source device and the receiving device may include one or more wireless transceivers for wireless communication.
[0075] The components of system 100 may include and / or may be implemented using circuitry or other electronic hardware, which may include one or more programmable circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.
[0076] Although system 100 is shown to include certain components, those skilled in the art will recognize that system 100 may include more than [other components]. Figure 1The components shown may include more or fewer components. For example, in some cases, system 100 may also include one or more memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices) in addition to storage devices 108 and 118; one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) that communicate with and / or are electrically connected to one or more memory devices; one or more wireless interfaces for performing wireless communication (e.g., including one or more transceivers and baseband processors for each wireless interface); one or more wired interfaces for performing communication over one or more hardwired connections (e.g., serial interfaces such as Universal Serial Bus (USB) inputs, lighting connectors, and / or other wired interfaces); and / or Figure 2 Other components not shown.
[0077] The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcasting or transmission, encoding digital video for storage on data storage media, decoding digital video stored on data storage media, or other applications. In some examples, system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.
[0078] Encoding device 104 (or encoder) can be used to encode video data using video decoding standards or protocols to generate an encoded video bitstream. Examples of video decoding standards include ITU-T H.261, ISO / IEC MPEG-1 visual, ITU-T H.262 or ISO / IEC MPEG-2 visual, ITU-T H.263, ISO / IEC MPEG-4 visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Decoding (SVC) and Multi-View Video Decoding (MVC) extensions), and High Efficiency Video Decoding (HEVC) or ITU-T H.265. Various extensions to HEVC exist that involve multi-layer video decoding, including range and screen content decoding extensions, 3D video decoding (3D-HEVC), multi-view extensions (MV-HEVC), and scalable extensions (SHVC). The ITU-T Video Decoding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) Joint Collaborative Working Group on Video Decoding (JCT-VC), as well as the Joint Collaborative Working Group on 3D Video Decoding Extensions (JCT-3V), have developed HEVC and its extensions.
[0079] MPEG and ITU-T VCEG have also formed the Joint Video Exploration Group (JVET) to explore and develop new video decoding tools for next-generation video decoding standards, named Universal Video Decoding (VVC). The reference software is called the VVC Test Model (VTM). VVC aims to provide significant improvements in compression performance compared to the existing HEVC standard, thereby facilitating the deployment of higher-quality video services and emerging applications (e.g., 360° omnidirectional immersive multimedia, high dynamic range (HDR) video, etc.). VP9 and AV1 are other video decoding standards that can be used.
[0080] Many of the embodiments described herein can be performed using video codecs such as VTM, VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein are also applicable to other decoding standards, such as MPEG, JPEG (or other decoding standards for still images), VP9, AV1, their extensions, or other suitable decoding standards that are already available or not yet available or under development. Therefore, although the techniques and systems described herein may be described with reference to a specific video decoding standard, it will be recognized by those skilled in the art that this description should not be construed as applicable only to that particular standard.
[0081] Reference Figure 1 Video source 102 can provide video data to encoding device 104. Video source 102 can be part of a source device, or it can be part of a device other than a source device. Video source 102 can include video capture devices (e.g., cameras, camera phones, video phones, etc.), video archive units containing stored video, video servers or content providers that provide video data, video feed interfaces that receive video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or any other suitable video source.
[0082] Video data from video source 102 may include one or more input images. Images may also be referred to as "frames." An image or frame is a still image that, in some cases, is part of the video. In some examples, data from video source 102 may be still images that are not part of the video. In HEVC, VVC, and other video decoding specifications, a video sequence may include a series of images. An image may include three sample arrays, denoted as SL, S... Cb and S Cr S L It is a two-dimensional array of brightness samples, S Cb It is a two-dimensional array of Cb chromaticity samples, and S CrIt is a two-dimensional array of chrominance samples. Chrominance samples may also be referred to as "chroma" samples in this paper. In other cases, the image may be monochrome and may consist only of an array of luminance samples.
[0083] Encoding device 104's encoder engine 106 (or encoder) encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. The decoded video sequence (CVS) comprises a series of access units (AUs) starting with an AU that has a random access point picture in the base layer and certain attributes, up to the next AU that has a random access point picture in the base layer and certain attributes, and excluding that next AU. For example, certain attributes of the random access point picture that starts the CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (where the RASL flag is equal to 0) does not start a CVS. An access unit (AU) comprises one or more decoded pictures and control information corresponding to decoded pictures sharing the same output time. Decoded slices of pictures are encapsulated at the bitstream level as data units called Network Abstraction Layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs, which include NAL units. Each NAL unit has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header use specified bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, Real-Time Transport (RTP) protocols, file formats, etc.
[0084] The HEVC standard contains two types of NAL units: Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that forms the decoded video bitstream. For example, a VCL NAL unit contains a sequence of bits that form the decoded video bitstream. A VCL NAL unit may include a slice or fragment of the decoded picture data (described below), and non-VCL NAL units contain control information associated with one or more decoded pictures. In some cases, NAL units may be referred to as packets. A HEVC AU includes: VCL NAL units containing decoded picture data, and non-VCL NAL units (if any) corresponding to the decoded picture data. Among other information, non-VCL NAL units may also contain a set of parameters with high-level information associated with the encoded video bitstream. For example, the parameter set may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Each slice or other part of the bitstream can reference a single active PPS, SPS, and VPS to allow the decoding device 112 to access information that can be used to decode the slice or other part of the bitstream.
[0085] NAL units can contain bit sequences that form a decoded representation of video data (e.g., a decoded video bitstream, a CVS of a bitstream, etc.), such as a decoded representation of images in a video. Encoder engine 106 generates a decoded representation of an image by dividing each image into multiple slices. Slices are independent of other slices, allowing information in that slice to be decoded without depending on data from other slices within the same image. A slice includes one or more segments, comprising independent segments and (if present) one or more dependent segments that depend on previous segments.
[0086] In HEVC, the slice is then divided into decoder tree blocks (CTBs) for luma and chroma samples. One or more CTBs for luma samples and one CTB for chroma samples, along with the syntax used for the samples, are called decoder tree units (CTUs). CTUs can also be called "tree blocks" or "maximum decoder units" (LCUs). A CTU is the basic processing unit used for HEVC encoding. A CTU can be subdivided into multiple decoder units (CUs) of different sizes. A CU contains an array of luma and / or chroma samples called a decoder block (CB).
[0087] Luminance and chrominance CBs can be further subdivided into prediction blocks (PBs). A PB is a sample block of the luminance or chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy prediction (when available or enabled for use). A luminance PB and one or more chrominance PBs, along with their associated syntax, form a prediction unit (PU). For inter-frame prediction, the set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU, along with the inter-frame prediction for the luminance PB and one or more chrominance PBs. Motion parameters can also be referred to as motion information. CBs can also be subdivided into one or more transform blocks (TBs). A TB represents a square block of samples of the chrominance component, where the same two-dimensional transform is applied to decode the prediction residual signal. A transform unit (TU) represents a TB of luminance and chrominance samples and the corresponding syntax elements.
[0088] The size of the CU corresponds to the size of the decoding mode and can be square. For example, the size of the CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other suitable size up to the corresponding CTU size. The phrase "N x N" is used herein to refer to the pixel size of the video block in both the vertical and horizontal dimensions (e.g., 8 pixels x 8 pixels). Pixels in a block can be arranged in rows and columns. In some embodiments, the block may not have the same number of pixels in the horizontal direction as it does in the vertical direction. The syntax data associated with the CU can describe, for example, the segmentation of the CU into one or more PUs. The segmentation mode can differ between whether the CU is encoded using intra-frame prediction mode or inter-frame prediction mode. PUs can be segmented into non-square shapes. The syntax data associated with the CU can also describe, for example, the segmentation of the CU into one or more TUs according to the CTU. TUs can be square or non-square shapes.
[0089] According to the HEVC standard, transform units (TUs) can be used to perform transforms. The TU can vary for different CUs. The size of the TU can be set based on the size of the PU within a given CU. The TU can be the same size as the PU or smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to TUs. The pixel differences associated with the TU can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.
[0090] Once the video data is segmented into Units (CUs), the encoder engine 106 uses a prediction mode to predict each Unit (PU). The prediction unit or block is then subtracted from the original video data to obtain the residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. Prediction modes can include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). Intra-frame prediction utilizes the correlation between spatially adjacent samples within a picture. For example, using intra-frame prediction, each PU is predicted based on, for example, DC prediction to find an average value for the PU, planar prediction to make a planar surface conform to the PU, orientation prediction to infer from adjacent data, or any other suitable prediction type, based on adjacent image data within the same picture. Inter-frame prediction uses temporal correlations between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted based on image data in one or more reference pictures (in the output order before or after the current picture) using motion-compensated prediction. For example, a decision can be made at the CU level whether to use inter-picture prediction or intra-picture prediction to decode a picture region.
[0091] Encoder engine 106 and decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, a video decoder (such as encoder engine 106 and / or decoder engine 116) segments the image into multiple decoder tree units (CTUs) (wherein, the CTB of a luminance sample and one or more CTBs of a chrominance sample, along with the syntax used for the samples, are referred to as a CTU). The video decoder can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple segmentation types, such as the distinction between CU, PU, and TU in HEVC. The QTBT structure includes two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoder units (CUs).
[0092] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. A ternary tree partition is a partition in which a block is split into three sub-blocks. In some examples, a ternary tree partition divides a block into three sub-blocks without partitioning the original block through the center. The partitioning type in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.
[0093] In some examples, the video decoder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video decoder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBT and / or MTT structures for the respective chrominance components).
[0094] Video decoders can be configured to use quadtree segmentation according to HEVC, QTBT segmentation, MTT segmentation, or other segmentation structures. For illustrative purposes, the description herein may refer to QTBT segmentation. However, it should be understood that the techniques of this disclosure can also be applied to video decoders configured to use quadtree segmentation or other types of segmentation.
[0095] In some examples, one or more slices of an image are assigned slice types. Slice types include intra-frame decoded slices (I-slices), inter-frame decoded P-slices, and inter-frame decoded B-slices. An I-slice (intra-frame decoded frame, independently decodeable) is a slice of an image that is decoded solely by intra-frame prediction, and is therefore independently decodeable because an I-slice only requires intra-frame data to predict any prediction unit or prediction block of the slice. A P-slice (one-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and one-way inter-frame prediction. Each prediction unit or prediction block within a P-slice is decoded using either intra-frame prediction or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted using only one reference image, and therefore the reference sample is from only one reference region within a frame. A B-slice (two-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and inter-frame prediction (e.g., dual prediction or single prediction). Bidirectional prediction of a B-slice's prediction unit or block can be performed based on two reference images, where each image contributes a reference region, and the sample sets of the two reference regions are weighted (e.g., with equal weights or different weights) to generate the prediction signal for the bidirectional prediction block. As explained above, a slice of an image is decoded independently. In some cases, an image may be decoded into only one slice.
[0096] As mentioned above, intra-image prediction utilizes the correlation between spatially adjacent samples within the image. Multiple intra-prediction modes exist (also referred to as "intra-frame modes"). In some examples, intra-frame prediction for a luma patch includes 35 modes, comprising a planar mode, a DC mode, and 33 angular modes (e.g., a diagonal intra-prediction mode and the angular modes adjacent to the diagonal intra-prediction mode). The 35 intra-prediction modes are indexed as shown in Table 1 below. In other examples, more intra-frame modes may be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with angular modes may differ from those used in HEVC.
[0097] Intra-prediction mode Associated Names 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34
[0098] Table 1 – Specification of Intra-Frame Prediction Modes and Associated Names
[0099] Inter-image prediction uses temporal correlations between images to derive motion-compensated predictions for image sample blocks. Using a translational motion model, the position of a block in a previously decoded image (reference image) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the current block's position, and Δy specifies the vertical displacement of the reference block relative to the current block's position. In some cases, the motion vector (Δx, Δy) can be integer sample precision (also referred to as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of the reference frame. In other cases, the motion vector (Δx, Δy) can have fractional sample precision (also referred to as fractional pixel precision or non-integer precision) to more accurately capture the motion of the underlying object without being limited to an integer pixel grid of the reference frame. The precision of the motion vector can be expressed by the quantization level of the motion vector. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample accuracy, interpolation is applied to the reference image to derive the predicted signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values at fractional positions. The previously decoded reference image is indicated by a reference index (refIdx) against a list of reference images. The motion vector and reference index can be referred to as motion parameters. Inter-image predictions can be performed, including single and double predictions.
[0100] In the case of inter-frame prediction using dual prediction, two sets of motion parameters (Δx0, y0, refIdx0 and Δx1, y1, refIdx1) are used to generate two motion-compensated predictions (from the same reference image or possibly from different reference images). For example, in the case of dual prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are then combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference images that can be used in dual prediction are stored in two separate lists (represented as list 0 and list 1). The motion parameters can be derived at the encoder using a motion estimation process.
[0101] When using single prediction for inter-frame prediction, a set of motion parameters (Δx0, y0, refIdx0) is used to generate motion-compensated predictions based on a reference image. For example, in the case of single prediction, each prediction block uses at most one motion-compensated prediction signal and generates P prediction units.
[0102] The prediction unit (PU) may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU is encoded using intra-frame prediction, the PU may include data describing the intra-frame prediction mode used for the PU. As another example, when the PU is encoded using inter-frame prediction, the PU may include data defining the motion vectors used for the PU. The data defining the motion vectors used for the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution used for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the reference index, a list of reference pictures used for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.
[0103] After performing intra-frame prediction and / or inter-frame prediction, encoding device 104 may perform transform and quantization. For example, after prediction, encoder engine 106 may compute a residual value corresponding to the PU. The residual value may include the pixel difference between the current pixel block (PU) being decoded and the prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating a prediction block (e.g., using inter-frame prediction or intra-frame prediction), encoder engine 106 may generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the differences between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of the pixel values.
[0104] Block transforms are used to transform any residual data that may remain after prediction is performed. These block transforms can be based on discrete cosine transform, discrete sine transform, integer transform, wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., sizes of 32x32, 16x16, 8x8, 4x4, or other suitable sizes) can be applied to the residual data in each CU. In some embodiments, TUs can be used for transform and quantization processes implemented by encoder engine 106. A given CU having one or more PUs may also include one or more TUs. As described further in detail below, residual values can be transformed into transform coefficients using block transforms, and then quantized and scanned using TUs to produce serialized transform coefficients for entropy decoding.
[0105] In some embodiments, after intra-frame prediction or inter-frame prediction decoding using the PU of the CU, the encoder engine 106 can compute residual data for the TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). The TU may include coefficients in the transform domain after applying the block transform. As previously described, the residual data may correspond to the pixel difference between a pixel in the uncoded image and the predicted value corresponding to the PU. The encoder engine 106 can form a TU including the residual data for the CU, and can then transform the TU to produce transform coefficients for the CU.
[0106] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by reducing the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.
[0107] Once quantization is performed, the decoded video bitstream includes quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other syntax data). The different elements of the decoded video bitstream can then be entropy-coded by encoder engine 106. In some examples, encoder engine 106 can scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy-coded. In some examples, encoder engine 106 can perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), encoder engine 106 can entropy-code that vector. For example, encoder engine 106 can use context-adaptive variable-length decoding, context-adaptive binary arithmetic decoding, syntax-based context-adaptive binary arithmetic decoding, probabilistic interval segmentation entropy decoding, or another suitable entropy coding technique.
[0108] The output 110 of encoding device 104 can send NAL units constituting the encoded video bitstream data to decoding device 112 of receiving device via communication link 120. The input 114 of decoding device 112 can receive the NAL units. Communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces, and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi). TM Radio frequency (RF), UWB, WiFi Direct, Cellular, Long Term Evolution (LTE), WiMax TM Wired networks can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable Ethernet, digital signal line (DSL), etc.). Various devices can be used to implement wired and / or wireless networks, such as base stations, routers, access points, bridges, gateways, switches, etc. Encoded video bitstream data can be modulated and transmitted to receiving devices according to communication standards such as wireless communication protocols.
[0109] In some examples, encoding device 104 may store encoded video bitstream data in storage unit 108. Output 110 may retrieve encoded video bitstream data from encoder engine 106 or from storage unit 108. Storage unit 108 may include any of a wide variety of distributed or locally accessible data storage media. For example, storage unit 108 may include hard disk drives, storage disks, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage unit 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In other examples, storage unit 108 may correspond to a file server or another intermediate storage device that may store encoded video generated by a source device. In such cases, receiving device including decoding device 112 may access the stored video data from the storage device via streaming or downloading. File server may be any type of server capable of storing encoded video data and sending such encoded video data to receiving devices. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data via any standard data connection, including an internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage unit 108 can be streaming, downloading, or a combination thereof.
[0110] Input 114 of decoding device 112 receives encoded video bitstream data and can provide the video bitstream data to decoder engine 116 or to storage unit 118 for later use by decoder engine 116. For example, storage unit 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including decoding device 112 can receive encoded video data to be decoded via storage unit 108. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the receiving device. The communication medium used to transmit the encoded video data can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, a switch, a base station, or any other means for facilitating communication from the source device to the receiving device.
[0111] Decoder engine 116 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements from one or more decoded video sequences that make up the encoded video data. Decoder engine 116 can then rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is then passed to the prediction stage of decoder engine 116. Decoder engine 116 then predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (the residual data).
[0112] Decoding device 112 can output decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, video destination device 122 may be part of a receiving device that includes decoding device 112. In some aspects, video destination device 122 may be part of a separate device, distinct from the receiving device.
[0113] Decoding device 112 can monitor the encoded video bitstream received from encoding device 104 and can detect when packets become lost or corrupted. For example, the video bitstream (or media file containing the video bitstream) may include corrupted or lost video frames (or images) in the encoded data. A lost frame may occur when all the encoded data of a lost frame is missing. Corrupted frames may occur in different ways. For example, a frame may become corrupted when a packet of a frame or a portion of the encoded data of that frame is missing. As another example, a frame may become corrupted when it is part of an inter-frame prediction chain and some other encoded data in the inter-frame prediction chain is missing or corrupted, making the frame unable to be decoded correctly. For example, if one or more reference frames are lost or corrupted, inter-frame decoding frames that rely on one or more reference frames for prediction may be undecodeable.
[0114] In response to the detection of frames, slices, or other video with missing packets, decoding device 112 may send feedback information 124 to encoding device 104. Feedback information 124 may indicate that a video frame, video slice, a portion thereof, or other video information is missing or corrupted (referred herein to as a "corrupted video frame," "corrupted video slice," or other type of corrupted video data). Encoder engine 106 may use feedback information 124 to determine whether to adaptively insert intra-predictive decoded frames (also referred to as intra-decoded frames or pictures) or intra-predictive decoded slices (also referred to as intra-decoded slices) into the encoded video bitstream. For example, as described in more detail below, encoder engine 106 may dynamically insert I-frames into the encoded video bitstream and / or dynamically insert I-slices (e.g., individual I-slices with intra-refresh periods) into the encoded video bitstream based on feedback information 124. In some cases, in response to the detection of corrupted video data, the receiving device, including the decoding device 112, may rely on error concealment (e.g., asynchronous time warp (ATW) error concealment) until error-free intra-decoded frames, slices, or other video data are received.
[0115] In some embodiments, the video encoding device 104 and / or the video decoding device 112 may be integrated with the audio encoding device and the audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary for implementing the above-described decoding techniques, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated as part of a combined encoder / decoder (codec) in the respective device.
[0116] Figure 1The example system shown is merely an illustrative example of the techniques that can be used herein. The techniques for processing video data using the techniques described herein can be implemented by any digital video encoding and / or decoding device. Although, in general, the techniques of this disclosure are implemented by video encoding or video decoding devices, the techniques can also be implemented by a combined video encoder / decoder, commonly referred to as a "CODEC". Furthermore, the techniques of this disclosure can also be implemented by a video preprocessor. The source device and receiving device are merely examples of such decoding devices, wherein the source device generates decoded video data for transmission to the receiving device. In some examples, the source device and receiving device can operate in a substantially symmetrical manner, such that each device includes both video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0117] As described above, in some examples, the source device may include a server, and the receiving device may include an extended reality (XR) client device in an XR system (e.g., an HMD, smart glasses, or other suitable device). XR includes augmented reality (AR), virtual reality (VR), mixed reality (MR), and so on. Each of these forms of XR allows users to experience or interact with virtual content, sometimes in combination with real content.
[0118] Split-rendering infinite XR systems are a type of XR system that splits the XR processing burden between the server side and the client side (e.g., the side with XR headsets such as HMDs). Figure 2 This is a block diagram illustrating an example of an XR split rendering system 200 including a server-side component 202 (corresponding to a server component) and a client-side component 220 (corresponding to a client device or receiving device component). The XR split rendering system 200 can split the processing burden of XR applications (e.g., virtual reality (VR), augmented reality (AR), mixed reality (MR), or other XR applications) between the server-side component 202 and the client-side component 220. For example, the client device on the client-side component 220 can perform on-device processing enhanced by the computing resources located on the server-side component 202 via a wireless connection (e.g., using a broadband connection such as 4G, 5G, etc., using a WiFi connection, or other wireless connection). In one example, the server-side component 202 may be located at the cloud edge in a wireless network (e.g., a broadband network such as a 5G network).
[0119] The split processing between the client-side 220 and the server-side 202 enables photorealistic, high-quality, and immersive experiences. Various advantages are gained by performing content rendering and other processing on the server-side 202. For example, the server-side 202 provides greater processing capacity and heat dissipation compared to client devices on the client-side 220, which is advantageous when the client device is a lightweight XR headset (e.g., a head-mounted display) that may have limited processing power, battery life, and / or heat dissipation capabilities. The higher processing power of the server-side 202 allows resource-intensive applications (e.g., multiplayer games, multi-user video conferencing, multi-user video applications, etc.) to be rendered on the server-side 202 and displayed on the client-side 220 with high quality and low latency.
[0120] Server-side 202 includes various components, including a media source engine 204, a video and audio encoder 206, a depth / motion encoder 208, and a rate / error adaptive engine 210. Client-side 220 also includes various components, including a video and audio decoder 222, a depth / motion decoder 224, a post-processing engine 226, and a display 228. Low-latency transport links 212, 214, and 216 are also used to transmit data between server-side 202 and client-side 220.
[0121] The components of the XR split rendering system 200 may include circuitry or other electronic hardware, and / or may be implemented using circuitry or other electronic hardware, which may include one or more programmable circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.
[0122] Although the XR split rendering system 200 is shown to include certain components, those skilled in the art will recognize that the XR split rendering system 200 may include components that are compatible with those in the present invention. Figure 2The components shown may be more or fewer than those shown. For example, the XR split rendering system 200 may also include one or more input devices and one or more output devices (not shown). In some cases, the XR split rendering system 200 may also include one or more memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) communicating with and / or electrically connected to one or more memory devices, one or more wireless interfaces for performing wireless communication (e.g., including one or more transceivers and baseband processors for each wireless interface), one or more wired interfaces for performing communication via one or more hardwired connections (e.g., serial interfaces such as Universal Serial Bus (USB) inputs, lighting connectors, and / or other wired interfaces), and / or Figure 2 Other components not shown.
[0123] In some examples, the client device on client-side 220 may include an XR headset, such as an HMD, XR glasses, or other suitable head-mounted device with a display. In some cases, the XR headset may perform client-side functions required to display XR content. In some examples, the client device may include multiple devices, such as an XR headset that communicates wired or wirelessly with computing devices (such as smartphones, tablets, personal computers, and / or other devices). For example, the computing device may perform client-side processing functions required to prepare XR content for rendering or display (e.g., error hiding, de-warping of one or more images in one or more video frames, processing to minimize motion-to-photon latency, and other functions), and the XR headset may display content based on the processing performed by the computing device.
[0124] Using a client device, when a user of the XR headset moves their head, on-device processing (e.g., on-device processing of the XR headset or a wirelessly connected or wired-connected computing device to the XR headset (not shown)) determines the head pose and transmits the pose to server-side 202 (e.g., media source engine 204, video and audio encoder 206, and / or depth / motion encoder 208) via low-latency transmission link 212. Low-latency transmission link 212 may include a high-quality-of-service communication link (e.g., a 5G link, a WiFi link, or other communication link). Media source engine 204 can use the head pose to partially render the next video and audio frames, and can output the video and audio frames to video and audio encoder 206. Media source engine 204 can also use the head pose to render depth and motion information, which can be output to depth / motion encoder 208.
[0125] Media source engine 204 may include any media data source, such as a game engine, streaming media source, video-on-demand content source, or other media source. In some implementations, media source engine 204 may provide immersive media content that offers an immersive user experience to users on client devices. Examples of immersive media content include 360-degree (or virtual reality (VR)) video, 360-degree (or VR) video game environments, and other virtual or synthetic visualizations. For such immersive media content, the user's head posture (e.g., corresponding to the position and / or orientation of the XR headset) reflects the viewing direction and / or field of vision relative to the immersive content. For example, when an XR headset user turns their head to the right, media source engine 204 may render immersive media content to adjust the virtual scene to where the user expects to see based on the new head position and orientation. In some cases, eye gaze and / or eye focus (e.g., for depth-based features) may be used to determine the user's interaction with the immersive environment. For example, the media source engine 204 can provide and use eye gaze information that indicates where the user is looking in a virtual scene to determine the selection of objects in the scene, determine the part of the scene to be emphasized, present enhanced content in the scene, make objects (e.g., characters) in the scene react, and / or perform other operations based on eye gaze information.
[0126] Video and audio encoder 206 encodes video and audio data and transmits the encoded audio-video data to client-side 220 via low-latency transmission link 214. The encoded data can also be transmitted to rate / error adaptive engine 210. Rate / error adaptive engine 210 can adjust the bit rate based on communication channel information. Depth / motion encoder 208 encodes depth and motion data and transmits the encoded depth and motion data to client-side 220 via low-latency transmission link 216.
[0127] On the client side 220, video and audio decoders 222 decode the received audio-video data. Depth / motion decoders 224 can decode the received depth and motion data. The decoded audio-video data, along with the decoded depth and motion data, can be provided to post-processing engine 226. Based on the latest head pose (generated at a high frequency by the client device), post-processing engine 226 can perform any further rendering and adjustments required. Post-processing can include error concealment (e.g., asynchronous time warp (ATW) or other types of error concealment) to hide errors in the video and / or audio data, spatial warp (frame rate smoothing), de-warping of images in video frames, processing to minimize motion-to-photon latency, and other post-processing functions. Post-processing can be performed on the client device to meet latency thresholds (e.g., 14 milliseconds, 20 milliseconds, or other appropriate thresholds) required to avoid user discomfort. For example, high motion-to-photon latency can prevent true immersion in a virtual environment and may cause user discomfort.
[0128] Motion-to-photon delay is the delay between the occurrence of user movement and the display of corresponding content. For example, motion-to-photon delay can include the length of time between when a user performs a movement (e.g., turning their head to the right or left) and when the display shows the appropriate content for that particular movement (e.g., the content on the HMD subsequently moves to the right or left). The term "photon" is used to specify that all parts of the display system are involved in this process until a photon is emitted from the display.
[0129] To provide users with an immersive experience, minimizing motion-to-photon latency is crucial. Humans are highly sensitive to such latency, and excessive lag can cause discomfort or disorientation. In some cases, users may not detect lag of up to 20 milliseconds for VR content. Low motion-to-photon latency is essential for providing a good user experience and preventing motion sickness or other adverse effects on client devices such as head-mounted displays or HMDs.
[0130] For infinitely split rendering systems (such as XR split rendering system 200), there may be even more latency because the client sends events, gestures, user interactions, and other information to the server, and the server renders the XR scene based on this information, encodes the content, and sends the rendered content to the client device for decoding and display. This latency can be referred to as "motion-to-render-to-photon latency." For example, motion-to-render-to-photon latency can be the length of time between the user performing a movement, the server side 202 rendering the appropriate content for that specific movement, providing the content to the client side 220, and the display showing the appropriate content for that specific movement.
[0131] The video encoding and decoding components of an XR system can affect its latency. For example, higher bitrate video content may require more bandwidth for transmission compared to lower bitrate video content. Due to the real-time nature of some XR content and the quality requirements of these systems, a constant bitrate (CBR) scheme can be used to ensure that a specific quality is maintained. In some cases, high bitrate video frames (e.g., I-frames and I-slices larger than inter-frame decoded frames) are periodically inserted into the encoded video bitstream, even if these frames may not be needed at such frequencies. As mentioned above, an I-frame is a frame that is decoded solely using intra-frame prediction of the data within the frame. I-frames are independently decodeable because they only require intra-frame data to predict any prediction units or blocks of the frame. P-frames can be decoded using intra-frame prediction and one-way inter-frame prediction, and B-frames can be decoded using intra-frame prediction, one-way inter-frame prediction, or two-way inter-frame prediction. A frame can be divided into multiple slices, where each slice comprises one or more blocks of the frame. Similarly, an I-slice (including an I-frame) is an independent, decodable slice of a frame that is decoded using only the blocks within the slice.
[0132] Techniques are described for adaptively controlling an encoding device (e.g., in a split-rendering XR system or other video-related systems) based on feedback information provided to the encoding device from a client device or other device (e.g., network device, such as an access point (AP), a server on server side 202, or other devices). Feedback information (such as feedback information 124 provided to encoding device 104 from decoding device 112) is also described. In an XR split-rendering system, encoding device 104 may be located on the server side (e.g., Figure 2 The server side 202 and the decoding device 112 can be located on the client side (e.g., ...). Figure 2 (Client-side 220 in the code). Feedback information can indicate to the encoder that video data is missing or corrupted. Although the examples of missing or corrupted video data described below use frames and slices as examples, those skilled in the art will recognize that any part of the video can be detected as having missing groups, such as groups of frames, blocks of frames (e.g., CU, PU, or other blocks), or other suitable video data.
[0133] Video data can be lost or corrupted due to various factors. Video communication systems can perform various steps from encoding to decoding of video data. For example, video can first be compressed by a video encoder (as described above) to reduce the video's data rate. Then, the encoded video bitstream can be segmented into fixed- or variable-length packets and multiplexed with other data types, such as audio and / or metadata. The packets can be transmitted directly over the network or can undergo a channel coding stage (e.g., using forward error correction (FEC) and / or other techniques) to protect the packets from transmission errors. At the receiving device (or client device), the received packets can be channel decoded (e.g., FEC decoding) and unpacked, and the resulting encoded video bitstream can be provided to a video decoder to reconstruct the original video.
[0134] Unless a dedicated link providing guaranteed Quality of Service (QoS) is available between the video source and receiving equipment, data packets may be missing or corrupted, for example, due to bit errors caused by traffic congestion or physical channel damage. A lost frame may occur when all encoded data (e.g., all packets) of the lost frame is missing. Corrupted frames (with corrupted packets or data) can occur under different circumstances. For example, a frame may become corrupted when a portion of the encoded data of a frame (e.g., some packets but not all packets) is missing. Furthermore, compressed video streams are sensitive to transmission errors because the encoding device uses predictive decoding at the source. For example, due to the use of spatiotemporal prediction, a single incorrectly recovered sample can cause errors in the same frame and subsequent samples in subsequent frames (received after the incorrectly recovered sample). In one example, a frame may become corrupted when it is part of an inter-frame prediction chain, and some other encoded data in the inter-frame prediction chain is missing (or corrupted), preventing the frame from being correctly decoded. If one or more reference frames are lost or corrupted, inter-frame decoded frames that rely on predictions made from those reference frames may be undecodeable. As the prediction chain continues, errors may continue to propagate. For example, an error in a previously decoded frame (reference frame) may cause the current frame to be incorrectly decoded, resulting in degradation in the video data of the decoded frame. This degradation in the video data will continue to propagate to subsequent inter-frame decoded frames until an I-frame or I-slice is encountered.
[0135] When packets are missing or corrupted, video quality is degraded and a poor user experience may occur. For example, a user of an XR headset playing a real-time multiplayer game may experience poor video quality when packets are missing or corrupted. In another example, a user watching live video or video served by a streaming service may experience degraded video quality when packets are missing or corrupted. Users can visualize poor or degraded video quality as freezing in the displayed content (where the displayed content is temporarily paused), jittery or unstable visual effects in the displayed content, motion blur, etc.
[0136] In such cases, error-free delivery of data packets can be achieved by allowing the retransmission of missing or corrupted packets using techniques such as Automatic Repeat Request (ARQ). However, retransmitting missing or corrupted packets can lead to unacceptable latency for some real-time applications, such as XR applications, broadcast applications, or other real-time applications. For example, broadcast applications may prevent the use of retransmission algorithms due to network flooding considerations.
[0137] The techniques described in this paper provide a video decoding scheme that allows for dynamic adjustment of the video bitstream to make the data resilient to transmission errors. As described in more detail below, the encoding device can use feedback information from the client device (or other devices, such as network devices) indicating that video data is lost or corrupted to determine when to adaptively insert I-frames, I-slices (with or without intra-frame refresh periods), or other video data into the encoded video bitstream. Since I-frames are predicted using only intra-frame video data, the insertion of I-frames or I-slices can terminate error propagation.
[0138] For any of the techniques described herein, the client device may rely on error hiding until it receives an error-free I-frame or I-slice. In an illustrative example, asynchronous time warp (ATW) error hiding may be performed by a client device in an XR-based system, such as a split-rendering XR system. ATW error hiding can use information from previous frames to hide missing groups. For example, using ATW, previous frames (frames before a corrupted or missing frame) can be warped to generate the current frame. The warp can be performed to reflect head movement since the previous frame was rendered. In one example, updated orientation information for the client device can be retrieved just before applying the time warp, and a transformation matrix can be computed that warps the eye buffers (e.g., based on the position of one or both of the user's eyes and / or a stereoscopic view of the rendered virtual scene) from where they were in the previous frame to where they should be when the current frame is to be displayed. Although the newly generated current image is not exactly the same as the current frame rendered by the rendering engine, displaying the distorted previous frame as the current frame will reduce jitter and other effects compared to displaying the previous frame again, since the previous frame has already been adjusted for head rotation.
[0139] In some applications where client-side buffering of content is permitted, a prolonged period of error hiding by the client device can be tolerated. In other cases, if the client can buffer content, it can request retransmission of corrupted packets, in which case error hiding may not be necessary (at the cost of increased latency). While error hiding (e.g., ATW error hiding) can be performed by the client device until error-free I-frames or I-slices are received, limiting the latency in correcting for frames, slices, or other video data with missing packets can be beneficial (and even critical in some applications). For example, in latency-sensitive systems, such as those delivering real-time content (e.g., in some split-rendering XR systems, live video streaming, and / or broadcasting), the client device may not be able to buffer content locally, and therefore there may be limitations regarding the amount of latency that can be tolerated. In such systems, the amount of time the client device must perform error hiding can be finite. For example, in some latency-sensitive systems, there may be a maximum acceptable amount of time that the client device can perform error hiding.
[0140] As described in more detail below, the systems and techniques described herein can limit the latency of receiving I-frames, I-slices, or other intra-frame decoded data. For example, to expedite the recovery of a video bitstream, a complete I-frame or I-slice can be inserted to immediately terminate error propagation. In another example, I-slices can be spread across the intra-frame refresh cycle to limit the bit rate spikes (and corresponding quality degradation) required for inserting an I-frame. The number of frames included in the intra-frame refresh cycle and / or the size of the I-slices within the intra-frame refresh cycle can be defined based on various factors described below, allowing for a compromise between limiting bit rate spikes and limiting the amount of time the client device must perform error hiding. Each frame in the intra-frame refresh cycle may include a single slice or may include multiple slices. Video content has a bitrate, which is the amount of data transmitted in the video per time interval (e.g., in bits per second). High bitrates result in higher bandwidth consumption and greater latency. For example, if the available bandwidth is less than the bitrate, video reception may be delayed or completely stopped. Therefore, bitrate spikes can cause delays when client devices receive data (e.g., video content sent from server side 202 in XR split rendering system 200 to client side 220). Avoiding bitrate spikes can reduce the bandwidth required for video content transmission, and consequently reduce latency, which can be important in latency-sensitive systems and applications.
[0141] Various techniques for dynamically inserting I-frames or I-slices (with or without intra-frame refresh cycles) into an encoded video bitstream based on feedback information indicating that missing or corrupted packets have been detected will now be described. In some cases, I-frames can be dynamically inserted into the bitstream in systems with a strictly constant bit rate (CBR) coding structure by limiting the frame size using a restricted video buffer validator (VBV) buffer size or a hypothetical reference decoder (HRD) buffer size. Such techniques assume that the encoder can force frames to be I-frames. Figure 3 This is a diagram illustrating an example of a video decoding structure using strict CBR. (See diagram for example.) Figure 3As shown in the strict CBR encoding structure, the default encoding structure using IPPPIPPPI… (where I indicates an I-frame and P indicates a P-frame) has a reference frame (hence, a P-frame). As mentioned above, a strict CBR encoding structure can be obtained using a restricted VBV buffer size or a restricted HRD buffer size. VBV and HRD (used for AVC, HEVC, etc.) are theoretical video buffer models and serve as constraints to ensure that the encoded video bitstream can be properly buffered and played back at the decoder device. By definition, VBV will not overflow or underflow when the input to VBV is a conforming stream, and therefore the encoder must conform to VBV requirements when encoding the bitstream. For the CBR encoding structure, the decoder device's buffer is filled at a constant data rate over time.
[0142] exist Figure 3 In a strict CBR encoding structure, an I-frame is periodically inserted every four frames. Using feedback from a client device (or other device) indicating that video data is missing or corrupted, the encoding device can relax (or even eliminate) the periodic insertion of I-frames into the encoding structure. For example, when the encoding device (e.g., in a server of an XR split rendering system) receives feedback indicating that one or more packets of a video frame are missing or corrupted, the encoding device can react by forcing the next frame in the encoded video bitstream to be encoded as an I-frame.
[0143] As mentioned above, dynamically inserting I-frames based on feedback information allows for a reduction, or in some cases even elimination, of the I-frame insertion period. For example, in some cases, the I-frame insertion period can be increased compared to a typical strict CBR coding structure (e.g., one I-frame every 35 frames instead of one I-frame every four frames). In some cases, I-frames are inserted only in response to feedback indicating that a missing or corrupted video frame has been detected; in this case, periodic insertion is eliminated because I-frames are not periodically inserted into the bitstream. The limited VBV (or HRD) buffer size in a CBR structure can cause a momentary drop in the peak signal-to-noise ratio (PSNR) of I-frames because I-frames have a higher bit rate compared to P-frames. However, given the high operating bit rate of some systems (such as XR systems), the drop in PSNR is negligible.
[0144] Figure 4 This is a diagram illustrating an example of dynamic I-frame insertion in an encoded video bitstream with a strict CBR coding structure. (See diagram for example.) Figure 4As shown, the first frame 402 in the CBR bitstream is a raw I-frame that is periodically inserted into the encoded video bitstream. The periodic rate at which I-frames are inserted into the encoded video bitstream is once every 35 frames (as shown in the gap between the first I-frame 402 and the next I-frame 408). Once a client device detects a missing packet, the encoding device can force the next frame to be an I-frame by performing intra-frame prediction on the next frame. For example, a client device can receive an encoded video bitstream and can detect that one or more packets of video frame 404 are missing or corrupted. In some cases, if packets in a slice are missing or corrupted, the entire slice is undecodeable. In some cases, header information can be used to detect whether one or more packets are missing or corrupted. For example, the encoding device (or other device) can add a header for each packet. The header can include information indicating which slice the header belongs to and also indicating which blocks (e.g., macroblocks, CTUs, CTUs, or other blocks) the slice covers. In an illustrative example, the information in the header can indicate the first and last blocks covered by the packet. The client device can parse the information in the packet header to detect lost packets. Then, the client device can send feedback information to the encoding device indicating that the video frame has lost or missing packets (e.g., Figure 1 (Feedback information 124). Even if the decoding order of video frames is not scheduled as I-frames according to periodic I-frame insertion, the encoding device can force the use of intra-frame prediction to encode video frame 406.
[0145] There is a delay from the time the client device detects missing or corrupted data in video frame 404 to the time the encoding device can insert the forced I-frame 404 into the bitstream. This delay can be based on the amount of time spent detecting the missing or corrupted data, the amount of time spent sending feedback information to the encoding device, and the amount of time spent by the encoding device performing intra-frame prediction and other decoding processes for video frame 406. Based on this delay, a gap exists between the detected missing or corrupted video frame 404 and the dynamically inserted forced I-frame 406. As described above, the client device can perform error hiding on the video frames of the bitstream (e.g., frames between the missing or corrupted frame 404 and the forced I-frame 406) until the forced I-frame 406 is received. Once the I-frame is received, the client device can stop performing error hiding.
[0146] Reducing the rate at which I-frames are inserted into the bitstream allows for the inclusion of more instances of lower bitrate frames (e.g., P-frames and / or B-frames) in the encoded video bitstream. Reducing the number of I-frames offers various benefits. For example, reducing the number of I-frames allows the system to operate at a lower overall average bitrate based on the lower bitrates of other frame types (e.g., P-frames and / or B-frames).
[0147] Another technique that can be implemented using feedback information is to dynamically insert I-slices with intra-frame refresh periods into the bitstream. This technique assumes that the encoder can enforce I-frames and intra-frame refresh periods. Intra-frame refresh periods spread the intra-frame decoded blocks of the I-frame across several frames. In one example of an I-frame with four I-slices, the I-slices of the I-frame could be included in each of four consecutive frames, where the other slices of the four consecutive frames include P-frames (or B-frames in some cases). In some cases, multiple I-slices can be included in a frame with an intra-frame refresh period (or across multiple frames). Intra-frame refresh periods can help prevent frame size spikes when inserting full I-frames into the bitstream. Frame size spikes can increase latency, cause jitter, and can lead to other problems, especially for XR-related applications. Preventing frame size spikes can be advantageous for XR-related applications (e.g., providing immersive media consumption), multi-user applications, and other applications that typically require higher bandwidth and / or are more latency-sensitive compared to other types of media. For example, some XR-related applications consume a lot of data and are sensitive to latency. In such cases, reducing the number of I-frames or I-slices can help reduce latency and bandwidth consumption.
[0148] Figure 5 This diagram illustrates an example of a video decoding structure with intra-frame refresh cycles in an error-free link (where no lost or corrupted frames are detected). As shown, the first frame 502 in the bitstream is the original I-frame decoded into the encoded video bitstream. Intra-frame refresh cycles (including intra-frame refresh cycles 504 and 506) can be periodically inserted into the bitstream. Intra-frame refresh cycle 504 distributes I-slices across four frames, including first frame 504a, second frame 504b, third frame 504c, and fourth frame 504d. The first slice (the topmost slice) of first frame 504a includes an I-slice, while the second, third, and fourth slices of first frame 504a include P or B slices. The second slice of second frame 504b (directly below the first slice) includes an I-slice, while the first, third, and fourth slices of second frame 504b include P or B slices. The third slice of frame 504c (directly below the second slice) includes an I slice, while the first, second, and fourth slices of frame 504c include either a P or B slice. The fourth slice of frame 504d (the bottommost slice) includes an I slice, while the first, second, and third slices of frame 504d include either a P or B slice.
[0149] Using feedback information, the periodic insertion of intra-frame refresh cycles can be relaxed (or even eliminated in some cases). For example, if each frame is divided into N slices, intra-frame refresh cycles can be inserted at longer intervals (e.g., M frames), where each cycle covers N frames and M >> N. In one example of feedback-based intra-frame refresh insertion, the encoding device (e.g., at a server in a split-rendering XR system) can identify slices with missing or corrupted groups based on feedback information received from client devices. In some cases, the encoding device can generate a mask for the missing slices. The mask for the missing slices can include the location of the missing block in the image. For example, the mask can include a binary value per pixel, where the binary value is true (e.g., value 1) for pixels in the missing slice and false (e.g., value 0) for pixels not in the missing slice.
[0150] The server can then force an intra-frame refresh period (including the period of the intra-frame refresh slice) and, in some cases, insert a complete forced I-frame into the next available frame or multiple available frames of the bitstream. In some cases, for a forced intra-frame refresh period, the encoding device can generate a slice with a slice size larger than the original slice size of the encoded video bitstream. Generating a slice larger than the original slice in the bitstream ensures that any errors propagating to other parts of the frame are compensated for as quickly as possible. For example, the slice size of the slice at position N in a frame with an intra-frame refresh period can be larger than the original slice at position N in a missing or corrupted frame, which ensures that the complete missing slice and any possible propagation motion are covered by the intra-frame decoded block of the slice in the intra-frame refresh period. The decision on how many slices to divide the frame into can be a per-frame decision. In the case of a forced intra-frame refresh period, the encoding device can determine how long the intra-frame refresh period will be. For example, refer to Figure 7A (Described in more detail below), the original frame 702a has six slices. The encoding device can choose to insert intra-frame refresh cycles on a certain number of frames (e.g., based on the location of missing or corrupted slices, based on the delay between detecting a missing or corrupted slice and inserting an intra-frame refresh cycle or an I-frame slice, and other factors). For example, based on the decision to insert intra-frame refresh cycles on three frames, the encoding device can divide the next three frames (including frames 706a, 708a, and 710a) into three slices each, and make one slice in each frame an I-slice, as shown below. Figure 7A As shown.
[0151] As described above, by generating a slice larger than the original slice in the bitstream, any errors propagating to other parts of the frame can be compensated for as quickly as possible. Propagated errors can include those that propagate to subsequent frames processed by the decoding device after a frame with missing or corrupted information. In some cases, errors can propagate because subsequent frames to be processed by the decoding device (when generating new forced I-slices and / or I-frames) are P-frames and / or B-frames that rely on inter-frame prediction. For example, in such a case, errors can propagate to future P-frames or B-frames because the decoding device uses a frame with a missing slice to predict one or more subsequent frames.
[0152] As described above, the length of the forced intra-refresh period (e.g., the number of frames), the number of slices in each intra-refresh frame, the size of the slices in the forced intra-refresh period, and / or whether a complete forced I-frame is inserted can be determined based on various factors. Examples of such factors include the maximum motion search range (also known as the motion search range), the number of slices in a frame that includes video slices with missing packets, the location of the missing or corrupted video slices in the video frame, the time at which a forced intra-refresh period or I-frame can be inserted into the updated video bitstream based on feedback information, the maximum acceptable amount of time that error hiding can be performed by the client device, any combination thereof, and / or other appropriate factors. One or more of these factors can be used to determine the length of the forced intra-refresh period, the size of the slices in the forced intra-refresh period, and whether a complete forced I-frame is inserted. The time for inserting a forced intra-refresh period or I-frame can be based on the delay between detecting a missing or corrupted slice and inserting the slice of the intra-refresh period or I-frame (e.g., in part based on how quickly the encoding device can react and insert the intra-refresh period or I-frame). As mentioned above, the delay can be based on the amount of time spent detecting lost or corrupted packets, the amount of time spent sending feedback information to the encoding device, and the amount of time spent by the encoding device performing intra-frame prediction and other encoding processes for video frame 406.
[0153] Decisions based on one or more of these factors can ensure that missing slices, excluding errors propagating from intra-frame decoded blocks (I-frames, I-slices, or other video data) are covered within the maximum acceptable amount of time that the client device can perform error hiding, except for errors due to movement across frames. Larger intra-frame refresh cycles can avoid instantaneous quality degradation but may require the client device to maintain error hiding for more frames (e.g., ATW error hiding). If the encoding device determines the length of the intra-frame refresh cycle, or inserts complete I-frames, it can remain within the maximum acceptable amount of time that the client device can perform error hiding. Such a solution can be useful in latency-sensitive systems (e.g., systems delivering real-time content, such as in some split-rendering systems, real-time video streaming, and / or broadcasting).
[0154] In an illustrative example, assuming a missing or corrupted slice spans from line X to line Y in the original frame, and the maximum motion search range is d, intra-frame decoded blocks (e.g., I-slices or other groups of intra-frame decoded blocks) can be added, covering Xd to Y+d one frame later, X-2d to Y+2d two frames later, and so on. The maximum motion search range (also known as the motion search extent) can be the maximum distance that motion estimation (inter-frame prediction) can use to search for similar blocks relative to the current block. Those skilled in the art will recognize that intra-frame decoded blocks can be added to the bitstream based on other multiples of the motion search range (e.g., X-2d to Y+2d one frame later, X-3d to Y+3d two frames later, or other multiples). In some cases, the maximum motion search range can be a parameter in the configuration of the video encoder. In some cases, the maximum motion search range can be set as a default value and / or can be provided as user input. In one example, depending on the factors described above, a larger slice at position N in the frame can be generated compared to the original slice at position N in the missing or corrupted frame. In another example, a complete forced I-frame can be inserted into the encoded video bitstream based on various factors. The following section discusses... Figure 6A – Figure 6D , Figure 7A – Figure 7F and Figure 8A – Figure 8H An example is described.
[0155] Figures 6A-6D This is a diagram illustrating an example of a video decoding structure using dynamic I-slices. Figures 6A-6D The configuration shown is for frames with a raw slice structure comprising four slices (with missing groups). Figure 6A As shown, the original frame 602a (which is a frame included in the encoded video bitstream) includes a missing or corrupted slice 604a. The client device can detect the missing or corrupted slice 604a and can send feedback information to the encoding device indicating that frame 602a includes a missing or corrupted slice. The encoding device can then begin inserting a forced intra-frame refresh period or an I-frame into the next available frame 606a. Frame 606a can be the next frame in the encoded video bitstream immediately following frame 602a, or it can be multiple frames after frame 602a (based on the delay required to receive the feedback information and generate the intra-frame refresh period).
[0156] As described above, various factors can be considered when determining the number of frames in a forced intra-refresh cycle, the number of I-slices in a forced intra-refresh cycle, the size of the I-slices in a forced intra-refresh cycle, and / or whether to insert a complete forced I-frame. In some examples, factors that can be considered include the maximum motion search range, the number of slices in a frame that includes lost or corrupted video slices, the location of the lost or corrupted video slices in the video frame, the delay between detecting a lost or corrupted frame and inserting a slice in the intra-refresh cycle or I-frame, any combination thereof (including one or more of the factors), and / or other appropriate factors. Figure 6A In this context, the missing or corrupted slice 604a is the first slice (the topmost slice) in the original frame 602a. Based on the fact that slice 604a is the topmost slice in the original frame 602a and that there are four slices in the original frame 602a, the encoding device can insert an intra-frame refresh period on two frames (including the first intra-frame refresh frame 606a and the second intra-frame refresh frame 608a). The intra-frame refresh period includes two I slices, namely slice 605a in the first intra-frame refresh frame 606a and slice 607a in the second intra-frame refresh frame 608a. Based on the determination to include an intra-frame refresh period on both frames, each of the two frames 606a and 608a includes two slices, each slice including an I slice (slice 605a and slice 607a) and a P slice or a B slice.
[0157] To compensate for error propagation caused by a lost or corrupted slice 604a, slices (slices 605a and 607a) for the intra-frame refresh period are generated such that they are larger than the original slice including one or more lost packets. For example, the encoding device may generate slice 605a for insertion into the first intra-frame refresh frame 606a such that slice 605a is sized to include multiple blocks (e.g., CTU, CTB, or other blocks) of the first intra-frame refresh frame 606a. In one example, slice 605a may include blocks from the upper half of the first intra-frame refresh frame 606a. Another slice in frame 606a (e.g., blocks from the lower half) may include a P-slice or a B-slice.
[0158] The encoding device can also generate slice 607a for insertion into the second intra-refresh frame 608a, wherein the size of slice 607a is defined to include the remaining blocks (e.g., macroblocks, CTUs, CTBs, or other blocks) of the second intra-refresh frame 608a not covered by slice 605a. Continuing the example above, where slice 605a includes blocks in the upper half of the first intra-refresh frame 606a, slice 607a may include blocks in the lower half of the second intra-refresh frame 608a. Another slice in frame 608a (e.g., blocks in the upper half) may include a P-slice or a B-slice. Frames following frame 608a (including frames 610a and 612a) may include P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-refresh period is inserted into the bitstream based on feedback from the client device.
[0159] In an illustrative example, frame 602a can have a resolution of 1440x1440 (in terms of pixel width x pixel height), making frame 602a have 1440 rows of pixels. In such an example, a missing or corrupted slice 604a can span from the first row (the topmost row) (X=1) of frame 602a to the 360th row (Y=360), and the maximum motion search range can be equal to 32 (d=32). As described above, for a missing or corrupted slice spanning from row X to row Y with a maximum motion search range of d, an intra-frame decoding block can be added that covers Xd to Y+d after one frame (e.g., the encoder needs one frame after receiving feedback to begin inserting an intra-frame refresh cycle, in which case there is no frame between the missing frame and the first frame of the intra-frame refresh cycle), X-2d to Y+2d after two frames (e.g., the encoder needs two frames after receiving feedback to begin inserting an intra-frame refresh cycle, in which case there is one frame between the missing frame and the first frame of the intra-frame refresh cycle), and so on. Using this example, and assuming the encoding device can insert a first intra-frame refresh frame 606a one frame later, the encoding device can generate slice 606a to include an intra-frame decoded block spanning from the first row to row 392 (Y+d = 360+32 = 392). Note that since the first row is the top row of frame 602a, the distance d is not subtracted from the first row (X = 1). If the encoding device inserts the first intra-frame refresh frame 606a two frames later, the encoding device can generate slice 606a to include an intra-frame decoded block spanning from the first row to row 424 (Y+2d = 360+64 = 424).
[0160] When a client device receives an additional inter-frame decoded frame (e.g., a P-frame or B-frame) before receiving an error-free I-frame or I-slice, the client device can perform error hiding on the inter-frame decoded frame until it receives the error-free I-frame or I-slice. Once the intra-frame decoded block of the I-frame or I-slice covers the missing slice, the client can stop performing error hiding. For example, in... Figure 6A In this process, the client device can stop error hiding after refreshing frame 606a in the first frame because it receives error-free I slice 605a in the refreshed frame 606a in the first frame.
[0161] exist Figure 6B In the original frame 602b, there is a missing or corrupted slice 604b. The client device can detect the missing or corrupted slice 604b and can send feedback information to the encoding device, so that the encoding device knows that frame 602b contains a missing or corrupted slice. The encoding device can begin inserting a forced intra-frame refresh period or I-frame into the next available frame 606b, which can be the next frame immediately following frame 602b in the encoded video bitstream or multiple frames following frame 602b (based on the delay required to receive the feedback information and generate the intra-frame refresh period). The missing or corrupted slice 604b is the second slice in the original frame 602b (the slice immediately below the topmost slice). Because there are four slices in the original frame 602b, and because slice 604b is not the topmost or bottommost slice in the original frame 602b, the error caused by the missing or corrupted slice 604b can propagate to the first and / or third slices (from the top of the frame) of one or more subsequent frames (including frame 606b). Because errors can propagate to the first and / or third slices, and therefore to the area covered by approximately three-quarters of the frame, it is not possible to do so as in... Figure 6A As in the example, intra-frame refresh cycles are inserted on both halves of the I-block. For example, an intra-frame refresh is a full cycle controlled by the frame's regular slicing scheme, which constraints can enforce the position of dynamic I-blocks. In this case, the encoding device can force frame 606b to be a complete I-frame, ensuring that propagation errors are accounted for. Frames following frame 606b (including frames 608b and 610b) can include P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh cycle is inserted into the bitstream based on feedback from the client device. Figure 6B In the example, if needed, the client device can perform error hiding until I-frame 606b is received, and can stop error hiding after frame 606b.
[0162] In an illustrative example, frame 602b can have a resolution of 1440x1440 (with 1440 rows of pixels), a missing or corrupted slice 604b can span from row 361 (X=361) to row 720 (Y=720) of frame 602b, and the maximum motion search range can be equal to 32 (d=32). Because errors can propagate to the first and / or third slices, as described above, the encoding device can generate frame 606b as a complete I-frame instead of adding an I-slice spanning from row 329 (Xd=361-32) to row 752 (Y+d=720+32) in frame 606b (assuming the encoding device can insert frame 606b after one frame), or for other sizes if the encoding device needs more time to insert frame 606b. For example, as described above, intra-frame refresh is a full cycle controlled by the frame's regular slicing scheme, which constraints can enforce the position of dynamic I-blocks. Due to this constraint, the encoding device cannot insert a slice that covers three-quarters of frame 602b.
[0163] exist Figure 6C In the original frame 602c, a missing or corrupted slice 604c can be detected by the client device. The client device can send feedback to the encoding device to indicate that frame 602c includes a missing or corrupted slice. The encoding device can begin inserting a forced intra-frame refresh period or I-frame into the next available frame 606c, which can be the next frame immediately following frame 602c in the encoded video bitstream or multiple frames following frame 602c. The missing or corrupted slice 604c is the third slice in the original frame 602c (the slice immediately above the bottommost slice). Because there are four slices in the original frame 602c, and because slice 604c is neither the topmost nor the bottommost slice in the original frame 602c, errors from the missing or corrupted slice 604c can propagate to the second and / or fourth slices (from the top of the frame) of one or more subsequent frames (including frame 606c). The encoding device can force frame 606c to be a complete I-frame to ensure that the propagated errors are taken into account. Frames following frame 606c (including frames 608c and 610bc) may include P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh period is inserted into the bitstream based on feedback from the client device. If necessary, the client device may perform error hiding until I-frame 606c is received, and may stop error hiding after frame 606c.
[0164] In an illustrative example, frame 602c can have a resolution of 1440x1440 (with 1440 rows of pixels), a missing or corrupted slice 604c can span from row 721 (X=721) to row 1080 (Y=1080) of frame 602c, and the maximum motion search range can be equal to 32 (d=32). Similar to the above regarding... Figure 6B For example, the encoding device could generate frame 606c as a complete I-frame instead of adding an I-slice from line 689 (Xd = 721-32) to 1112 (Y+d = 1080+32) in frame 606c (assuming the encoding device can insert frame 606c after a frame), or for a different size if the encoding device needs more time to insert frame 606b.
[0165] exist Figure 6D In the original frame 602d, there is a missing or corrupted slice 604d. The client device can detect the missing or corrupted slice 604d and can send feedback information to the encoding device, so that the encoding device knows that frame 602a includes a missing or corrupted slice. The encoding device can begin inserting a forced intra-frame refresh period or an I-frame into the next available frame 606d. Frame 606d can be the next frame immediately following frame 602d, or it can be one of several frames following frame 602d based on the delay required for receiving feedback information and generating an intra-frame refresh period.
[0166] generate Figure 6D The intra-frame refresh cycle slices (slices 605d and 607d) are made larger than the original slice in frame 602d to compensate for propagated errors (due to motion errors across frames). Figure 6D In this example, the missing or corrupted slice 604d is the fourth slice (the bottommost slice) in the original frame 602d. Because slice 604d is the bottommost slice in the original frame 602d, the encoding device can insert intra-frame refresh cycles on two frames (including the first intra-frame refresh frame 606d and the second intra-frame refresh frame 608d). In this case, each of the two frames 606d and 608d includes two slices, namely an I slice (slice 605d and slice 607d) and a P slice or a B slice. To account for propagation errors, the encoding device can generate slice 605d for the first intra-frame refresh frame 606d, which has the size of multiple blocks (e.g., macroblocks, CTUs, CTBs, or other blocks) that include the first intra-frame refresh frame 606d. In one example, slice 605d may include blocks from the upper half of the first intra-frame refresh frame 606d. Another slice in frame 606d (e.g., blocks from the lower half) may include a P slice or a B slice.
[0167] The encoding device can generate slice 607d for insertion into the second intra-refresh frame 608d. The size of slice 607d can be defined as including the remaining blocks (e.g., macroblocks, CTUs, CTBs, or other blocks) of the second intra-refresh frame 608d not covered by slice 605d. Continuing the example above, where slice 605d includes blocks from the upper half of the first intra-refresh frame 606d, slice 607d can include blocks from the lower half of the second intra-refresh frame 608d. Another slice in frame 608d (e.g., blocks from the upper half) can include a P-slice or a B-slice. Frames following frame 608d (including frames 610d and 612d) can include P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-refresh period is inserted into the bitstream based on feedback from the client device. If necessary, the client device can perform error hiding until it receives I-slice 607d, and can stop error hiding after frame 608d, which is the point where the intra-frame decoded block covers the missing slice 604d.
[0168] In an illustrative example, frame 602d may have a resolution of 1440x1440. In such an example, a missing or corrupted slice 604d can span from line 1081 (X = 1081) of frame 602a to the bottom line 1440 (Y = 1440), and the maximum motion search range can be equal to 32 (d = 32). In one example, assuming the encoding device can insert an intra-frame refresh of frame 606d after a frame, the encoding device can generate slice 606d to include an intra-frame decoded block spanning from line 1049 (Xd = 1081 - 32 = 1049) to 1440. Note that the distance d is not added to the last line of frame 602d (Y = 1440). If the encoding device inserts an intra-frame refresh frame 606d after two frames, the encoding device can generate slice 606d to include an intra-frame decoded block spanning from line 1049 (Xd = 1081 - 64 = 1017) to 1440.
[0169] Figures 7A-7F This is a diagram illustrating an additional example of a video decoding structure using dynamic I-slices. Figures 7A-7F The configuration shown is for a lost or corrupted frame that comprises a raw slice structure of six slices. Figure 7AIn the encoded video bitstream, the original frame 702a includes a missing or corrupted slice 704a. The client device can detect the missing or corrupted slice 704a and can send feedback to the encoding device indicating that frame 702a includes a missing or corrupted slice. The encoding device can then begin inserting a forced intra-frame refresh period or I-frame into the next available frame 706a, which can be the next frame in the encoded video bitstream immediately following frame 702a, or it can be multiple frames following frame 702a (based on the delay required to receive the feedback and generate the intra-frame refresh period).
[0170] Since the lost or damaged slice 704a is the first slice (the topmost slice) in the original frame 702a, and based on the existence of six slices in the original frame 702a, the encoding device can insert an intra-refresh period on three frames (including the first intra-refresh frame 706a, the second intra-refresh frame 708a, and the third intra-refresh frame 710a). Based on the determination to include an intra-refresh period on three frames, each of the three frames 706a, 708a, and 710a includes three slices, including I slices (slices 705a, 707a, and 709a) and P or B slices. For example, an intra-refresh period includes three slices, including slices 705a, 707a, and 709a. To account for error propagation caused by the lost or damaged slice 704a, the slices of the intra-refresh period (slices 705a, 707a, and 709a) are larger than the original slices including one or more missing groups, so as to cover the propagated motion. For example, the encoding device may generate a slice 705a for insertion into the first intraframe refresh frame 706a, such that the size of slice 705a includes a first number of blocks (e.g., macroblocks, CTUs, CTBs, or other blocks) of the first intraframe refresh frame 706a, such as the first third of the blocks of the first intraframe refresh frame 706a. The remaining two slices in frame 706a (e.g., the bottom two-thirds of the blocks of frame 706a) may include P slices or B slices.
[0171] The encoding device can generate slice 707a for insertion into a second intra-frame refresh frame 708a, the size of which is defined to include a second number of blocks (e.g., macroblocks, CTUs, CTBs, or other blocks) of the second intra-frame refresh frame 708a, such as the middle third of the second intra-frame refresh frame 708a. The remaining two slices in the second intra-frame refresh frame 708a (e.g., the top third and bottom third blocks) can include P-slices or B-slices. The encoding device can also generate slice 709a for insertion into a third intra-frame refresh frame 710a, wherein the size of slice 709a is defined to include the remaining blocks (e.g., macroblocks, CTUs, CTBs, or other blocks) of the third intra-frame refresh frame 710a not covered by slice 705a and slice 707a. Continuing the example above, where slice 705a includes the top third of the first intra-frame refresh frame 706a, and slice 707a includes the middle third of the second intra-frame refresh frame 708a, slice 709a can include the bottom third of the third intra-frame refresh frame 710a. The remaining two slices in the third intra-frame refresh frame 710a (e.g., the top two-thirds block) may include either a P-slice or a B-slice. Frames following 710a (including frames 712a and 714a) may include either P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh period is inserted into the bitstream based on feedback from the client device. Figure 7A In the example, if needed, the client device can perform error hiding until I slice 705a is received, and can stop error hiding after frame 706a.
[0172] In an illustrative example, frame 702a can have a resolution of 1440x1440 (with 1440 rows of pixels). A missing or corrupted slice 704a can span from the first row (topmost row) (X=1) of frame 702a to row 240 (Y=240), and the maximum motion search range can be equal to 32 (d=32). Assuming the encoding device can insert a first intra-frame refresh frame 706a one frame later, the encoding device can generate slice 706a to include an intra-frame decoded block spanning from the first row to row 272 (Y+d=240+32). Because the first row is the top row of frame 702a, the distance d is not subtracted from the first row (X=1). If the encoding device inserts a first intra-frame refresh frame 706a two frames later, the encoding device can generate slice 706a to include an intra-frame decoded block spanning from the first row to row 304 (Y+2d=240+64).
[0173] Reference Figure 7BThe original frame 702b includes a missing or corrupted slice 704b. Based on feedback received from the client device indicating that frame 702b includes a missing or corrupted slice, the encoding device can begin inserting a forced intra-frame refresh period or I-frame (e.g., the next frame immediately following frame 702b or multiple frames following frame 702b) in the next available frame 706b. Since the missing or corrupted slice 704b is the second slice in the original frame 702b (the slice immediately below the topmost slice), and based on the presence of six slices in the original frame 702b, the encoding device can insert intra-frame refresh periods on two frames (including the first intra-frame refresh frame 706b (with slice 705b) and the second intra-frame refresh frame 708b (with slice 707b)). Based on the intra-frame refresh period spanning two frames, the first intra-frame refresh frame 706b and the second intra-frame refresh frame 708b each include two slices, including an I-slice (slice 705b and slice 707b) and a P-slice or a B-slice.
[0174] To account for error propagation caused by a lost or corrupted slice 704b, the encoding device may generate slice 705b such that it includes a first number of blocks from the first intra-frame refresh frame 706b (e.g., blocks in the upper half of the first intra-frame refresh frame 706b). Another slice in frame 706b (e.g., blocks in the lower half of frame 706b) may include a P-slice or a B-slice. The encoding device may generate slice 707b for insertion into the second intra-frame refresh frame 708b, the size of which is defined as including a second number of blocks from the second intra-frame refresh frame 708b (e.g., blocks in the lower half of the second intra-frame refresh frame 708b). Another slice in the second intra-frame refresh frame 708b (e.g., blocks in the upper half) may include a P-slice or a B-slice. Frames following frame 708b (including frames 710b and 712b) may include P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh period is inserted into the bitstream based on feedback from the client device. Figure 7B In the example, if needed, the client device can perform error hiding until I slice 705b is received, and can stop error hiding after frame 706b.
[0175] In an illustrative example, frame 702b may have a resolution of 1440x1440. A lost or corrupted slice 704b may span from line 241 (X = 241) to line 480 (Y = 480) of frame 702b, and the maximum motion search range may be equal to 32 (d = 32). Assuming the encoding device can insert an intra-frame refresh frame 706b after a frame, the encoding device can generate slice 706b to include an intra-frame decoded block spanning from line 209 (Xd = 241 - 32) to line 512 (Y + d = 480 + 32).
[0176] exist Figure 7C In the original frame 702c, there is a missing or corrupted slice 704c. Based on the received feedback indicating that frame 702c includes a missing or corrupted slice, the encoding device can begin inserting a forced intra-frame refresh period or I-frame into the next available frame 706c. Because there are six slices in the original frame 702c, and because slice 704c is not one of the top two or bottom two slices of the original frame 702c, errors from the missing or corrupted slice 704c can propagate to the entirety of one or more subsequent frames (including frame 706c). In such a case, the encoding device can force frame 706c to be a complete I-frame to ensure that the propagated errors are taken into account. Frames after frame 706c (including frames 708c and 710c) can include P-frames or B-frames until the client device detects another missing or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh period is inserted into the bitstream based on feedback from the client device. If necessary, the client device can perform error hiding until I-frame 706c is received, and can stop error hiding after frame 706c.
[0177] In an illustrative example, frame 702c can have a resolution of 1440x1440, a missing or corrupted slice 704c can span from line 481 (X=481) to line 720 (Y=720) of frame 702c, and the maximum motion search range can be equal to 32 (d=32). The encoding device can generate frame 706c as a complete I-frame instead of adding an I-slice spanning from line 449 (Xd=481-32) to 752 (Y+d=720+32) in frame 706c (assuming the encoding device can insert frame 706c after one frame), or for a different size if inserting frame 706c takes longer.
[0178] exist Figure 7DIn the original frame 702d, there is a missing or corrupted slice 704d. Based on feedback received from the client device indicating that frame 702d includes a missing or corrupted slice, the encoding device can begin inserting a forced intra-frame refresh period or I-frame into the next available frame 706d. Because there are six slices in the original frame 702d, and slice 704d is not one of the top two or bottom two slices of the original frame 702d, the encoding device can force frame 706d to be a complete I-frame to ensure that errors that could propagate to one or more subsequent frames are taken into account. Frames after frame 706d (including frames 708d and 710d) can include P-frames or B-frames until the client device detects another missing or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh period is inserted into the bitstream based on feedback from the client device. If necessary, the client device can perform error hiding until I-frame 706d is received, and can stop error hiding after frame 706d.
[0179] In an illustrative example, frame 702d can have a resolution of 1440x1440 (with 1440 rows of pixels), a missing or corrupted slice 704d can span from row 721 (X=721) to row 960 (Y=960) of frame 702d, and the maximum motion search range can be equal to 32 (d=32). (Regarding...) Figure 7C The example discussed is similar; the encoding device can generate frame 706d as a complete I-frame, instead of adding I-slices that span a subset of the lines in frame 706d.
[0180] Reference Figure 7E The original frame 702e includes a missing or corrupted slice 704e. Based on the received feedback indicating that frame 702e includes a missing or corrupted slice, the encoding device can begin inserting a forced intra-frame refresh period or I-frame (e.g., the next frame immediately following frame 702e or multiple frames following frame 702e) into the next available frame 706e. Since the missing or corrupted slice 704e is the fifth slice in the original frame 702e (the slice immediately above the bottom slice), and based on the existence of six slices in the original frame 702e, the encoding device can insert intra-frame refresh periods on two frames (including the first intra-frame refresh frame 706e (with slice 705e) and the second intra-frame refresh frame 708e (with slice 707e)).
[0181] To account for error propagation caused by the lost or corrupted slice 704b, the encoding device may generate slice 705e such that it includes a first number of blocks from the first intra-frame refresh frame 706e (e.g., blocks in the upper half of the first intra-frame refresh frame 706e). Another slice in frame 706e (e.g., blocks in the lower half of frame 706e) may include a P-slice or a B-slice. The encoding device may generate slice 707e having a size defined as including a second number of blocks from the second intra-frame refresh frame 708e (e.g., blocks in the lower half of the second intra-frame refresh frame 708e). Another slice in the second intra-frame refresh frame 708e (e.g., blocks in the upper half) may include a P-slice or a B-slice. Frames following frame 708e (including frames 710e and 712e) may include P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh period is inserted into the bitstream based on feedback from the client device. If necessary, the client device can perform error hiding until I slice 705e is received, and can stop error hiding after frame 706e.
[0182] In an illustrative example, frame 702e can have a resolution of 1440x1440. A lost or corrupted slice 704e can span from line 961 (X = 961) to line 1200 (Y = 1200) of frame 702e, and the maximum motion search range can be equal to 32 (d = 32). Assuming the encoding device can insert an intra-frame refresh frame 706e after a frame, the encoding device can generate slice 706e to include an intra-frame decoded block spanning from line 929 (Xd = 961 - 32) to line 1232 (Y + d = 1200 + 32).
[0183] exist Figure 7FIn the original frame 702f, there is a missing or corrupted slice 704f. Based on the received feedback indicating that frame 702f contains a missing or corrupted slice, the encoding device can begin inserting a forced intra-frame refresh period or I-frame in the next available frame 706f. Since the missing or corrupted slice 704f is the sixth slice (the bottommost slice) in the original frame 702f, and based on the existence of six slices in the original frame 702f, the encoding device can insert intra-frame refresh periods on three frames (including the first intra-frame refresh frame 706f (including slice 705f), the second intra-frame refresh frame 708f (including slice 707f), and the third intra-frame refresh frame 710f (including slice 709f)). Since the intra-frame refresh period spans three frames, frames 706f, 708f, and 710f each have three slices, including I-slices (slices 705f, 707f, and 709f) and P-slices or B-slices. To account for error propagation caused by a missing or corrupted slice 704f, the encoding device may generate a slice 705f for insertion into the first intraframe refresh frame 706f, such that the size of slice 705f includes a first number of blocks of the first intraframe refresh frame 706f (e.g., the top third of the blocks of the first intraframe refresh frame 706f). The remaining two slices in frame 706f (e.g., the bottom two-thirds of the blocks of frame 706f) may include either a P slice or a B slice.
[0184] The encoding device can generate a slice 707f to be inserted into the second intra-frame refresh frame 708a, the size of which is defined to include a second number of blocks of the second intra-frame refresh frame 708f (e.g., the middle third of the blocks of the second intra-frame refresh frame 708F). The remaining two slices in the second intra-frame refresh frame 708f (e.g., the top third and top third blocks) can include P-slices or B-slices. The encoding device can also generate a slice 709f to be inserted into the third intra-frame refresh frame 710f, wherein the size of slice 709f is defined to include the remaining blocks in the third intra-frame refresh frame 710f not covered by slice 705f and slice 707f. Continuing the example above, slice 709f can include the bottom third of the blocks of the third intra-frame refresh frame 710f. The other two slices in the third intra-frame refresh frame 710f (e.g., the top two-thirds blocks) can include P-slices or B-slices. Frames following frame 710f (including frames 712f and 714f) may include P-frames or B-frames until the client device detects another lost or corrupted frame or slice and an I-frame, I-slice, or intra-frame refresh cycle is inserted into the bitstream based on feedback from the client device. If necessary, the client device may perform error hiding until I-slice 709f is received, which occurs when an intra-frame decoded block overwrites the missing slice 704f. The client device may stop error hiding after frame 706f.
[0185] In an illustrative example, frame 702f can have a resolution of 1440x1440 (with 1440 rows of pixels). A missing or corrupted slice 704f can span from the first row (topmost row) (X=1) of frame 702a to row 240 (Y=240), and the maximum motion search range can be equal to 32 (d=32). Assuming the encoding device can insert an intra-frame refresh of frame 706f after a previous frame, the encoding device can generate slice 706f to include an intra-frame decoded block spanning from the first row to row 272 (Y+d=240+32). Distance d is not added to the last row of frame 702f (Y=1440).
[0186] Figures 8A-8H This is a diagram illustrating an additional example of a video decoding structure using dynamic I-slices. Figures 8A-8H The configuration shown is for a lost or corrupted frame, which consists of an original slice structure of eight slices. Figures 8A-8H The frames in the dataset can have a resolution of 1440x1440 (with 1440 rows of pixels), and the maximum motion search range can be equal to 32 (d = 32). (This is related to the previous text about...) Figures 6A-6D and Figures 7A-7F The example described above is similar; one or more of the factors mentioned above can be used to determine the number of frames in a forced intra-refresh cycle, the number of slices in each intra-refresh frame, the size of the I-slice in the intra-refresh frame of the forced intra-refresh cycle, and whether a complete forced I-frame is inserted. For example, in Figure 8A In this process, the encoding device can insert intra-frame refresh cycles (including the first intra-frame refresh frame 806a and the second intra-frame refresh frame 808a) on four frames based on the fact that the lost or damaged slice is the first slice (the topmost slice) in the original frame 802a and that there are eight slices in the original frame. Figure 8A (The third and fourth intra-frame refresh frames are not shown). In such an example, when the intra-frame refresh cycle spans four frames, intra-frame refresh frames 806a and 808a will each include four slices, including I slices (e.g., slices 805a and 807a) and P or B slices. In an illustrative example, in Figure 8AIn this context, a lost or corrupted slice can span from the first row (X=1) to row 180 (Y=180) in the original frame 802a. Using the formulas above, such as Xd to Y+d (after one frame) and X-2d to Y+2d (after two frames), the lost or corrupted slice covers rows 1 to 180 in the original frame 802a, and the error can propagate to row 212 (180+32=212). An I-slice 805a can be generated such that its intra-frame block spans from row 1 to row 360 in the first frame refreshed frame 806a, ensuring the elimination of errors (from row 1 to row 212). An I-slice 807a can be generated such that its intra-frame block spans from row 361 to row 720 in the second frame refreshed frame 808a.
[0187] Figure 8H A similar configuration is shown, where the encoding device can insert intra-frame refresh cycles across four frames based on the fact that the lost or corrupted slice is the last slice (the bottommost slice) in the original frame. For example, the lost or corrupted slice can span from line 1261 (X=1261) to line 1440 (Y=1440) in the original frame, and I slices can be generated in four intra-frame refresh frames such that the intra-frame blocks of the I slices each span 360 lines (to cover the entire 1440 lines of the original frame).
[0188] Reference Figure 8B The encoding device can insert intra-frame refresh cycles over three frames based on the fact that the lost or damaged slice 804b is the second slice in the original frame 802b (the slice immediately below the topmost slice) and that there are eight slices in the original frame. Figure 8B Only the first intra-frame refresh frame 806b and the second intra-frame refresh frame 808b are shown. In an illustrative example, in Figure 8B In this frame, a lost or damaged slice 804b can span from line 181 (X=181) to line 360 (Y=360) in the original frame 802b, and can generate an I slice 805b in the first intra-refresh frame 806b, such that the intra-block of I slice 805b spans from line 1 to line 480 in the first intra-refresh frame 806b. An I slice 807b can be generated in the second intra-refresh frame 808b, such that the intra-block of I slice 807b spans from line 481 to line 960 in the second intra-refresh frame 808b. An I slice can be generated in a third intra-refresh frame (not shown), such that the intra-block of the I slice spans from line 961 to line 1440 in the third intra-refresh frame. Figure 8GA similar configuration is shown, where the encoding device can insert intra-frame refresh cycles across three frames based on the fact that the lost or corrupted slice is the seventh slice in the original frame (the slice immediately above the bottommost slice). For example, the lost or corrupted slice can span from line 1081 (X=1081) to line 1260 (Y=1260) in the original frame, and I slices can be generated in three intra-frame refresh frames such that the intra-frame blocks of the I slices each span 480 lines (to cover the full 1440 lines of the original frame).
[0189] exist Figure 8C In this context, the encoding device can insert intra-frame refresh cycles on two frames (including the first intra-frame refresh frame 806c and the second intra-frame refresh frame 808c) based on the fact that the lost or corrupted slice 804c is the third slice in the original frame 802c (the slice immediately below the top two slices) and that there are eight slices in the original frame 802c. In an illustrative example, in Figure 8C In the original frame 802c, a missing or corrupted slice 804c can span from line 361 (X=361) to line 540 (Y=540) in the original frame 802c. An I-slice 805c can be generated in the first intra-refresh frame 806c such that its intra-block spans the upper half of the first intra-refresh frame 806c (from line 1 to line 720). An I-slice 807c can be generated in the second intra-refresh frame 808c such that its intra-block spans the lower half of the second intra-refresh frame 808c (from line 721 to line 1440). Figure 8F A similar configuration is shown, where the encoding device can insert intra-frame refresh cycles on two frames based on the fact that the lost or corrupted slice is the sixth slice in the original frame (the slice immediately above the two bottommost slices). For example, the lost or corrupted slice can span from line 901 (X=901) to line 1080 (Y=1080) in the original frame, and three intra-frame refresh frames can be generated such that the intra-frame blocks of the I slices each span 720 lines (to cover the full 1440 lines of the original frame).
[0190] exist Figure 8D In this process, the encoding device can insert a complete I-frame 806d based on the fact that the lost or corrupted slice 804d is the fourth slice in the original frame 802d (the slice immediately below the top three slices) and that there are eight slices in the original frame. For example, since the lost or corrupted slice 804D is located in the middle of frame 802D, the motion compensation error caused by the lost or corrupted slice 804D can propagate to the entire portion of one or more subsequent frames. In an illustrative example, refer to... Figure 8DA missing or corrupted slice 804d can span from line 541 (X=541) to line 720 (Y=720) in the original frame 802d. An I-frame 806d can be generated such that the intra-frame blocks of I-frame 806d span the entire frame (from line 1 to line 1440). Figure 8E A similar configuration is shown, where a complete I-frame is generated based on the missing or corrupted slice being the fifth slice in the original frame (the slice immediately above the bottom three slices). For example, a missing or corrupted slice can span from line 721 (X=721) to line 900 (Y=900) in the original frame, and an I-frame can be generated such that the intra-frame block of the I-frame spans the entire frame (from line 1 to line 1440).
[0191] As described above, any one or more of the various factors (in any combination) can be used to determine the number of frames in a forced intra-refresh cycle, the number of slices in each intra-refresh frame, the size of the I-slices in the intra-refresh frames of a forced intra-refresh cycle, and / or whether to insert a complete forced I-frame. In one example, if an error occurs in the first slice of a video frame, the intra-refresh cycle can be used to provide a complete I-frame over multiple subsequent intra-refresh cycle frames, regardless of the number of slices in the video frame where corrupted or lost data occurs. In another example, if the lost or corrupted slice is not the first slice of the video frame, a complete forced I-frame can be added to the bitstream. In such an example, the number of slices in the video frame is not a factor determining the characteristics of the intra-decoded data to be inserted into the bitstream.
[0192] Another technique that can be performed in response to received feedback information is to dynamically insert individual I-slices into video frames. For example, if the encoding device allows, it can insert individual I-slices into the error-affected portions of the bitstream. Such a technique assumes that the encoder can force a slice to be an intra-frame decoded slice. As previously mentioned, errors in a given frame may be due to missing or corrupted slices, and in some cases may also include prediction errors propagating due to missing or corrupted slices. The propagating motion depends on the motion search range (where the maximum search range is d, as described above), when the encoding device receives feedback indicating a missing or corrupted slice, and when the encoding device begins inserting the I-slice. Inserting individual I-slices does not change the slice structure of the video frame and can work on any encoding configuration (e.g., strict CBR or periodic intra-frame refresh). In some cases, the encoding device may decide to insert the desired I-slices on multiple frames to avoid introducing a momentary degradation in quality. For example, if the encoding device inserts a complete I-frame into the bitstream, it will introduce bitrate spikes to maintain video quality, which increases the amount of video data transmitted and can cause delays in receiving video data. In another example, the encoding device can insert low-quality I-frames (e.g., with a smaller frame size) to avoid bitrate spikes that degrade video quality. Both of these inefficiencies can be addressed by inserting I-slices across multiple frames. When inserting I-slices across multiple frames, the encoding device can account for the need to cover propagating errors (e.g., by increasing the size of the I-slices). As previously mentioned, the client device can perform error hiding (e.g., ATW error hiding) until missing slices with propagating errors are removed.
[0193] Figure 9A and Figure 9B This is a diagram illustrating an example of a video coding structure using dynamically individual I-slices. Figure 9A In the video bitstream, frame 902a includes a P-slice and / or a B-slice. In frame 906a, the client device detects a slice 904a with missing packets. The client device may send feedback information to the encoding device indicating that frame 906a includes a corrupted slice. Based on the delay between the detection of the corrupted slice 904a and the time when the encoding device can begin inserting a forced I-slice into the bitstream, the encoding device begins inserting an I-slice 909a at frame 912a.
[0194] Because there is a two-frame delay between the detection of the corrupted slice 904a and the insertion of the I-slice 909a by the encoding device, the client device continues to receive frames with P-slices and / or B-slices (including frames 908a and 910a). The prediction error caused by the corrupted slice 904a propagates to subsequent frames, including frames 908a and 910a. Any region of a video frame that can be predicted from the corrupted slice 904a may have corrupted video data (e.g., video decoding artifacts) based on corrupted packet data in the corrupted slice 904a. As shown, frame 908a includes the propagated error 905a, and frame 910a includes the propagated error 907a. For example, the propagated error 905a is caused by the client's decoding device making predictions using the corrupted slice 904a (including lost packets). Similarly, the propagated error 907a is caused by the client's decoding device making predictions using the corrupted slice 904a and / or other corrupted slices (which may be caused by the use of the corrupted slice 904a).
[0195] As described above, the encoding device begins inserting I slice 909a at frame 912a. As... Figure 9A As shown in the example, the encoding device decides to use I-slices within a single frame to clear all propagated errors (instead of spreading the I-slices across multiple frames). This solution prevents errors from propagating further into subsequent frames and limits the number of I-slices required to stop error propagation. While in Figure 9A The diagram shows that I slice 909a includes three I slices, but a single I slice can be inserted in frame 912a (or other frames) to cover prediction errors.
[0196] Reference Figure 9B Frame 902b in the video bitstream includes a P-slice and / or a B-slice. In frame 906b, the client device detects a slice 904b with missing packets. The client device may send feedback information to the encoding device indicating that frame 906b includes a corrupted slice. Based on the delay between the detection of the corrupted slice 904b and the time when the encoding device can begin inserting a forced I-slice into the bitstream, the encoding device begins inserting an I-slice 909b at frame 910b.
[0197] Because there is a frame delay from the detection of the corrupted slice 904b to the insertion of the I slice 907b by the encoding device, the client device receives frame 908b with P slices and / or B slices. The prediction error caused by the lost or corrupted slice 904b propagates to subsequent frames, including frames 908b and 910b. As shown, frame 908b includes the propagated error 905b. Because the I slice 909b consists of only two slices, frame 910b includes the residual propagated error 911b. The propagated error 905b is caused by the client's decoding device making predictions using slice 904b, and the propagated error 907b is caused by making predictions using slice 904b and / or other corrupted slices (caused by the use of slice 904b).
[0198] The encoding device begins inserting I-slice 909b at frame 910b. (As follows) Figure 9B As shown, the encoding device decides to use I-slices on two frames to clear all propagation errors, resulting in residual propagation error 911b. The I-slice 911b in frame 912b comprises two slices to compensate for the residual propagation error 911b. (Compared to...) Figure 9A Compared to the previous example, additional I-slices (such as I-slice 911b in frame 912b, which includes two slices) are required to stop error propagation. Inserting I-slices 909b and 911b across multiple frames can prevent the introduction of transient quality degradation in the encoded video bitstream (e.g., based on bitrate spikes that might be caused by inserting an entire I-frame).
[0199] In some examples, more advanced schemes (e.g., standards-compliant schemes) can modify the slice structure on each frame or a set number of frames. For instance, when an encoding device decides to use I-slices, it can define I-slices of optimal size that cover error propagation based on its knowledge of error propagation, if it has sufficient information. For example, regions of frames affected by error propagation can be defined as slices, and intra-frame blocks can be generated for slices that cover the defined regions. Regions of frames unaffected by error propagation can be encoded as one or more P-slices or B-slices.
[0200] In some examples, the encoding device can analyze the history of motion vectors from past frames. Motion vectors corresponding to regions affected by propagating errors are not used for inter-frame prediction (e.g., regions affected by propagating errors are not used as references for subsequent frames). For example, given an error in a previous frame, the encoding device can buffer motion vectors from the past M frames (e.g., motion vectors from the last 2-3 frames) and possible error propagation from frame n-1, from frame n-2, etc. (where n is the current frame being processed). In one example, motion vectors from errors occurring in frame n-2 can be tracked, and possible locations of errors can be marked in frame n-1. These possible error locations can be avoided (not used as references). In such examples, when selecting motion vectors for the current frame n, reference blocks with propagating errors in frame n-1 are avoided and not used for inter-frame prediction. In some cases, when no match is found in the search region and there is no error, the current block to be decoded in frame n (e.g., macroblock, CTU, CTB, CU, CB, or other blocks) is encoded as an intra-frame decoded block, even within a P-slice. An error-free region can occur when errors from one or more previous frames n-1, n-2, etc., do not propagate to the error-free region of the frame.
[0201] In some examples, the encoding device can implement forward error mitigation techniques. For instance, the future slice structure can be made dynamic based on recent errors. If recent packets have seen a large number of missing frames, the encoding device can reduce the period of inserted I-frames or the size of the intra-frame refresh period. More frequent I-frames require higher bit rates to achieve the same quality, are more difficult to maintain a lower peak-to-average frame size ratio, and provide stronger robustness to packet missing frames. A trade-off based on the history of recent packet missing frames can be achieved.
[0202] An example of the process performed using the dynamic I-frame or I-slice insertion technique described herein will now be described. Figure 13 This is a flowchart illustrating an example of a process 1300 for processing video data. At block 1302, process 1300 includes: a computing device (e.g., Figure 1 The decoding device 112) determines that at least a portion of a video slice of a video frame in the video bitstream is missing or corrupted. In one example, the computing device may parse the packet header of a group of video slices to determine whether any packets in the group are missing or corrupted. The packet header may indicate which slice the packet belongs to, and may also indicate the first and last blocks covered by the packet or slice. In other examples, the computing device may refer to other informational portions of the video data, such as slice headers, parameter sets (e.g., Picture Parameter Set (PPS) or other parameter sets), Supplemental Enhancement Information (SEI) messages, and / or other information.
[0203] At box 1304, process 1300 includes sending feedback information to an encoding device. The feedback information indicates that at least a portion of a video slice is missing or corrupted. For example, the feedback information may include information indicating that a video slice has one or more missing groups. In some examples, the information may include an image of a slice including the missing or corrupted slice and an indication of one or more slices including the missing groups.
[0204] At block 1306, process 1300 includes: receiving an updated video bitstream from an encoding device in response to feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the lost or corrupted video slice. The updated video bitstream may include subsequent portions of the video bitstream, and at least one intra-frame decoded video slice is included in a later frame in the video bitstream than a video frame that includes a slice (e.g., a group) with lost or corrupted information. For example, at least one intra-frame decoded slice in the updated video bitstream may have a higher POC, output order, and / or decoding order compared to the video frames in the video bitstream that include slices (e.g., groups) with lost or corrupted information. In an illustrative example, reference is made to... Figure 6A Frame 602a of the video bitstream includes a lost or corrupted slice 604a, and in response to feedback information, the video bitstream can be updated to include the forced intra-frame decoded slice 605a in the next available frame 606a (e.g., which may be the next frame immediately following frame 602a in the encoded video bitstream, or it may be multiple frames following frame 602a).
[0205] The size of at least one intra-frame decoded video slice is determined to cover (or compensate for) missing or corrupted slices and the propagation error in the video frame caused by the missing or corrupted slices. The propagation error in the video frame may be caused by missing or corrupted slices based on the motion search range. For example, as described above, any region of a video frame that may have corrupted video data can be predicted using data from missing or corrupted slices. In some examples, the missing or corrupted slice spans from the first row to the second row in the video frame, and the size of at least one intra-frame decoded video slice is defined to include at least the first row minus the motion search range to the second row plus the motion search range. In some cases, at least one intra-frame decoded video slice may be larger than the first row minus the motion search range to the second row plus the motion search range. For example, refer to the above regarding... Figure 8AFor example, using formulas such as Xd to Y+d (after one frame), X-2d to Y+2d (after two frames), a missing or corrupted slice can span from row 1 to row 180 in the original frame 802a, and the error can propagate to row 212 (180+32=212). Intra-decoded video slices (e.g., I-slice 805a) can be generated such that intra-blocks of the intra-decoded video slice can span from row 1 to row 360 in an intra-refresh frame (e.g., intra-refresh frame 806a), which ensures the removal of errors (from row 1 to row 212). Those skilled in the art will recognize that other multiples of the motion search range (such as X-2d to Y+2d after one frame, X-3d to Y+3d after two frames, or any other multiple that can be used to cover (e.g., remove) propagated errors) can be added to the bitstream.
[0206] In some examples, in response to determining that at least a portion of a video slice is missing or corrupted, the computing device may perform error concealment (e.g., asynchronous time warp (ATW) error concealment or other types of error concealment) on one or more video frames until an error-free intra-decoded video slice is received in an updated video bitstream. For example, a certain amount of time may be required between the computing device detecting a missing or corrupted slice and the computing device receiving the intra-decoded video slice. The computing device may perform error concealment for any frames received before receiving the intra-decoded video slice.
[0207] In some implementations, at least one intra-decoded video slice includes an intra-decoded frame. For example, in some cases, as described above, at least one intra-decoded slice is a complete forced intra-decoded frame. In some implementations, at least one intra-decoded video slice may be included as part of an intra-refresh cycle. For example, an intra-refresh cycle includes at least one video frame, wherein each of the at least one video frame includes one or more intra-decoded video slices. As described above, the number of at least one video frame in an intra-refresh cycle can be based on various factors, such as the maximum search range described above, the number of slices in a video frame including video slices with missing or corrupted data, the position of the video slices in the video frame, the time at which the intra-refresh cycle is inserted into the updated video bitstream based on feedback information, any combination thereof, and / or other factors. As described above, the time of inserting the intra-refresh cycle can be based on the delay between detecting a missing or corrupted slice and inserting the slice of the intra-refresh cycle or I-frame (e.g., in part based on how quickly the encoding device can react and insert the intra-refresh cycle or I-frame).
[0208] In an illustrative example, when the video slice is the first slice in a video frame, at least one video frame in the intra-frame refresh cycle comprises at least two frames. For example, refer to... Figure 6A As an example, slice 604a may have missing packets. In this case, it can be determined that the intra-frame refresh cycle has two frames, where the first intra-frame refresh frame 606a has a first I-slice 605a, and the second intra-frame refresh frame 608a has a second I-slice 608a. In such an example, the computing device can perform error concealment on the first frame of at least two frames (e.g., frame 606a) but not on the second frame of at least two frames (e.g., frame 608a). The second frame is in the video bitstream after the first frame.
[0209] In another illustrative example, when the video slice is not the first slice in a video frame, at least one video frame in the intra-frame refresh cycle includes an intra-frame decoded frame. For example, refer to... Figure 6B As an example, slice 604b may have missing or corrupted packets; in this case, a complete I-frame 606b can be inserted to cover the propagation errors. In another example, refer to... Figure 8D Slice 804d can have missing or corrupted packets, and a complete I-frame 806d can be inserted to cover propagation errors.
[0210] In another example, when the video slice is the last slice in a video frame, at least one video frame in the intra-frame refresh cycle comprises at least two frames. For example, refer to... Figure 6D As an example, slice 604d may have missing packets, and it can be determined that the intra-frame refresh cycle has two frames, where the first intra-frame refresh frame 606d has a first I-slice 605d, and the second intra-frame refresh frame 608d has a second I-slice 608d. In such an example, the computing device can perform error hiding on the first frame (e.g., frame 606d) and the second frame (e.g., frame 608d) of at least two frames based on the fact that the video slice is the last slice in the video frame. The second frame is in the video bitstream after the first frame.
[0211] Figure 14 This is a flowchart illustrating an example of a process 1400 for processing video data. At block 1402, process 1400 includes: at an encoding device, from a computing device (e.g., ... Figure 1 The decoding device 112) receives feedback information. The feedback information indicates that at least a portion of a video slice of a video frame in the video bitstream is missing or corrupted. For example, the feedback information may include information indicating that a video slice has one or more missing packets. In some examples, the information may include a picture of a slice that includes a missing or corrupted slice and an indication of one or more slices that include missing packets.
[0212] At box 1402, process 1400 includes: generating an updated video bitstream in response to feedback information. The updated video bitstream includes at least one intra-frame decoded video slice having a size larger than the lost or corrupted video slice. (As mentioned above regarding...) Figure 13 As mentioned, the updated video bitstream may include subsequent portions of the video bitstream, and at least one intra-decoded video slice is included in a later frame of the video bitstream than a video frame that includes a slice (e.g., a group) with missing or corrupted information. For example, at least one intra-decoded slice in the updated video bitstream may have a higher POC, output order, and / or decoding order compared to the video frames in the video bitstream that include slices (e.g., groups) with missing or corrupted information.
[0213] The size of at least one intra-frame decoded video slice is determined to cover (or compensate for) lost or corrupted slices and the propagation error in the video frame caused by the lost or corrupted slices. The propagation error in the video frame may be caused by lost or corrupted slices based on the motion search range. For example, as described above, any region of the video frame that can be predicted from lost or corrupted slices may have corrupted video data. In some examples, the lost or corrupted slice spans from the first row to the second row in the video frame. The encoding device may determine the size of at least one intra-frame decoded video slice to include at least the first row minus the motion search range to the second row plus the motion search range. In some cases, at least one intra-frame decoded video slice may be larger than the first row minus the motion search range to the second row plus the motion search range. For example, refer to the above regarding... Figure 8A For example, using formulas such as Xd to Y+d (after one frame), X-2d to Y+2d (after two frames), a lost or corrupted slice can span from row 1 to row 180 in the original frame 802a, and the error can propagate to row 212 (180+32=212). Intra-decoded video slices (e.g., I-slice 805a) can be generated such that intra-blocks of the intra-decoded video slice can span from row 1 to row 360 in an intra-refreshed frame (e.g., intra-refreshed frame 806a), which ensures that the error (from row 1 to row 212) is cleared.
[0214] As described above, the computing device can perform error hiding on one or more video frames in response to at least a portion of a video slice being lost or corrupted, until an error-free intra-decoded video slice is received in an updated video bitstream.
[0215] In some implementations, at least one intra-decoded video slice includes an intra-decoded frame. For example, in some cases, at least one intra-decoded slice is a complete forced intra-decoded frame. In some implementations, at least one intra-decoded video slice may be included as part of an intra-refresh period. For example, an intra-refresh period includes at least one video frame, wherein each of the at least one video frame includes one or more intra-decoded video slices. For example, the encoder may determine to insert I-slices (e.g., as shown in the image) on multiple frames of the intra-refresh period. Figure 6A , Figure 6D , Figure 7A , Figure 7B (as shown in the example). In other examples, the encoder can determine the insertion of a complete I-frame (e.g., as shown in the example). Figure 6B and Figure 6C As shown above, whether to insert a complete I-frame or multiple video frames within an intra-frame refresh period can be based on various factors, such as the maximum search range mentioned above, the number of video slices in a video frame including video slices with missing or corrupted data, the position of the video slices in the video frame, the time to insert the intra-frame refresh period into the updated video bitstream based on feedback information, any combination thereof, and / or other factors.
[0216] In some cases, the encoding device may store updated video bitstreams (e.g., in a decoded picture buffer (DPB), in a storage location used for retrieval for decoding and display, and / or other storage). In other cases, the encoding device may send updated video bitstreams to a computing device.
[0217] In some implementations, as described in more detail below, the encoding device can synchronize with a common reference clock along with other encoding devices. In such an implementation, process 1400 may include adding intra-frame decoded video data to a video bitstream based on a reference clock shared with at least one other encoding device. The reference clock defines a scheduling for interleaving intra-frame decoded video from the encoding device and at least one other encoding device. Process 1400 may include sending a request to adapt the reference clock in response to feedback information, allowing the encoding device to add intra-frame decoded video data to the video bitstream in unscheduled time slots. Process 1400 may include receiving an indication that the reference clock has been updated to define an updated schedule; and adding intra-frame decoded video slices to the video bitstream based on the updated schedule and the updated reference clock. The following section discusses... Figure 10 – Figure 12 and Figure 15 Further details describe encoder synchronization.
[0218] In some examples, processes 1300 and / or 1400 may be powered by a computing device or apparatus (such as having...) Figure 16 The process 1300 is executed by a computing device of computing device architecture 1600 (as shown). In some examples, process 1300 may be executed by a computing device of computing device architecture 1600 having a decoding device (e.g., decoding device 112) or a client device including or communicating with a decoding device. In some examples, process 1400 may be executed by a computing device of computing device architecture 1600 having an encoding device (e.g., encoding device 104) or a server including or communicating with an encoding device. In an illustrative example, the computing device (e.g., executing process 1300) may include an extended reality display device, and the encoding device (e.g., executing process 1400) may be part of a server. In such an example, the encoding device is configured to generate a video bitstream for processing (e.g., decoding) and display by the extended reality display device based on motion information (e.g., pose, orientation, motion, etc.) provided by the extended reality display device and received by the encoding device (and / or the server). For example, an extended reality display device (or a device connected to an extended reality display device, such as a mobile device or other device) can transmit or send motion information (e.g., posture information, orientation, eye position, and / or other information) to a server. The server or a portion thereof (e.g., from...) Figure 2 The media source engine 204 can generate extended reality content based on motion information, and the encoding device can generate a video bitstream that includes decoded frames (or images) with extended reality content.
[0219] In some cases, a computing device or apparatus may include an input device, an output device, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of processes 1300 and / or 1400. Components of the computing device (e.g., one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in circuitry. For example, a component may include circuitry or other electronic hardware, and / or may be implemented using circuitry or other electronic hardware, which may include one or more programmable circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The computing device may also include a display (as an example of an output device or in addition to an output device), a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP)-based data or other types of data.
[0220] Processes 1300 and 1400 are shown as logic flowcharts, the operations of which represent a series of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media, which, when executed by one or more processors, perform the described operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0221] Additionally, process 1300 and / or process 1400 can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As noted above, the code can be stored, for example, in the form of a computer program comprising multiple instructions executable by one or more processors on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.
[0222] As previously mentioned, this paper also describes systems and techniques for synchronizing encoding devices with a common reference clock, along with other encoding devices. Figure 10 This is a diagram illustrating a system comprising multiple encoding devices (including encoding devices 1030a, 1030b to 1030n, where n can be any integer value) communicating with network device 1032 via wireless network 1031. Network entity 1032 may include a wireless access point (AP), a server (e.g., a server containing the encoding devices, such as...) Figure 2 The server side 202), base station, router, bridge, network gateway, or other network device or system. The wireless network 1031 may include a broadband network (e.g., 5G, 4G or LTE, 3G or other broadband networks), a WiFi network, or other wireless networks.
[0223] Network device 1032 communicates with one or more client devices (e.g., personal computer 1034a, XR headset 1034b, and mobile smartphone 1034c) via wireless network 1033. Wireless network 1033 may include a broadband network (e.g., 5G, 4G, or LTE, 3G, or other broadband networks), a WiFi network, or other wireless networks. In one illustrative example, wireless network 1031 includes a broadband network, and wireless network 1033 includes a WiFi network. In another illustrative example, wireless network 1031 and wireless network 1033 include broadband networks (e.g., the same broadband network or different broadband networks).
[0224] Encoding devices 1030a, 1030b through 1030n are synchronized with a common reference clock. This common reference clock is set by network device 1032. For example, network device 1032 may provide a time reference (e.g., frame insertion scheduling) to each encoding device 1030a, 1030b through 1030n for transmitting encoded data, where each time reference is unique for a given encoding device. The encoded data may represent a complete frame, an eye buffer, or a slice (or group of slices). Although Figure 10 Three encoding devices are shown, but those skilled in the art will recognize that any number (fewer or more than three) of encoding devices can be synchronized with a common reference clock.
[0225] In some cases, encoding devices 1030a, 1030b, to 1030n can utilize periodic (non-dynamic) I-frames / I-slices to synchronize with a common clock, so that I-frames / I-slices from different users can be interleaved to minimize the total number of I-frames / I-slices at any given time. For example, when each encoding device 1030a, 1030b, to 1030n is connected to network device 1032, network device 1032 can assign an index to each encoding device 1030a, 1030b, to 1030n. Network device 1032 can provide a global reference clock for the different encoding devices 1030a, 1030b, to 1030n to send I-frames / I-slices at separate time intervals based on their assigned indices. Such a technique can be beneficial in various use cases, such as in strict CBR coding structures, periodic intra-frame refresh coding structures, etc.
[0226] Figure 11 This shows two encoding devices synchronized with a reference clock (e.g., from...). Figure 10 A diagram illustrating an example of a video decoding structure for two of the encoding devices 1030a, 1030b to 1030n. Figure 11The decoding structure shown is illustrated with periodic I-frame insertion. When each encoding device connects (e.g., at initial connection) to the network device, the network device (e.g., network device 1032) can assign unique indices to the first and second encoding devices, respectively. The network device can provide a global common reference clock for all encoding devices, which defines a frame insertion schedule that instructs the first and second encoding devices (and any other encoding devices connected to it) when to transmit I-frames, I-slices, and / or other intra-frame decoded video data at separate time intervals based on their assigned indices. Each encoding device can then transmit I-frames and / or I-slices (or other intra-frame decoded data) based on its assigned indices and the reference clock, resulting in interleaved I-frames / slices from different encoding devices.
[0227] As shown in the figure, the first encoding device (corresponding to the first user) inserts I-frames (including I-frames 1102 and 1104) every four frames according to a common reference clock set for frame insertion scheduling. According to the reference clock, the second encoding device (corresponding to the second user) is scheduled to periodically insert I-frames every four frames, but starting from a different frame in the time domain. For example, the second encoding device inserts the first I-frame 1106 at a different time than the first I-frame 1102 inserted by the first encoding device, based on different indices assigned to the first and second encoding devices.
[0228] In some cases, encoding devices (e.g., encoding devices 1030a, 1030b to 1030n) can be synchronized with a common reference clock and can dynamically insert I-frames and / or I-slices. For example, when an encoding device receives feedback indicating one or more missing or corrupted packets and needs to force an I-frame or I-slice, any non-urgent (e.g., non-feedback-based) I-frame and / or I-slice insertions from other encoding devices can be delayed so that an encoding device can insert an I-frame or I-slice as soon as possible. In an illustrative example, based on feedback information, the affected encoding device can request a network device (e.g., network device 1032) to adapt to the common reference clock to allow the encoding device to insert an I-frame or I-slice immediately, for example, at the next frame of the bitstream (with or without an intra-frame refresh period). Based on the updated reference clock from the network device, other encoding devices can accordingly adapt their scheduling for non-urgent I-frames / I-slices so that there is no overlap in the I-frames / I-slices. Such techniques can be beneficial in a variety of use cases, such as in strict CBR coding structures, periodic intra-frame refresh coding structures, when inserting dynamic individual I slices, and so on.
[0229] Figure 12 This shows two encoding devices synchronized with a common reference clock (e.g., from...). Figure 10A diagram showing another example of a video decoding structure for two of the encoding devices 1030a, 1030b to 1030n. Figure 12 The decoding structure shown is illustrated as having periodic and dynamic insertions of I-frames and / or one or more I-slices. The reference clock is provided by the network device (e.g., from...). Figure 10 Network device 1032 is configured with an initial (or original) frame insertion schedule, which schedules when each encoding device will insert an I-frame and / or one or more I-slices (with or without an intra-frame refresh period). As shown, the first encoding device (corresponding to the first user) inserts an I-frame every four frames based on the index assigned to the first encoding device, including I-frame 1202 and I-frame 1204. Similar to... Figure 11 For example, the second encoding device (corresponding to the second user) is scheduled to periodically insert I-frames every four frames based on an index assigned to the second encoding device, but starting from a different frame in the time domain. For instance, the second encoding device inserts the first I-frame 1206 at a different time than the first encoding device inserts the first I-frame 1202 (two frames after the periodically scheduled I-frame inserted by the first encoding device).
[0230] At frame 1208, the client device receiving the bitstream generated by the first encoding device (e.g., Figure 10 The XR headset 1034b can detect the presence of a lost or corrupted packet from frame 1208, or that frame 1208 is lost. The client device (or, in some cases, the network device) can send feedback information to the first encoding device. Upon receiving the feedback information, the affected encoding device can request the network device to adapt the reference clock to allow the encoding device to immediately insert an I-frame, I-slice, or intra-frame refresh period in the next available frame (which may be an unscheduled time slot not scheduled in the initial frame insertion schedule). The network device can then update the reference clock to the updated schedule based on the request, and all remaining encoding devices can stop inserting I-frames according to their initially assigned insertion schedule and adapt their I-frame / I-slice schedules according to the updated reference clock from the network device.
[0231] For example, based on the updated reference clock set to the updated schedule, the first encoding device can force I-frame 1210 into the bitstream, and then continue periodically inserting I-frames every four frames thereafter (starting from I-frame 1212). As described above, once the network device updates the reference clock to the updated schedule based on a request, all remaining encoding devices can adapt their I-frame / I-slice schedules according to the updated reference clock. Figure 12As shown, because the first encoding device dynamically inserts I-frame 1210 in this time slot, the second encoding device inserts P-frame 1216 instead of inserting a periodically scheduled I-frame four frames after the periodically scheduled I-frame 1214 (as defined by the initial frame insertion schedule). Based on the updated reference clock and frame insertion schedule, the second encoding device continues to insert periodically scheduled I-frames every four frames (starting from I-frame 1218, two frames after the first encoding device inserts the dynamically inserted I-frame 1210).
[0232] In a multi-user environment, synchronizing multiple encoding devices with a common reference clock can be helpful. Synchronization with a common reference clock can also help reduce bit rate fluctuations on the radio link, regardless of the encoding configuration. For example, because multiple I-frames or I-slices will not be transmitted during the same time slot (based on synchronization), bit rate fluctuations on the radio link can be reduced.
[0233] Figure 15 This is a flowchart illustrating an example of a process 1500 for processing video data. At block 1502, process 1500 includes: by an encoding device (e.g., Figure 1 Encoding device 104, Figure 10 The encoding device 1030a or other encoding devices generate a video bitstream. (For example, the encoding device inserts intra-frame decoded video data into the video bitstream according to a reference clock shared with at least one other encoding device.) The reference clock defines the scheduling for interleaving intra-frame decoded video from the encoding device and at least one other encoding device. In some cases, multiple encoding devices are synchronized with the reference clock. In such cases, a different time reference (such as an index) can be assigned to each of the multiple encoding devices, and the encoded data can be sent according to the time reference (e.g., the first encoding device is assigned a first time reference, the second encoding device is assigned a second time reference, the third encoding device is assigned a third time reference, and so on). The first time reference assigned to an encoding device is different from the second time reference assigned to at least one other encoding device. In some examples, the reference clock can be generated by a network device (e.g., from...). Figure 10 The network device (1032) may be configured, or configured by another device or system. The network device may include a wireless access point (AP), a server including an encoding device or another encoding device, or a separate server excluding one of the encoding devices that adheres to a reference clock. The device may provide each encoding device with a time reference for transmitting encoded data, wherein each time reference is unique for a given encoding device. The encoded data may represent a complete frame, an eye buffer, or a slice (or group of slices).
[0234] At block 1504, process 1500 includes: obtaining feedback information from the encoding device indicating that at least a portion of a video slice, or at least a portion of a video frame or picture, of the video bitstream is missing or corrupted. For example, the feedback information may include information indicating that a video slice has one or more missing packets. In some examples, the information may include indications of frames or pictures that include missing or corrupted slices and of one or more slices that include missing packets.
[0235] At block 1506, process 1500 includes: sending a request to adapt a reference clock in response to feedback information to allow the encoding device to insert intra-frame decoded video data into the video bitstream in unscheduled time slots. The request may be sent to a network device (e.g., an access point, a server, etc.) that sets the reference clock. At block 1508, process 1500 includes: (e.g., from the device that sets the reference clock) receiving an indication that the reference clock has been updated to define an updated schedule.
[0236] At box 1510, process 1500 includes: inserting intra-frame decoded video data into the video bitstream based on an updated scheduling, according to an updated reference clock. Unscheduled time slots requested by the encoding device deviate from multiple time slots defined by the encoding device's reference clock. For example, as... Figure 12 As shown, the first encoding device dynamically inserts I-frame 1210 into a time slot different from the time slot originally scheduled for the first encoding device according to the original scheduling based on the reference clock. When the clock is reset based on a request from the first encoding device, the first encoding device will continue to periodically insert I-frames every four frames until it encounters a lost or corrupted frame or resets the clock again based on a request from another encoding device.
[0237] An updated reference clock can be shared with at least one other encoding device (e.g., one or more encoding devices from encoding devices 1030b to 1030n). In some examples, based on the updated scheduling, at least one other encoding device delays the scheduling of intra-frame decoded video relative to the previously scheduled time slots defined by the reference clock. For example, referencing again... Figure 12 Since the first encoding device dynamically inserts I-frame 1210 in this time slot, the second encoding device inserts P-frame 1216 based on the updated reference clock, instead of inserting periodically scheduled I-frames four frames after periodically scheduled I-frame 1214 (as defined by the initial frame insertion schedule).
[0238] In some examples, intra-frame decoded video data includes one or more intra-frame decoded video frames. For example, intra-frame decoded video data may include one or more intra-frame decoded video slices. In another example, intra-frame decoded video data may include an intra-frame refresh period, which includes at least one video frame. For example, each video frame in at least one video frame may include one or more intra-frame decoded video slices. In yet another example, intra-frame decoded video data may include a complete I-frame.
[0239] In some implementations, the feedback information is provided from a computing device. In some implementations, the computing device may include an extended reality display device, and the encoding device performing process 1500 may be part of a server. The encoding device is configured to generate a video bitstream for processing (e.g., decoding) and display by the extended reality display device based on motion information (e.g., pose, orientation, movement, etc.) provided by the extended reality display device and received by the encoding device (and / or the server). For example, the extended reality display device (or a device connected to the extended reality display device, such as a mobile device or other device) may transmit or send motion information (e.g., pose information, orientation, eye position, and / or other information) to the server. The server or part of the server (e.g., from...) Figure 2 The media source engine 204 can generate extended reality content based on motion information, and the encoding device can generate a video bitstream that includes decoded frames (or images) with extended reality content.
[0240] In some examples, process 1500 may be performed by a computing device or apparatus (such as having...) Figure 16The computing device (the computing device architecture 1600 shown) is used to perform the process 1500. In some examples, the process 1500 may be performed by a computing device of the computing device architecture 1600 that implements an encoding device (e.g., encoding device 104) or a server that includes or communicates with an encoding device. In some cases, the computing device or apparatus may include input devices, output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of the process 1500. Components of the computing device (e.g., one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in circuitry. For example, components may include circuitry or other electronic hardware, and / or may be implemented using circuitry or other electronic hardware, which may include one or more programmable circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The computing device may also include a display (as an example of an output device or in addition to an output device), a network interface configured to transmit and / or receive data, any combination thereof and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP) based data or other types of data.
[0241] Process 1500 is shown as a logic flowchart, the operations of which represent a series of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media, which, when executed by one or more processors, perform the described operations. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0242] Furthermore, process 1500 can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As noted above, the code can be stored, for example, in the form of a computer program comprising multiple instructions executable by one or more processors on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.
[0243] Figure 16 An example computing device architecture 1600 is shown, which can implement the various technologies described herein. Components of the computing device architecture 1600 are shown to be in electrical communication with each other via a connection 1605 (such as a bus). The example computing device architecture 1600 includes a processing unit (CPU or processor) 1610 and a computing device connection 1605 that couples various computing device components, including computing device memories 1615 (such as read-only memory (ROM) 1620 and random access memory (RAM) 1625), to the processor 1610.
[0244] The computing device architecture 1600 may include a cache of high-speed memory that is directly connected to, close to, or integrated into the processor 1610. The computing device architecture 1600 may copy data from memory 1615 and / or storage device 1630 to cache 1612 for fast access by the processor 1610. In this way, the cache can provide performance improvements by preventing latency for the processor 1610 while waiting for data. These and other modules may control or be configured to control the processor 1610 to perform various actions. Other computing device memory 1615 may also be available. Memory 1615 may include multiple different types of memory with different performance characteristics. The processor 1610 may include any general-purpose processor and hardware or software services configured to control the processor 1610 (such as services 1 1632, 2 1634, and 3 1636 stored in storage device 1630), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1610 can be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.
[0245] To enable user interaction with the computing device architecture 1600, the input device 1645 can represent any number of input mechanisms, such as a microphone for voice, a touchscreen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The output device 1635 can also be one or more of several output mechanisms known to those skilled in the art, such as a monitor, projector, television, speaker device, etc. In some instances, a multi-mode computing device allows the user to provide multiple types of input to communicate with the computing device architecture 1600. The communication interface 1640 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore, the basic features described here can be readily replaced by improved hardware or firmware arrangements (when they are developed).
[0246] Storage device 1630 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as magnetic tape, flash memory cards, solid-state storage devices, digital multifunction disks, magnetic tape cassettes, random access memory (RAM) 1625, read-only memory (ROM) 1620, and combinations thereof. Storage device 1630 may include services 1632, 1634, and 1636 for controlling processor 1610. Other hardware or software modules are contemplated. Storage device 1630 may be connected to computing device connection 1605. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components (such as processor 1610, connection 1605, output device 1635, etc.) to perform that function.
[0247] The technology disclosed herein is not necessarily limited to wireless applications or setups. The technology can be applied to video decoding to support any of the various multimedia applications described, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video (e.g., HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding digital video stored on data storage media, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0248] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and excludes carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as compact optical discs (CDs) or digital versatile optical discs (DVDs), flash memory, memory, or memory devices. Computer-readable media may have code and / or machine-executable instructions stored thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or sent via any suitable unit, including memory sharing, messaging, token passing, network transmission, etc.
[0249] In some embodiments, computer-readable storage devices, media, and memories may include cable or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.
[0250] Specific details have been provided in the foregoing description to provide a full understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that the embodiments described can be practiced without these specific details. For clarity, in some instances, the techniques described herein may be presented as comprising individual functional blocks including devices, device components, steps or routines in methods embodied in software, or combinations of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0251] The various embodiments described above can be presented as processes or methods, depicted as flowcharts, schematic diagrams, data flow diagrams, structural diagrams, or block diagrams. While a flowchart may describe operations as a sequential process, many of these operations can be performed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates upon completion of its operations, but may have additional steps not included in the diagram. A process can correspond to a method, function, process, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0252] The processes and methods described in the examples above can be implemented using computer-executable instructions, which are stored in or otherwise made available from a computer-readable medium. Such instructions may include, for example, instructions or data that cause a general-purpose computer, special-purpose computer, or processing device to perform or otherwise configure it to perform a particular function or a particular set of functions. A portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include hard disks or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, etc.
[0253] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take on any of a wide variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mount devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By further example, such functionality may also be implemented on a circuit board between different chips or different processes executed in a single device.
[0254] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example units for providing the functionality described in this disclosure.
[0255] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it is to be understood that the inventive concept may be embodied and employed differently in other ways, and the appended claims are intended to be interpreted as including such variations, except those limited by the prior art. Various features and aspects of the above-described applications may be used individually or in combination. Furthermore, embodiments may be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be recognized that, in alternative embodiments, the methods may be performed in a different order than that described.
[0256] Those skilled in the art will recognize that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.
[0257] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operation, programming programmable electronic circuits (e.g., microprocessors or other suitable circuits) to perform the operation, or any combination thereof.
[0258] The phrase “coupled to” refers to any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0259] The language of a claim that states "at least one of the members in the set" or other languages indicates that one or more members in the set satisfy the claim. For example, the language of a claim that states "at least one of A and B" means A, B, or A and B.
[0260] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above in relation to their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application; however, such implementation decisions should not be construed as a departure from the scope of this application.
[0261] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a wide variety of devices, such as general-purpose computers, mobile phones with wireless communication capabilities, or integrated circuit devices with multiple uses (including applications in mobile phones with wireless communication capabilities and other devices). Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code that, when executed, performs one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively or concurrently, the technology may be implemented, at least in part, by a computer-readable communication medium (such as a propagating signal or wave) that carries or transmits program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer.
[0262] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Accordingly, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
Claims
1. A method for processing video data, comprising: The computing device determines that at least a portion of the video slices in the video frame of the video bitstream is missing or corrupted; Sending feedback information to the encoding device, the feedback information indicating that at least a portion of the video slice is missing or corrupted, wherein the feedback information sent to the encoding device causes the encoding device to force an intra-frame refresh cycle to be inserted into the next one or more available frames of the video bitstream; and In response to the feedback information, an updated video bitstream is received from the encoding device, the updated video bitstream comprising at least one intra-frame decoded video slice having a size larger than the lost or corrupted video slice in the next one or more available frames of the updated video bitstream, wherein the size of the at least one intra-frame decoded video slice is determined to cover the lost or corrupted video slice and the propagation error in the video frame caused by the lost or corrupted video slice.
2. The method according to claim 1, wherein, The propagation error in the video frame caused by the lost or damaged video slice is based on the motion search range.
3. The method according to claim 1, wherein, The lost or damaged video slice spans from the first row to the second row in the video frame, and the size of the at least one intra-frame decoded video slice is defined as including the first row minus a multiple of the motion search range to the second row plus the multiple of the motion search range.
4. The method according to claim 1, further comprising: In response to determining that at least a portion of the video slice is missing or corrupted, error concealment is performed on one or more video frames until an error-free intra-decoded video slice is received in the updated video bitstream.
5. The method according to claim 1, wherein, The at least one intra-frame decoded video slice is a complete forced intra-frame decoded frame.
6. The method according to claim 1, wherein, The at least one intra-frame decoded video slice is included as part of an intra-frame refresh cycle, which distributes the intra-frame decoded block of an I-frame across several frames.
7. The method according to claim 6, wherein, The number of at least one video frame in the intra-frame refresh cycle is based on at least one of the following: the number of slices in the video frame including the video slice, the position of the lost or damaged video slice in the video frame, and the delay between the intra-frame refresh cycle inserted into the updated video bitstream and the corresponding feedback information.
8. The method according to claim 7, wherein, When the location of the lost or damaged video slice is the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle includes at least two frames.
9. The method according to claim 8, further comprising: Error hiding is performed on the first of the at least two frames but not on the second of the at least two frames, the second frame being after the first frame in the video bitstream.
10. The method according to claim 7, wherein, When the position of the lost or damaged video slice is not the first slice in the video frame, the intra-frame refresh cycle includes the intra-frame decoded frame.
11. The method according to claim 7, wherein, When the position of the lost or damaged video slice is the last slice in the video frame, the intra-frame refresh cycle includes at least two frames.
12. The method of claim 11, further comprising: Error hiding is performed on the first and second frames of the at least two frames.
13. The method according to claim 1, wherein, The computing device includes an extended reality display device configured to provide motion information to the encoding device to generate the video bitstream for display by the extended reality display device.
14. The method according to claim 3, wherein, The multiple of the motion search range includes the value 1.
15. The method according to claim 3, wherein, The multiple of the motion search range includes the value 2.
16. An apparatus for processing video data, the apparatus comprising: A memory configured to store video data; as well as A processor, which is implemented in a circuit and configured as follows: It is determined that at least a portion of the video slices in the video frames of the video bitstream is missing or corrupted; Sending feedback information to the encoding device, the feedback information indicating that at least a portion of the video slice is missing or corrupted, wherein the feedback information sent to the encoding device causes the encoding device to force an intra-frame refresh cycle to be inserted into the next one or more available frames of the video bitstream; and In response to the feedback information, an updated video bitstream is received from the encoding device, the updated video bitstream comprising at least one intra-frame decoded video slice having a size larger than the lost or corrupted video slice in the next one or more available frames of the updated video bitstream, wherein the size of the at least one intra-frame decoded video slice is determined to cover the lost or corrupted video slice and the propagation error in the video frame caused by the lost or corrupted video slice.
17. The apparatus according to claim 16, wherein, The propagation error in the video frame caused by the lost or damaged video slice is based on the motion search range.
18. The apparatus according to claim 16, wherein, The lost or damaged video slice spans from the first row to the second row in the video frame, and the size of the at least one intra-frame decoded video slice is defined as including the first row minus a multiple of the motion search range to the second row plus the multiple of the motion search range.
19. The apparatus according to claim 16, wherein, The processor is also configured to: In response to determining that at least a portion of the video slice is missing or corrupted, error concealment is performed on one or more video frames until an error-free intra-decoded video slice is received in the updated video bitstream.
20. The apparatus according to claim 16, wherein, The at least one intra-frame decoded video slice is included as part of an intra-frame refresh cycle, which distributes the intra-frame decoded block of an I-frame across several frames.
21. The apparatus according to claim 20, wherein, The number of at least one video frame in the intra-frame refresh cycle is based on at least one of the following: the number of slices in the video frame including the video slice, the position of the lost or damaged video slice in the video frame, and the delay between the intra-frame refresh cycle inserted into the updated video bitstream and the corresponding feedback information.
22. The apparatus according to claim 21, wherein, When the location of the lost or damaged video slice is the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle includes at least two frames.
23. The apparatus according to claim 22, wherein, The processor is also configured to: Error hiding is performed on the first of the at least two frames but not on the second of the at least two frames, the second frame being after the first frame in the video bitstream.
24. The apparatus according to claim 21, wherein, When the position of the lost or damaged video slice is not the first slice in the video frame, the intra-frame refresh cycle includes the intra-frame decoded frame.
25. The apparatus according to claim 21, wherein, When the position of the lost or damaged video slice is the last slice in the video frame, the intra-frame refresh cycle includes at least two frames.
26. The apparatus according to claim 25, wherein, The processor is also configured to: Error hiding is performed on the first and second frames of the at least two frames.
27. The apparatus according to claim 16, wherein, The apparatus includes an extended reality display device configured to provide motion information to the encoding device to generate the video bitstream for display by the extended reality display device.
28. The apparatus according to claim 18, wherein, The multiple of the motion search range includes the value 1.
29. The apparatus according to claim 18, wherein, The multiple of the motion search range includes the value 2.
30. A non-transitory computer-readable medium having instructions stored thereon, the instructions causing the one or more processors, when executed, to perform the following operations: It is determined that at least a portion of the video slices in the video frames of the video bitstream is missing or corrupted; Sending feedback information to the encoding device, the feedback information indicating that at least a portion of the video slice is missing or corrupted, wherein the feedback information sent to the encoding device causes the encoding device to force an intra-frame refresh cycle to be inserted into the next one or more available frames of the video bitstream; and In response to the feedback information, an updated video bitstream is received from the encoding device, the updated video bitstream comprising at least one intra-frame decoded video slice having a size larger than the lost or corrupted video slice in the next one or more available frames of the updated video bitstream, wherein the size of the at least one intra-frame decoded video slice is determined to cover the lost or corrupted video slice and the propagation error in the video frame caused by the lost or corrupted video slice.
31. The non-transitory computer-readable medium according to claim 30, wherein, The lost or damaged video slice spans from the first row to the second row in the video frame, and the size of the at least one intra-frame decoded video slice is defined as including the first row minus a multiple of the motion search range to the second row plus the multiple of the motion search range.
32. The non-transitory computer-readable medium according to claim 30, wherein, The at least one intra-frame decoded video slice is included as part of an intra-frame refresh cycle, which distributes the intra-frame decoded block of an I-frame across several frames.
33. A method for processing video data, comprising: At the encoding device, feedback information is received from the computing device, the feedback information indicating that at least a portion of a video slice of a video frame in the video bitstream is missing or corrupted; as well as In response to the feedback information, an updated video bitstream is generated by forcing an intra-frame refresh period to be inserted into the next one or more available frames of the video bitstream. The updated video bitstream includes at least one intra-frame decoded video slice in one or more intra-frame refresh period frames having a size larger than the lost or corrupted video slice, wherein the size of the at least one intra-frame decoded video slice is determined to cover the lost or corrupted video slice and the propagation error in the video frame caused by the lost or corrupted video slice.
34. The method according to claim 33, wherein, The propagation error in the video frame caused by the lost or damaged video slice is based on the motion search range.
35. The method according to claim 33, wherein, The lost or corrupted video slice spans from the first line to the second line of the video frame, and also includes: The size of the at least one intra-frame decoded video slice is determined to include the first row minus a multiple of the motion search range to the second row plus the multiple of the motion search range.
36. The method according to claim 33, wherein, In response to at least a portion of the video slice being lost or corrupted, error concealment is performed on one or more video frames until an error-free intra-decoded video slice is received in the updated video bitstream.
37. The method according to claim 33, wherein, The at least one intra-frame decoded video slice is included as part of an intra-frame refresh cycle, which distributes the intra-frame decoded block of an I-frame across several frames.
38. The method of claim 37, further comprising: The number of at least one video frame in the intra-frame refresh cycle is determined based on at least one of the following: the number of slices in the video frame including the video slice, the location of the lost or damaged video slice in the video frame, and the delay between the intra-frame refresh cycle inserted into the updated video bitstream and the corresponding feedback information.
39. The method according to claim 38, wherein, When the location of the lost or damaged video slice is the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle is determined to include at least two frames.
40. The method according to claim 39, wherein, Error hiding is performed on the first of the at least two frames but not on the second of the at least two frames, the second frame being after the first frame in the video bitstream.
41. The method according to claim 38, wherein, When the position of the lost or damaged video slice is not the first slice in the video frame, the intra-frame refresh period is determined to include the intra-frame decoded frame.
42. The method according to claim 38, wherein, When the position of the lost or damaged video slice is the last slice in the video frame, the intra-frame refresh period is determined to include at least two frames.
43. The method according to claim 42, wherein, Error hiding is performed on the first and second frames of the at least two frames.
44. The method according to claim 33, wherein, The computing device includes an extended reality display device, and wherein the encoding device is part of a server, the encoding device being configured to generate the video bitstream based on motion information received by the encoding device from the extended reality display device for display by the extended reality display device.
45. The method of claim 33, further comprising: Intra-decoded video data is added to the video bitstream according to a reference clock shared with at least one other encoding device, the reference clock being defined for interleaving the intra-decoded video from the encoding device and the at least one other encoding device; In response to the feedback information, a request to adapt to the reference clock is sent to allow the encoding device to add intra-frame decoded video data to the video bitstream in unscheduled time slots; Receive an indication that the reference clock has been updated to define an updated schedule; as well as Based on the updated scheduling, at least one intra-frame decoded video slice is added to the video bitstream according to the updated reference clock.
46. The method of claim 33, further comprising: The updated video bitstream is sent to the computing device.
47. The method of claim 33, further comprising: Store the updated video bitstream.
48. The method according to claim 35, wherein, The multiple of the motion search range includes the value 1 or the value 2.
49. An apparatus for processing video data, the apparatus comprising: A memory configured to store video data; as well as A processor, which is implemented in a circuit and configured as follows: Receive feedback information from a computing device, the feedback information indicating that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted; as well as In response to the feedback information, an updated video bitstream is generated by forcing an intra-frame refresh period to be inserted into the next one or more available frames of the video bitstream. The updated video bitstream includes at least one intra-frame decoded video slice in one or more intra-frame refresh period frames having a size larger than the lost or corrupted video slice, wherein the size of the at least one intra-frame decoded video slice is determined to cover the lost or corrupted video slice and the propagation error in the video frame caused by the lost or corrupted video slice.
50. The apparatus according to claim 49, wherein, The propagation error in the video frame caused by the lost or damaged video slice is based on the motion search range.
51. The apparatus according to claim 49, wherein, The lost or corrupted video slice spans from the first line to the second line of the video frame, and the processor is further configured to: The size of the at least one intra-frame decoded video slice is determined to include the first row minus a multiple of the motion search range to the second row plus the multiple of the motion search range.
52. The apparatus according to claim 49, wherein, In response to determining that at least a portion of the video slice is missing or corrupted, error concealment is performed on one or more video frames until an error-free intra-decoded video slice is received in the updated video bitstream.
53. The apparatus according to claim 49, wherein, The at least one intra-frame decoded video slice is included as part of an intra-frame refresh cycle, which distributes the intra-frame decoded block of an I-frame across several frames.
54. The apparatus according to claim 53, wherein, The processor is configured to: The number of at least one video frame in the intra-frame refresh cycle is determined based on at least one of the following: the number of slices in the video frame including the video slice, the location of the lost or damaged video slice in the video frame, and the delay between the intra-frame refresh cycle inserted into the updated video bitstream and the corresponding feedback information.
55. The apparatus according to claim 54, wherein, When the location of the lost or damaged video slice is the first slice in the video frame, the at least one video frame in the intra-frame refresh cycle is determined to include at least two frames.
56. The apparatus according to claim 55, wherein, Error hiding is performed on the first of the at least two frames but not on the second of the at least two frames, the second frame being after the first frame in the video bitstream.
57. The apparatus according to claim 54, wherein, When the position of the video slice is not the first slice in the video frame, the intra-frame refresh period is determined to include the intra-frame decoded frame.
58. The apparatus according to claim 54, wherein, When the position of the lost or damaged video slice is the last slice in the video frame, the intra-frame refresh period is determined to include at least two frames.
59. The apparatus according to claim 58, wherein, Error hiding is performed on the first and second frames of the at least two frames based on the fact that the video slice is the last slice in the video frame.
60. The apparatus according to claim 49, wherein, The computing device includes an extended reality display device, and wherein the apparatus includes an encoding device as part of a server, the encoding device being configured to generate the video bitstream based on motion information received by the encoding device from the extended reality display device for display by the extended reality display device.
61. The apparatus of claim 49, further comprising: A transmitter configured to send the updated video bitstream to the computing device.
62. The apparatus according to claim 49, wherein, The memory is configured to store the updated video bitstream.
63. The apparatus according to claim 51, wherein, The multiple of the motion search range includes the value 1 or the value 2.
64. A non-transitory computer-readable medium having instructions stored thereon, the instructions causing the one or more processors to perform the following operations when executed by one or more processors: Receive feedback information from a computing device, the feedback information indicating that at least a portion of a video slice of a video frame in a video bitstream is missing or corrupted; and In response to the feedback information, an updated video bitstream is generated by forcibly inserting an intra-frame refresh period into the next one or more available frames of the video bitstream. The updated video bitstream includes at least one intra-frame decoded video slice in one or more intra-frame refresh period frames that has a size larger than the lost or corrupted video slice. The size of the at least one intra-frame decoded video slice is determined to cover the lost or damaged video slice and the propagation error in the video frame caused by the lost or damaged video slice.
65. The non-transitory computer-readable medium according to claim 64, wherein, The lost or corrupted video slice spans from the first line to the second line of the video frame, and also includes instructions that, when executed by the one or more processors, cause the one or more processors to perform the following operations: The size of the at least one intra-frame decoded video slice is determined to include the first row minus a multiple of the motion search range to the second row plus the multiple of the motion search range.
66. The non-transitory computer-readable medium according to claim 64, wherein, The at least one intra-frame decoded video slice is included as part of an intra-frame refresh cycle, which distributes the intra-frame decoded block of an I-frame across several frames.
67. The non-transitory computer-readable medium according to claim 65, wherein, The multiple of the motion search range includes the value 1 or the value 2.
Citation Information
Patent Citations
Transmitted / received data processing method and recording medium thereof
CN1307430A
Moving image decoding apparatus and processing method thereof
US20090245385A1
Method, System and Apparatus for Intra-Refresh in Video Signal Processing
US20130114697A1
Motion-adaptive intra-refresh for high-efficiency, low-delay video coding
US20170318308A1