Decentralized video coding using reliability data
Distributed video coding offloads encoding tasks to a receiving device with more resources, addressing the resource-intensive nature of modern video encoding processes by using error correction data for efficient video reconstruction in XR devices.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2024-04-22
- Publication Date
- 2026-05-26
AI Technical Summary
Modern video encoding processes are resource-intensive and require complex processors, high-speed memory, and significant energy consumption, which is not suitable for wireless communication systems with limited bandwidth constraints, especially for XR devices like XR headsets that need to minimize weight and power consumption.
A distributed video coding (DVC) process is employed where a transmitting device performs a limited video coding process and transmits error correction data, allowing the receiving device to estimate and fully encode missing data using error correction, reducing the complexity and resource consumption at the transmitting end.
This approach enables efficient video reconstruction at the receiving device while minimizing resource usage and power consumption at the transmitting device, maintaining low latency and high-quality video transmission.
Smart Images

Figure 2026516654000001_ABST
Abstract
Description
[Technical Field]
[0001]
[0001] This application claims priority to U.S. Patent Application No. 18 / 640,692, filed on April 19, 2024, and U.S. Provisional Patent Application No. 63 / 497,979, filed on April 24, 2023, both of which are incorporated by reference in their entirety. U.S. Patent Application No. 18 / 640,892, filed on April 19, 2024, claims the benefit of U.S. Provisional Patent Application No. 63 / 497,979, filed on April 24, 2023.
[0002]
[0002] This disclosure relates to video coding and video decoding. [Background technology]
[0003]
[0003] The popularity of virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies is rapidly increasing and is expected to be widely adopted in non-gaming applications such as healthcare, education, social media, retail, and many others. VR, AR, and MR are sometimes collectively referred to as extended reality (XR). This growing popularity is driving demand for XR devices, such as XR goggles, which feature high-quality 3D graphics, higher video resolution, and low-latency response. [Overview of the Initiative]
[0004]
[0004] This disclosure describes techniques for processing video data in a transmitting device and a receiving device. The transmitting device may be an XR device or other type of device. The receiving device may be a user equipment (UE) device such as a smartphone or tablet. The transmitting device may perform a limited video coding process on the video data to generate encoded video data. The transmitting device may apply channel coding to the encoded video data to generate error correction data. The transmitting device may send the error correction data and at least a portion of the encoded video data to the receiving device. The receiving device may estimate the video data based on one or more previously reconstructed pictures. The receiving device may then encode the estimated video data. When the receiving device performs a limited video coding process on the video data, it may use one or more coding tools to encode the estimated video data that was not used by the transmitting device. The receiving device may use the error correction data and the estimated video data to play back the portion of the encoded video data that was not sent by the transmitting device. This process may eliminate the need to send the portion of the encoded video data.
[0005]
[0005] In one example, the present disclosure describes a method for decoding video data, which includes: receiving error correction data from a transmitting device in a receiving device, which provides error correction information and is generated based on encoded video data of one or more blocks of picture in the video data; generating prediction data for picture, which includes predictions of blocks of picture based at least in part on blocks of picture that have been reconstructed previously in one or more video data, using one or more coding tools not used to generate encoded video data of one or more blocks in the receiving device; generating encoded video data based on the prediction data for picture in the receiving device; generating error-corrected encoded video data using the error correction data in order to perform an error correction operation on the encoded video data in the receiving device; and performing a reconstruction operation in the receiving device to reconstruct blocks of picture based on the error-corrected encoded video data, wherein the reconstruction operation is controlled by the values of one or more parameters.
[0006]
[0006] In another example, the Disclosure describes a method for encoding video data, which includes: a transmitting device acquiring video data from a video source; a transmitting device generating encoded video data of a first picture and encoded video data of a second picture of the video data based on a set of parameters; a transmitting device performing channel coding on the encoded video data of the first picture and the encoded video data of the second picture in order to generate error correction data for the first picture and error correction data for the second picture; and a transmitting device transmitting the encoded video data of the first picture, the error correction data for the first picture, and the error correction data for the second picture.
[0007]
[0007] In another example, the Disclosure describes a method for encoding video data, comprising: a transmitting device acquiring video data from a video source; a transmitting device generating transformation blocks based on the video data; a transmitting device determining which of the transformation blocks are anchor transformation blocks; a transmitting device computing a correlation matrix for the set of transformation blocks; a transmitting device generating a bit-reduced non-anchor transformation matrix; and a transmitting device transmitting the anchor transformation blocks, the non-anchor transformation blocks, and the correlation matrix to a receiving device.
[0008]
[0008] In another example, the present disclosure describes a device comprising a memory configured to store video data, a communication interface, and one or more processors implemented in a circuit and coupled to the memory, the one or more processors being configured to perform any of the methods of the claims.
[0009]
[0009] In another example, the Disclosure describes a device for processing video data, comprising a memory configured to store video data, a communication interface configured to receive error correction data from a transmitting device, which provides error correction information relating to pictures of video data, and one or more processors implemented in the circuit and coupled to the memory, the one or more processors being configured to generate prediction data for pictures, which includes predictions of blocks of pictures based at least in part on one or more previously reconstructed pictures of video data; to generate encoded video data, which includes encoded blocks, which include transform coefficients, based on the prediction data for pictures; to scale the bits of the transform coefficients of the transform blocks based on confidence values for bit positions; to generate error-corrected encoded video data using the error correction data to perform error correction operations on the scaled bits of the transform coefficients of the transform blocks; and to reconstruct pictures based on the error-corrected encoded video data.
[0010]
[0010] In another example, the Disclosure describes a device for processing video data, which includes a memory configured to store video data, and one or more processors implemented in a circuit and coupled to the memory, which are configured to acquire video data, acquire predictive quality feedback, which is based on the reliability of estimated pictures produced by a receiving device, adapt one or more of the video coding parameters or channel coding parameters based on the predictive quality feedback, and to produce coded video data based on one or more pictures of the acquired video data, and to produce channel coded data, which are configured to produce channel coded data, which are configured to produce channel coded data, and a communication interface configured to transmit the channel coded data to a receiving device.
[0011]
[0011] In another example, the Disclosure describes a method for processing video data, which includes: receiving error correction data from a transmitting device in a receiving device, which provides error correction information relating to pictures of video data; generating prediction data for pictures in a receiving device, which includes predictions of blocks of pictures based at least in part on pictures that have been reconstructed one or more times before the video data; generating encoded video data in a receiving device based on the prediction data for pictures, which includes encoded video data that includes a transformation block containing transformation coefficients; scaling the bits of the transformation coefficients of the transformation block in a receiving device based on confidence values for bit positions; generating error-corrected encoded video data in a receiving device using the error correction data to perform error correction operations on the scaled bits of the transformation coefficients of the transformation block; and reconstructing a picture in a receiving device based on the error-corrected encoded video data.
[0012]
[0012] In another example, the Disclosure describes a method for processing video data, which includes acquiring video data; acquiring predictive quality feedback, which is based on the reliability of an estimated picture produced by a receiving device; adapting one or more video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process, controlled by video coding parameters, to generate encoded video data based on one or more pictures of the acquired video data; performing a channel coding process, controlled by channel coding parameters, on the encoded video data to generate channel coded data; and transmitting the channel coded data to a receiving device.
[0013]
[0013] In another example, the present disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, the one or more processors acquiring a first set of multiview pictures of video data, wherein the first set of multiview pictures comprises a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint, and transmitting to a receiving device first encoded video data, wherein the first encoded video data is based on the first set of multiview pictures, and receiving from the receiving device multiview encoded video data The system is configured to receive a stream, obtain a second set of multiview pictures of video data, the second set of multiview pictures including a third picture and a fourth picture, where the third picture is from a first viewpoint and the fourth picture is from a second viewpoint, and, based on the multiview coding queue received from the receiving device, to generate second coded video data, by performing a multiview coding process on the second set of multiview pictures, the multiview coding process reducing interview redundancy between the third picture and the fourth picture, and then sending the second coded video data to the receiving device.
[0014]
[0014] In another example, the present disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, the one or more processors receiving from the transmitting device first encoded video data, wherein the first encoded video data is based on a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint. The system is configured to determine a multiview coding queue based on the first coded video data, send the multiview coding queue to a transmitting device, and retrieve from the transmitting device a second coded video data, the second coded video data being based on a second set of multiview pictures including a third picture and a fourth picture, and the second coded video data being coded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.
[0015]
[0015] In another example, the present disclosure obtains a first set of multi-view pictures of video data, where the first set of multi-view pictures includes a first picture and a second picture, the first picture is from a first viewpoint, and the second picture is from a second viewpoint; transmits to a receiving device first encoded video data, where the first encoded video data is based on the first set of multi-view pictures; receives a multi-view encoding queue from the receiving device; obtains a second set of multi-view pictures of video data, where the second set of multi-view pictures includes a third picture and a fourth picture, the third picture is from the first viewpoint, and the fourth picture is from the second viewpoint; performs, on the second set of multi-view pictures, a multi-view encoding process that reduces inter-view redundancy between the third picture and the fourth picture, to generate second encoded video data based on the multi-view encoding queue received from the receiving device; and transmits the second encoded video data to the receiving device, to describe a method for processing video data.
[0016]
[0016] In another example, the Disclosure describes a method for processing video data, which includes: obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of video data, the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; determining a multiview coding queue based on the first encoded video data; transmitting the multiview coding queue to the transmitting device; and obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multiview pictures including a third picture and a fourth picture, and the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.
[0017]
[0017] In another example, the present disclosure provides means for obtaining a first set of multi-view pictures of video data, where the first set of multi-view pictures includes a first picture and a second picture, the first picture is from a first viewpoint, and the second picture is from a second viewpoint; means for transmitting to a receiving device first encoded video data, where the first encoded video data is based on the first set of multi-view pictures; means for receiving a multi-view encoding queue from the receiving device; means for obtaining a second set of multi-view pictures of video data, where the second set of multi-view pictures includes a third picture and a fourth picture, the third picture is from the first viewpoint, and the fourth picture is from the second viewpoint; means for performing a multi-view encoding process on the second set of multi-view pictures to generate second encoded video data based on the multi-view encoding queue received from the receiving device, where the multi-view encoding process reduces inter-view redundancy between the third picture and the fourth picture; and means for transmitting the second encoded video data to the receiving device, to describe a device.
[0018]
[0018] In another example, the present disclosure describes a device that includes means for obtaining from a transmitting device first encoded video data, wherein the first encoded video data is based on a first set of multiview pictures of video data, the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; means for determining a multiview coding queue based on the first encoded video data; means for transmitting the multiview coding queue to the transmitting device; and means for obtaining from the transmitting device second encoded video data second encoded video data, wherein the second encoded video data is based on a second set of multiview pictures including a third picture and a fourth picture, and the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.
[0019]
[0019] In another example, the present disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, the one or more processors being configured to encode a first set of pictures of video data to generate first encoded video data, transmit the first encoded video data to a receiving device, receive from the receiving device a decimation pattern instruction indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data, encode a second set of pictures of video data to generate second encoded video data, apply the decimation pattern to the second encoded video data to generate decimated video data, and transmit the decimated video data to a receiving device.
[0020]
[0020] In another example, the present disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, the one or more processors receiving first encoded video data from a transmitting device, performing a decoding process to reconstruct a first set of pictures based on the first encoded video data, determining a decimation pattern indicating a non-transmitted pattern of the encoded video data based on the first set of pictures, transmitting a decimation pattern instruction to the transmitting device indicating the determined decimation pattern, receiving decimated video data from the transmitting device, wherein the decimated video data comprises second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data, and is configured to perform a decoding process to reconstruct a second set of pictures based on the second encoded video data.
[0021]
[0021] In another example, the present disclosure describes a method comprising: encoding a first set of pictures of video data to generate a first encoded video data; transmitting the first encoded video data to a receiving device; receiving a decimation pattern instruction from the receiving device indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; encoding a second set of pictures of video data to generate a second encoded video data; applying the decimation pattern to the second encoded video data to generate decimated video data; and transmitting the decimated video data to a receiving device.
[0022]
[0022] In another example, the present disclosure describes a method comprising: receiving first encoded video data from a transmitting device; applying a decoding process to reconstruct a first set of pictures based on the first encoded video data; determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; transmitting a decimation pattern instruction to the transmitting device indicating the determined decimation pattern; receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.
[0023]
[0023] In another example, the present disclosure describes a device comprising: means for encoding a first set of pictures of video data to generate a first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving a decimation pattern instruction from the receiving device, which indicates a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; means for encoding a second set of pictures of video data to generate a second encoded video data; means for applying the decimation pattern to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to a receiving device.
[0024]
[0024] In another example, the present disclosure describes a device comprising means for receiving first encoded video data from a transmitting device; means for performing a decoding process to reconstruct a first set of pictures based on the first encoded video data; means for determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; means for transmitting a decimation pattern instruction to the transmitting device indicating the determined decimation pattern; means for receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and means for performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.
[0025]
[0025] In another example, the present disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, the one or more processors being configured to encode a first picture of video data to generate a first encoded video data, transmit the first encoded video data to a receiving device, receive from the receiving device encoded selection data for a second picture of video data, the encoded selection data for the second picture indicating an encoding selection used to encode an estimate of the second picture, the second picture following the first picture in decoding order, encode the second picture based on the encoded selection data for the second picture to generate a second encoded video data, and transmit the second encoded video data to a receiving device.
[0026]
[0026] In another example, the present disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, the one or more processors being configured to receive first encoded video data from a transmitting device, reconstruct a first picture of the video data based on the first encoded video data, estimate a second picture of the video data based on the first picture, the second picture being a picture that follows the first picture in the decoding order, generate encoding selection data for the second picture, the encoding selection data for the second picture indicating the encoding selection used to encode the second picture, transmit the encoding selection data for the second picture to the transmitting device, receive second encoded video data from the transmitting device, and reconstruct a second picture based on the second encoded video data.
[0027]
[0027] In another example, the present disclosure describes a method for processing video data, which includes encoding a first picture of video data to generate a first encoded video data; transmitting the first encoded video data to a receiving device; receiving from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in the decoding order; encoding a second picture based on the encoded selection data for the second picture to generate a second encoded video data; and transmitting the second encoded video data to a receiving device.
[0028]
[0028] In another example, the present disclosure describes a method for processing video data, which includes receiving a first encoded video data from a transmitting device; reconstructing a first picture of the video data based on the first encoded video data; estimating a second picture of the video data based on the first picture, wherein the second picture is a picture that follows the first picture in the decoding order; generating encoding selection data for the second picture, wherein the encoding selection data for the second picture indicates the encoding selection used to encode the second picture; transmitting the encoding selection data for the second picture to the transmitting device; receiving a second encoded video data from the transmitting device; and reconstructing a second picture based on the second encoded video data.
[0029]
[0029] In another example, the present disclosure describes a device comprising: means for encoding a first picture of video data to generate a first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in the decoding order; means for encoding a second picture based on the encoded selection data for the second picture to generate a second encoded video data; and means for transmitting the second encoded video data to a receiving device.
[0030]
[0030] In another example, the present disclosure describes a device comprising: means for receiving a first encoded video data from a transmitting device; means for reconstructing a first picture of the video data based on the first encoded video data; means for estimating a second picture of the video data based on the first picture, wherein the second picture is a picture that follows the first picture in the decoding order; means for generating encoding selection data for the second picture, wherein the encoding selection data for the second picture indicates the encoding selection used to encode the second picture; means for transmitting the encoding selection data for the second picture to a transmitting device; means for receiving a second encoded video data from a transmitting device; and means for reconstructing a second picture based on the second encoded video data.
[0031]
[0031] Details of one or more examples are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]
[0032] [Figure 1]
[0032] This is a block diagram illustrating an exemplary system using the technique of the present disclosure. [Figure 2]
[0033] This is a block diagram showing exemplary components of a transmitting device and a receiving device according to the techniques of the present disclosure. [Figure 3]
[0034] Figure 3A is a conceptual diagram illustrating an exemplary channel coding process using the technique of this disclosure.
[0035] Figure 3B is a block diagram showing an exemplary channel decoding process using the technique of the present disclosure. [Figure 4]
[0036] This is a flowchart illustrating exemplary operation of a transmitting device using the techniques of this disclosure. [Figure 5]
[0037] This is a flowchart illustrating exemplary operation of a receiving device using the techniques of this disclosure. [Figure 6]
[0038] This is a conceptual diagram illustrating an exemplary decimation pattern using the technique of this disclosure. [Figure 7]
[0039] This flowchart illustrates the exemplary operation of a transmitting device for hybrid decimation of a conversion block using the technique of the present disclosure. [Figure 8]
[0040] This flowchart illustrates the exemplary operation of a receiving device for hybrid decimation of a conversion block using the technique of the present disclosure. [Figure 9]
[0041] This is a conceptual diagram showing exemplary decimation patterns adaptively selected by a receiving device using one or more of the techniques of the present disclosure. [Figure 10]
[0042] This is a block diagram showing exemplary components of a transmitting device and a receiving device according to the techniques of the present disclosure. [Figure 11]
[0043] This figure shows exemplary error probabilities and corresponding absolute log-likelihood ratios (LLRs) using one or more techniques of the present disclosure. [Figure 12]
[0044] This flowchart illustrates exemplary operation of transmitting a device using scaled bits with the technique of the present disclosure. [Figure 13]
[0045] This flowchart illustrates exemplary operation of a receiving device using scaled bits according to the technique of this disclosure. [Figure 14]
[0046] This flowchart illustrates an exemplary exchange of data between a transmitting device and a receiving device for multiview processing using one or more techniques of the present disclosure. [Figure 15]
[0047] This flowchart illustrates exemplary operation of a transmitting device for multiview processing using the technique of the present disclosure. [Figure 16]
[0048] This flowchart illustrates exemplary operation of a receiving device for multiview processing using the technique of the present disclosure. [Figure 17]
[0049] This block diagram shows exemplary components of a transmitting and receiving device that perform decimation on encoded video data using the technique of the present disclosure. [Figure 18]
[0050] This is a conceptual diagram illustrating an exemplary exchange of information, including decimation pattern instructions, using the technique of the present disclosure. [Figure 19]
[0051] This flowchart illustrates the exemplary operation of a transmitting device in which it receives a decimation pattern instruction using the technique of the present disclosure. [Figure 20]
[0052] This flowchart illustrates the exemplary operation of a receiving device in which the receiving device transmits a decimation pattern instruction using the technique of the present disclosure. [Figure 21]
[0053] This block diagram shows exemplary components of a transmitting device and a receiving device that transmits encoded selection data to the transmitting device, according to the technique of the present disclosure. [Figure 22]
[0054] This is a communication diagram illustrating an exemplary exchange of data between a transmitting device and a receiving device, including the transmission and reception of coded selection data using the techniques of the present disclosure. [Figure 23]
[0055] This flowchart illustrates the exemplary operation of a transmitting device in which it receives encoded selection data using the technique of the present disclosure. [Figure 24]
[0056] This flowchart illustrates exemplary operation of a receiving device, in which the receiving device transmits encoded selection data using the technique of the present disclosure. [Figure 25]
[0057] This is a conceptual diagram illustrating an exemplary hierarchy of encoded video data using the techniques of this disclosure. [Figure 26]
[0058] This block diagram shows exemplary alternative components for a transmitting device using one or more techniques of the present disclosure. [Figure 27]
[0059] A block diagram showing exemplary alternative components of a receiving device using one or more techniques of the present disclosure. [Modes for carrying out the invention]
[0033]
[0060] Modern video encoding processes can significantly reduce the amount of data required to represent video data, but such processes are typically resource-intensive and can involve a lot of memory activity. Therefore, modern video encoding processes can require complex processors, high-speed memory, and consume considerable energy. However, some modern and planned future wireless communication systems, such as 5G and 6G wireless communication systems, may have fewer constraints on wireless transmission bandwidth, especially when communicating over short distances, such as between devices in contact with a person's body.
[0034]
[0061] This disclosure describes a technique that can reduce the complexity of video coding at a transmitting device by using error correction performed as part of channel decoding using error correction data. A transmitting device may perform a limited video coding process that generates coded video data. A limited video coding process typically uses coding tools such as intra-prediction, which are relatively resource-intensive. Because the video coding process uses less complex coding tools, the resulting coded video data may be larger than the coded video data coded using more complex and resource-intensive coding tools. Error correction data is based on coded video data. A transmitting device may transmit error correction data to a receiving device. A transmitting device may not need to transmit all of the coded video data for one or more pictures to the receiving device.
[0035]
[0062] The receiving device may estimate the picture in the video data based on one or more previously reconstructed pictures. In some cases, to estimate the picture, the receiving device may extrapolate the contents of a block from a previously reconstructed picture. The receiving device may then perform a full video encoding process on the estimated picture to generate estimated encoded video data for the picture. When performing a full video encoding process, the receiving device may use more complex coding tools, such as interpretation, than the limited video encoding process performed by the transmitting device. The receiving device may perform a channel decoding process to generate error-corrected encoded video data based on the estimated encoded video data for the picture and error-corrected data for the picture. In some situations, the channel decoding process may generate error-corrected encoded video data based on the error-corrected data for the picture and a combination of the estimated encoded video data for the picture and the encoded video data for the picture sent by the transmitting device. The receiving device may reconstruct the picture based on the error-corrected encoded video data. In this way, the receiving device may be able to reconstruct each picture of the video data even if the transmitting device did not transmit all of the encoded video data of the picture.
[0036]
[0063] As further described in this disclosure, various techniques can be applied, such as the application of decimation patterns, to specify which transform blocks of lightly encoded video data are not signaled or have reduced bit depth. Furthermore, in some examples of this disclosure, confidence values may be determined for bit positions, and the bits of the transform coefficients of the transform blocks may be scaled using the confidence values, and the scaled values may be used in channel coding and channel decoding.
[0037]
[0064] As further described in this disclosure, a receiving device may determine a decimation pattern based on a first set of pictures. The decimation pattern is a non-transmitted pattern of encoded video data. The receiving device may transmit a decimation pattern instruction to a transmitting device indicating the determined decimation pattern. The transmitting device receives the decimation pattern instruction from the receiving device and may apply the indicated decimation pattern to the encoded video data to generate decimated video data. The transmitting device may transmit the decimated video data to the receiving device. In this way, the technique of this disclosure can further reduce resource consumption in the transmitting device while still avoiding the transmission of excessive amounts of data. This can further improve coding efficiency.
[0038]
[0065] Figure 1 is a block diagram illustrating an exemplary system 100 using the techniques of the present disclosure. In the example of Figure 1, system 100 includes a transmitting device 102, a receiving device 104, and a base station 106. The transmitting device 102 may be a device configured to include an augmented reality (XR) device (e.g., an XR headset), a mobile device, a wearable device, a sensor device, an Internet of Things (IoT) device, an intermediate networking device, or another type of device. In some examples, the transmitting device 102 may be included in a robot or a vehicle. The receiving device 104 may be a computing device such as a mobile device (e.g., a cell phone or tablet computer), a personal computer, a vehicle-based computing device, a wireless base station, a wearable computing device, an intermediate networking device, a dedicated device, an Internet of Things (IoT) device, or another type of device. In some examples, the receiving device 104 may be a device that the user of the transmitting device 102 may have in addition to the transmitting device 102.
[0039]
[0066] The transmitting device 102 and the receiving device 104 can communicate with the base station 106. In some examples, the transmitting device 102 and the receiving device 104 can communicate with the base station 106 using a fifth-generation (5G) wireless communication protocol, a sixth-generation (6G) wireless communication protocol, a WiFi protocol, a Bluetooth protocol, or another type of wireless communication protocol. The base station 106 can transmit data from the network 115 to the transmitting device 102 and the receiving device 104 via wireless downlink channels 108A and 108B (collectively, "wireless downlink channels 108"). The base station 106 can receive data from the transmitting device 102 and the receiving device 104 for transmission to other devices connected to the network 115 via wireless uplink channels 110A and 110B (collectively, "wireless uplink channels 110"). The transmitting device 102 and the receiving device 104 can communicate directly with each other via a wireless sidelink channel 112. In other examples, the transmitting device 102 and the receiving device 104 may communicate via other types of channels.
[0040]
[0067] In the example in Figure 1, the transmitting device 102 includes one or more processors 114, memory 116, a communication interface 118, a video source 120, and a display system 122. The receiving device 104 includes one or more processors 130, memory 132, and a communication interface 134. Processors 114 and 130 may include circuits configured to perform various information processing tasks, including the execution of computer-readable instructions. Processors 114 and 130 may include microprocessors, digital signal processors, and other types of circuits. Memories 116 and 132 may be configured to store data such as computer-readable instructions, video data, and other types of data. Communication interfaces 118 and 134 may be configured to transmit and receive data, for example, via a wireless downlink channel 108, a wireless uplink channel 110, and a wireless sidelink channel 112. In other examples, the transmitting device 102 and the receiving device 104 may communicate via other types of channels. The processor 114 may be coupled to the memory 116, for example, via one or more communication channels. Similarly, the processor 130 may be coupled to the memory 132, for example, via one or more communication channels.
[0041]
[0068] Generally, video source 120 represents a source of video data (e.g., raw, unencoded video data). Video source 120 may include one or more video capture devices such as video cameras, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, video source 120 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video.
[0042]
[0069] In an example where the transmitting device 102 is an XR device that presents MR and AR images to the user, video data from the video source 120 may need to be analyzed so that the display system 122 of the transmitting device 102 can display virtual elements in the correct locations. Processing video data in this way can require considerable computing resources. In other words, a powerful processor and a lot of energy may be used when processing video data. Since the transmitting device 102 may be designed to be worn on the user's head, it may be important to minimize the weight and power consumption of the transmitting device 102 while supporting high-quality, low-latency video.
[0043]
[0070] Furthermore, in some examples, the transmitting device 102 may be an XR headset and may be configured to process a picture of video data to generate virtual element data. The receiving device 104 may be configured to transmit virtual element data (and the transmitting device 102 may be configured to receive virtual element data). The transmitting device 102 may include a display system 122 configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0044]
[0071] Therefore, it may be desirable to offload the processing of video data to a device other than the transmitting device 102, such as the receiving device 104. The receiving device 104 may have more resources than the transmitting device 102, either permanently or temporarily. For example, the receiving device 104 may have a larger battery and a relatively more powerful processor. However, for the receiving device 104 to process the video data, the transmitting device 102 may need to transmit the video data to the receiving device 104 via the wireless sidelink channel 112. Since a very large number of bits may be required to represent unencoded high-quality video data, it would take a considerable amount of time and energy for the transmitting device 102 to transmit unencoded high-quality video data to the receiving device 104. The time required for transmission may undermine the goal of providing low-latency video to the user. The energy required for transmission may undermine the goal of minimizing power consumption. Encoding video data using video coding specifications such as H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), or H.266 / Versatile Video Coding (VVC) can significantly reduce the amount of data required to represent the video data. However, the encoding process itself can introduce its own inherent delays and power consumption requirements.
[0045]
[0072] This disclosure describes techniques that can address these problems. According to the techniques of this disclosure, the transmitting device 102 and the receiving device 104 may use a distributed video coding (DVC) process. The DVC process reduces the amount of coding work performed by the transmitting device 102 and offloads some of the coding work to the receiving device 104. The receiving device 104 may have more resources (e.g., computing power, access to power, etc.) than the transmitting device 102 and therefore may be more capable of performing the coding work. In some examples, the DVC process may be used to load balance computing tasks across devices. For example, the system may determine that it may be more efficient overall for the receiving device 104 to perform a particular video-related computing task than for the transmitting device 102 to perform it.
[0046]
[0073] In addition to the video encoding process, the transmitting device 102 may perform a channel encoding process to prepare the encoded video data for transmission to the receiving device 104. The channel encoding process may generate error correction data for sequences of data within the encoded video data. Typically, the receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, according to the techniques of this disclosure, the transmitting device 102 may send error correction data for some encoded video data, but the error correction data may not send the corresponding encoded video data. The receiving device 104 may estimate one or more subsequent pictures. The receiving device 104 may perform a video encoding process on the subsequent pictures to generate estimated encoded video data. The receiving device may use the estimated encoded video data and the received error correction data to generate error-corrected encoded video data. The receiving device 104 may then decode the error-corrected encoded video data to reconstruct the video data that the transmitting device 102 did not send.
[0047]
[0074] Therefore, in some examples, the receiving device 104 may receive first encoded video data and first error correction data from the transmitting device 102. The first encoded video data may represent one or more blocks of first pictures of video data. The first error correction data may provide error correction information regarding the blocks of first pictures. The receiving device 104 may use the first error correction data to generate first error-corrected encoded video data in order to perform an error correction operation on the first encoded video data. In addition, the receiving device 104 may perform a first reconstruction operation to reconstruct the blocks of first pictures based on the first encoded video data. The first reconstruction operation may be controlled by the values of one or more parameters.
[0048]
[0075] Furthermore, the receiving device 104 may obtain second error correction data from the transmitting device 102. The second error correction data may provide error correction information for one or more blocks of the second picture of the video data. The receiving device 104 may generate prediction data for the second picture. The prediction data for the second picture may include predictions of blocks of the second picture of the video data, based at least in part on blocks of one or more previously reconstructed pictures, such as the first picture. The receiving device 104 may use one or more coding tools to generate prediction data that was not used to generate encoded video data for the second picture. Based on the predictions of blocks of the second picture, the receiving device 104 may generate second encoded video data. The receiving device 104 may use the second error correction data to generate second error-corrected encoded video data in order to perform error correction operations on the second encoded video data. The receiving device 104 may perform a second reconstruction operation in which it reconstructs a second block of picture based on the second error-corrected encoded video data. The second reconstruction operation is controlled by the value of a parameter.
[0049]
[0076] Furthermore, according to one or more techniques of the present disclosure, a receiving device 104 may receive a decimation pattern instruction from the receiving device 104. The receiving device 104 may determine a decimation pattern instruction determined based on a previously reconstructed picture. The decimation pattern instruction may indicate a pattern of non-transmission of encoded video data. For example, a decimation pattern may indicate a pattern of skipping the transmission of encoded video data for the entire picture. In some examples, a decimation pattern may indicate a pattern of skipping the transmission of encoded video data for a specified area within the picture. In some examples where the video data is multi-view video data, a decimation pattern may indicate a pattern of skipping the transmission of encoded video data for the picture from a specified view.
[0050]
[0077] The transmitting device 102 may perform a video encoding process on the pictures of the video data. This video encoding process can compress fewer pictures than "heavy" or more complex compression operations, such as those described in the H.264, H.265, and H.266 video coding standards. In addition to the video encoding process, the transmitting device 102 may perform a channel encoding process to prepare the encoded video data for transmission to the receiving device 104. The channel encoding process can generate error correction data for sequences of data within the encoded video data. Typically, the receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, the receiving device 104 may also use the error correction data to recover information that was not intentionally transmitted to the receiving device. Therefore, the transmitting device 102 may apply a decimation pattern to the second encoded video data to generate decimated video data. The transmitting device 102 may transmit error correction data (generated based on undecimated encoded video data) and decimated video data to the receiving device 104.
[0051]
[0078] The receiving device 104 may obtain first encoded video data and first error correction data from the transmitting device 102. The first encoded data may represent one or more blocks of first pictures of video data. The first error correction data may provide error correction information regarding the blocks of first pictures. The receiving device 104 may use the first error correction data to generate first error-corrected encoded video data in order to perform an error correction operation on the first encoded video data. In addition, the receiving device 104 may perform a first reconstruction operation to reconstruct the blocks of first pictures based on the first encoded video data. The first reconstruction operation may be controlled by the values of one or more parameters.
[0052]
[0079] Furthermore, the receiving device 104 may obtain first error correction data and first encoded video data from the transmitting device 102. The receiving device 104 may apply an error correction process to correct the first encoded video data based on the first error correction data in order to generate first error-corrected encoded video data. The receiving device 104 may also apply a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data. Based on the first set of pictures, the receiving device 104 may determine a decimation pattern indicating a pattern of non-transmission of the encoded video data. The receiving device 104 may transmit a decimation pattern instruction to the transmitting device 102 indicating the determined decimation pattern. The receiving device 104 may receive second error correction data and decimated video data from the transmitting device 102. The decimated video data may include second encoded video data to which the decimation pattern has been applied. The second encoded video data is generated based on a second set of pictures in the video data. The receiving device 104 may apply an error correction process to correct the second encoded video data based on the second error correction data in order to generate the second error-corrected encoded video data. The receiving device 104 may apply a decoding process to reconstruct the second set of pictures based on the second error-corrected encoded video data.
[0053]
[0080] Figure 2 is a block diagram showing exemplary components of a transmitting device and a receiving device according to the technique of the present disclosure. System 200 includes a transmitting device 102 and a receiving device 104. The transmitting device 102 is configured to transmit encoded video data to the receiving device 104. In the example of Figure 2, the transmitting device 102 includes a video encoder 210, a channel encoder 212, and a puncturing unit 214. The receiving device 104 includes a depuncturing unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, and a video encoder 228. In other examples, the transmitting device 102 and the receiving device 104 may include more, fewer, or different units. The processor 114 of the transmitting device 102 (Figure 1) may implement the video encoder 210, the channel encoder 212, and the puncturing unit 214. The processor 130 of the receiving device 104 may implement a depunchling unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, and a video encoder 228. The communication interface 118 (Figure 1) may transmit and receive data on behalf of the transmitting device 102. The communication interface 134 (Figure 1) may transmit and receive data on behalf of the receiving device 104.
[0054]
[0081] The video encoder 210 of the transmitting device 102 may receive video data from a video source (for example, video source 120 (Figure 1)). The video data may include, for example, raw, unencoded video pictures from video source 120. In some examples, the memory of the transmitting device 102 (for example, memory 116 (Figure 1)) may store the video data. The video encoder 210 may perform a video encoding process on the video data to produce encoded video data. The video encoding process may be "limited" in the sense that it is relatively fast and consumes fewer resources than more robust video compression processes such as H.264 / AVC, H.265 / HEVC, or H.266 / VVC. The video encoding process may not reduce the number of bits representing the video data to the same extent as a more robust or full video encoding process.
[0055]
[0082] The video encoder 210 may perform a limited video encoding process in one of several ways. For example, in some cases, the video encoder 210 may perform a prediction process, such as an intra-prediction process, for each picture of the video data to generate prediction data. The video encoder 210 may generate residual data based on the prediction data. For example, the video encoder 210 may subtract a sample of the prediction data from the corresponding sample of the original picture to determine the sample of the residual data. The sample may be a value representing a color value (such as a Y, Cb, or Cr value in the YCbCr color region, or a red, green, or blue value in the RGB color region).
[0056]
[0083] The video encoder 210 may apply a transformation, such as a discrete cosine transform (DCT), to the residual data to produce a transformation block containing transformation coefficients. In addition, the video encoder 210 may quantize the transformation coefficients. The video encoder 210 may apply entropy coding, such as context adaptive binary arithmetic coding (CABAC) coding or exponential Golomb-Rice coding, to the syntax elements representing the quantized transformation coefficients. The encoded video data may contain entropy-coded syntax elements. In some examples, the video encoder 210 directly applies transformation and / or quantization to the video data without first using intra-prediction. In some examples where the video encoder 210 does not apply entropy coding, the encoded video data contains syntax elements representing quantized transformation coefficients, unquantized transformation coefficients, or residual data.
[0057]
[0084] In an example where the video encoder 210 does not use picture-to-picture prediction, fewer memory read requests may be required compared to a more robust video compression process that may need to read data about previously coded pictures from memory. Such memory read requests may be relatively time-intensive and energy-intensive.
[0058]
[0085] In some examples where the video data is multiview video data, the video encoder 210 may perform multiview video coding to generate prediction data. For example, the video encoder 210 may use interview prediction to generate prediction data for blocks of non-anchor pictures (e.g., macroblocks, coding units, etc.). In some cases, interview prediction may involve determining a disparity vector for the block, which indicates the lateral displacement between the block and the corresponding block in the picture of one or more reference views.
[0059]
[0086] The channel encoder 212 of the transmitting device 102 may apply a channel coding process to the encoded video data. The channel coding process prepares the encoded video data for transmission over a wireless communication channel, such as channel 230. Channel 230 may be wireless sidelink channel 112 (Figure 1) or another communication channel. The channel coded video data may include error correction data. The channel encoder 212 may generate error correction data in various ways. For example, the channel encoder 212 may generate error correction data as a convolutional code or a turbo code. The error correction data may help the receiving device 104 determine whether the received coded video data has been altered during transmission over channel 230, and may help the receiving device 104 correct such alterations. A more detailed explanation of channel coding and channel decoding is provided below with respect to Figure 3.
[0060]
[0087] Furthermore, in the example in Figure 2, the puncturing unit 214 of the transmitting device 202 may apply a bit puncturing process to the error correction data in order to generate bit-punctured error correction data. The bit puncturing process can reduce the number of bits in the error correction data. For example, the puncturing unit 214 may perform an operation to remove bits from the error correction data according to a puncturing pattern.
[0061]
[0088] The transmitting device 102 may transmit data such as encoded video data and error-corrected data (e.g., bit-punctured error-corrected data) to the receiving device 104 via channel 230. Channel 230 may introduce noise into the transmitted data. In some examples, channel 230 is a multipath channel, and the data transmitted within channel 230 may change over time. The receiving device 104 may receive the noise-corrected data. The receiving device 104 may store the noise-corrected data at least temporarily in a memory such as memory 132 (Figure 1).
[0062]
[0089] The depunching unit 220 may perform a depunching operation on the received bit-punctured error-corrected data in order to reconstruct the error-corrected data. The depunching operation may replace punctured symbols with neutral values as indicated by the puncture pattern. The depunching operation may generate erase bits that indicate the presence of neutral symbols in the error-corrected data.
[0063]
[0090] The channel decoder 222 may apply a channel decoding process to generate error-corrected encoded video data based on error-corrected data and encoded video data such as encoded video data received from the transmitting device and / or encoded video data generated by the receiving device 104. For example, the channel decoder 222 may modify the values of the bit-encoded video data according to one of various error correction schemes such as low-density parity-check (LDPC) coding or forward error correction (FEC).
[0064]
[0091] The video decoder 224 may perform a video decoding process to reconstruct a picture based on the error-corrected encoded video data. For example, the video decoder 224 may apply an entropy decoding process to the bits of the error-corrected encoded video data to obtain quantized transformation coefficients. The video decoder 224 may apply an inverse quantization operation to the quantized transformation coefficients, apply an inverse transform to the inversely quantized transformation coefficients to generate residual data, generate prediction data to generate prediction data, and use the prediction data and residual data to reconstruct a picture of the video data. The video decoder 224 may generate prediction blocks in the same manner as the video encoder 210.
[0065]
[0092] The picture estimation unit 226 can generate an estimate of the next picture in the video data. For example, the picture estimation unit 226 can extrapolate the next picture from two or more previously reconstructed pictures. For example, in this example, the picture estimation unit 226 can divide the first previously reconstructed picture into blocks. For each block of the first previously reconstructed picture, the picture estimation unit 226 can determine one or more corresponding blocks for that block in one or more further previously reconstructed pictures. The corresponding blocks for a block may be the best available counterpart for that block. Based on one or more corresponding blocks for a block, the picture estimation unit 226 can generate a prediction for the block. The picture estimation unit 226 can use unidirectional or bidirectional prediction to generate the prediction. Thus, by generating a prediction for each block of the next picture, the picture estimation unit 226 can generate an estimate of the next picture. In some examples, the picture estimation unit 226 generates the next picture by applying global motion to the previously reconstructed picture.
[0066]
[0093] In some examples, the picture estimation unit 226 may re-encode the current picture decoded by the video decoder 224. The next picture in the video data may be the picture that follows the picture in the decoding order just decoded by the video decoder 224. In this example, the picture estimation unit 226 may perform intra-prediction or inter-prediction on a block of the current picture. When performing inter-prediction on a block, the picture estimation unit 226 may determine one or more motion vectors for the block. For example, the picture estimation unit 226 may determine that a particular block of the current picture has a motion vector of magnitude m relative to a reference block in a reference picture whose picture order count (POC) distance from the current picture is p1. In this example, the current picture and the next picture may have a POC distance of p2. The picture estimation unit 226 may determine a scaling factor s as p2 / p1. The picture estimation unit 226 may then scale the motion vector of a particular block by s (for example, s * m). The picture estimation unit 226 may determine the location in the next picture indicated by the scaled motion vector and set the sample at the determined location as a sample in a particular block of the current picture. The picture estimation unit 226 may repeat this process for each interpreted block of the current picture.
[0067]
[0094] In some examples, the picture estimation unit 226 may apply one or more filters to the prediction data. For example, the picture estimation unit 226 may apply one or more deblocking filters, smoothing filters, adaptive loop filters, or other types of filters to the prediction data.
[0068]
[0095] The video encoder 228 may perform the same limited video encoding processes as the video encoder 210 on the video data generated by the picture estimation unit 226. For example, the video encoder 228 may perform intra-prediction to generate prediction data. The video encoder 228 may use the prediction data and corresponding blocks of the video data generated by the picture estimation unit 226 to generate residual data. The video encoder 228 may apply transformations (e.g., DCT transformation, DST transformation, etc.) to the residual data to generate transformation coefficients. The video encoder 228 may apply quantization to the transformed coefficients. In addition, the video encoder 228 may apply entropy coding to the syntax elements representing the transformation coefficients.
[0069]
[0096] As briefly mentioned above, the channel decoder 222 can apply the channel decoding process to channel-encoded video data. Figures 3A and 3B provide further information about the channel encoding process performed by the channel encoder 212 and the channel decoding process performed by the channel decoder 222.
[0070]
[0097] Specifically, Figure 3A is a block diagram illustrating an exemplary channel coding process using the technique of the present disclosure. For each picture of the coded video data, the channel encoder 212 of the transmitting device 102 may apply a systematic coding operation, such as a low-density parity check (LDPC) coding operation, to the systematic bits for the picture in order to generate error correction data for the picture. The systematic bits of the picture may include coded video data of the picture generated by the video encoder 210.
[0071]
[0098] In the example in Figure 3A, the error correction data is labeled as “error correction bit”. For picture n, the channel encoder 212 may generate error correction data 300A based on systematic bit 302A. Similarly, for picture n+1, the channel encoder 212 may generate error correction data 300B based on systematic bit 302B. The channel encoder 212 may classify the pictures in the video data as anchor video pictures and non-anchor pictures. The channel encoder 212 may classify the pictures so that anchor pictures are present periodically within the pictures. In some examples, the channel encoder 212 may classify a picture as an anchor picture if the channel decoder 222 receives an indication (for example, from the receiving device 104) that there is an error in the picture. For each of the anchor pictures, the transmitting device 102 may transmit the encoded anchor picture and error correction data for the anchor picture. However, for non-anchor pictures, the transmitting device 102 may only transmit error correction data for non-anchor pictures.
[0072]
[0099] For example, in the example in Figure 3A, picture n may be an anchor picture, and picture n+1 may be a non-anchor picture. Therefore, the transmitting device 102 may transmit systematic bit 302A for picture n, error correction data 300A for picture n, and error correction data 300B for picture n+1, but may not transmit systematic bit 302B for picture n+1.
[0073]
[0100] Figure 3B is a block diagram illustrating an exemplary channel decoding process according to the technique of the present disclosure. As mentioned above, the channel decoder 222 may perform a channel decoding process on the encoded video data in order to reconstruct the encoded video data. When processing an anchor picture (e.g., picture n), the channel decoder 222 processes the systematic bits (SI in Figure 3B) of the anchor picture. n(This is written as) and the error correction data for the anchor picture (in Figure 3B, y n The anchor picture (denoted as) can be obtained from the depuncturing unit 220. The systematic bits of the anchor picture may represent the encoded video data of the anchor picture. The channel decoder 222 may use error correction data for the anchor picture to detect and / or correct errors in the systematic bits of the anchor picture. The video decoder 224 may use the obtained error-corrected encoded video data for the anchor picture to reconstruct the anchor picture. The receiving device 104 may store the reconstructed picture, including the reconstructed anchor picture and the reconstructed non-anchor picture, in the decoded picture buffer 350.
[0074]
[0101] When processing non-anchor pictures, the channel decoder 222 represents the encoded video data of the non-anchor picture (SI n+1 Systematic bits (represented as y in Figure 3B) can be obtained. The video encoder 228 of the receiving device 104 can generate encoded video data for non-anchor pictures based on the video data generated by the picture estimation unit 226. The channel decoder 222 generates error correction data for non-anchor pictures (y in Figure 3B). n+1The channel decoder 222 can then obtain the transmitted error correction data for the non-anchor picture from the depunching unit 220. The channel decoder 222 can then perform the same channel decoding process that the channel decoder 222 applied when processing the anchor picture. Thus, the channel decoder 222 can use the error correction data for the non-anchor picture to detect and / or correct any “errors” in the systematic bits for the non-anchor picture. However, the “errors” in the systematic bits for the non-anchor picture are not due to noise in channel 230 (as in the case of errors in the systematic bits for the anchor picture). Rather, the “errors” in the systematic bits for the non-anchor picture may be due to the difference between the predicted version of the non-anchor picture and the original version of the non-anchor picture. Thus, the channel decoder 222 can use the transmitted error correction data for the non-anchor picture as a mechanism for “correcting” prediction errors.
[0075]
[0102] The video encoder 210 of the transmitting device 102, the video decoder 224 of the receiving device 104, and the video encoder 228 of the receiving device 104 may perform video encoding and video decoding processes based on the values of one or more sets of parameters. In other words, the parameter values may control various aspects of the video encoding process performed by the video encoders 210, 228, and 224. In some examples, the parameters may include one or more of the following: Parameters that indicate the color space (e.g., red-green-blue, Y-Cb-Cr, etc.) Pixel Decimation Parameters A parameter indicating the DCT size Transmitted DCT coefficient A parameter indicating the number of bits per DCT coefficient. Quantization parameters (for example, parameters indicating the quantization scheme, such as linear or Max-Lloyd)
[0076]
[0103] Each of the video encoder 210, video decoder 224, and video encoder 228 may need to use the same parameter value. Accordingly, according to one or more techniques of this disclosure, the transmitting device 102 may transmit the parameter value to the receiving device 104. The receiving device 104 may receive the transmitted parameter value. The video decoder 224 and video encoder 228 may use the parameter value in the video decoding process and the video encoding process.
[0077]
[0104] In some examples, the transmitted parameter values are static or semi-static. For example, in an example where the transmitted parameter value is static, the transmitting device 102 may transmit the parameter value to the receiving device 104 once, and the receiving device 104 may operate with the parameter value for an unspecified period. In an example where the transmitted parameter value is semi-static, the transmitting device 102 may update the parameter value from time to time and retransmit the updated parameter value to the receiving device 104.
[0078]
[0105] The transmitting device 102 may transmit the parameter value in one of several ways. For example, in some cases, the transmitting device 102 may transmit the parameter value to the receiving device 104 using an Uplink Control Information (UCI) / Media Access Control-Control Element (MAC-CE) message, a Radio Resource Control (RRC) message, or another type of message.
[0079]
[0106] In an example where the transmission device 102 transmits the value of a parameter to the reception device 104, the video encoder 210 of the transmission device 102 may divide each of the color components (e.g., R, G, and B components, Y, Cb, Cr components) of a picture of video data into equally sized (MxM) blocks. Examples of such blocks may include macroblocks (MBs) and largest coding units (LCUs). The video encoder 210 of the transmission device 102 may calculate a transform (e.g., 2D-DCT) for each of the blocks, resulting in M 2 transform coefficients. The video encoder 210 may assign an ordering to the transform coefficients of the blocks. For example, the video encoder 210 may order the transform coefficients of the blocks according to a zigzag scan order that starts from the most important transform coefficients (e.g., the lowest frequencies) and ends at the least important transform coefficients (e.g., the highest frequencies). The video encoder 210 may select the first N c transform coefficients, where N c is a parameter value indicating the amount of transform coefficients to be transmitted. The video encoder 210 may discard the transform coefficients that were not selected.
[0080]
[0107] In addition, the video encoder 210 may quantize the selected transform coefficients. For example, if the selected transform coefficients have an index i in the range from 0 to N c -1, the parameter may include bitwidth parameters (e.g., B i , i = 0, 1,..., N c -1) corresponding to different index values. For each of the selected transform coefficients d i , the video encoder 210 may quantize the selected transform coefficient d i using the following formula.
[0081]
Equation
[0082] In the above formula, c i is the transform coefficient di This is a quantized version of where α is the scaling constant, and B i is the bit width parameter for index i, and round is the function that rounds to the nearest integer. Therefore, in the example where B0 is 8, B1 is 4, and B2 is 4, the quantized conversion coefficients could be, for example, c0=00100011, c1=0110, c2=1001, etc.
[0083]
[0108] The video encoder 228 of the receiving device 104 may generate prediction data for a picture, generate residual data based on the prediction data, and apply one or more transformations to the residual data to generate a transformation block containing transformation coefficients. The video encoder 228 may need to use the same bit width parameters as the video encoder 210 so that the channel decoder 222 can correctly associate specific systematic bits with the corresponding error correction data received from the depunchaging unit 220.
[0084]
[0109] In some examples of this disclosure, the receiving device 104 may determine the value of one or more of the parameters without the transmitting device 102 transmitting the values of these parameters to the receiving device 104. Examples in which the receiving device 104 determines the value of one or more of the parameters without the transmitting device 102 transmitting the values to the receiving device 104 may enable a better compression-distortion trade-off with lower control signaling overhead. For example, the receiving device 104 may determine the number of DCT coefficients (N c ) to, N c The value of can be determined without transmitting it to the receiving device 104. For example, in this example, the receiving device 104 can determine N to achieve the desired peak signal-to-noise ratio (P-SNR). c The minimum value of can be determined. In other words, the receiving device 104 can determine the N that achieves the desired P-SNR. c The minimum value can be determined as follows:
[0085]
number
[0086] In the above formula, c i This is a transformation coefficient with index i (for example, a DCT coefficient), and the SNR d This is the desired P-SNR, and M 2 -1 is the maximum number of conversion factors. In some examples, the receiving device 104 calculates the number of conversion factors (N) once for each picture (or other segment) based on robust prediction of the picture. c ) can be evaluated. Periodic resetting of parameter values may be applied to avoid error propagation.
[0087]
[0110] In another example, instead of receiving the number of quantized bits from the transmitting device 102, the receiving device 104 may determine the number of quantized bits based on the predicted picture. For example, in this example, the receiving device 104 may calculate the probability distributions for the quantized and unquantized coefficients.
[0088]
number
[0089] In the above formula,
[0090]
number
[0091] is the probability distribution of the quantized coefficients, and p i c is the probability distribution of the unquantized coefficients, i The conversion coefficient d is i It is a quantized version of [the original].
[0092]
[0111] The receiving device 104 determines the number of quantized bits B based on the entropy ratio of the quantized coefficients to the unquantized coefficients, as follows: i It is possible to determine this.
[0093]
number
[0094] Therefore, the receiving device 104, based on the picture prediction generated by the picture estimation unit 226, once for each picture (or segment), B i It can be evaluated.
[0095]
[0112] Figure 4 is a flowchart illustrating exemplary operation of a transmitting device 102 using the techniques of the present disclosure. In the example of Figure 4, the video encoder 210 of the transmitting device 102 may acquire video data (400). For example, the video encoder 210 may acquire video data from a video source 120. In addition, the video encoder 210 may perform video coding on the video data to generate encoded video data (402). For example, the video encoder 210 may apply intra-prediction to generate prediction data, generate residual data based on the prediction data and the original video data, and apply a transformation (e.g., DCT) to the block of residual data to generate a transformation block. The video encoder 210 may quantize the transformation coefficients of the transformation block. Furthermore, in some examples, the video encoder 210 may apply entropy coding to the syntax elements representing the quantized transformation coefficients. In some examples, the video encoder 210 may implement a reconstruction loop to reconstruct the residual data, which may apply entropy decoding, inverse quantization, and one or more inverse transformations. The video encoder 210 may apply the prediction data and the reconstructed residual data to reconstruct the video data. In some examples, the video encoder 210 applies one or more filters, such as a deblocking filter, an adaptive loop filter, or a sample-adaptive offset filter, to the reconstructed video data. The video encoder 210 may use the reconstructed video data as reference data for intra-prediction.
[0096]
[0113] The video encoder 210 may perform a video encoding process based on the values of one or more parameters. For example, the video encoder 210 may quantize conversion coefficients according to specific quantization parameters, use a specific color space, and so on.
[0097]
[0114] The channel encoder 212 of the transmitting device 102 may perform channel coding on the encoded video data to generate error correction data (404). The transmitting device 102 may transmit the encoded video data and error correction data to the receiving device 104, for example, via channel 230 (406). In some examples, the transmitting device 102 may selectively transmit portions of the encoded video data and transmit other portions of the encoded video data. For example, the transmitting device 102 may transmit the encoded video data of some pictures and not the encoded video data of other pictures. In another example, the transmitting device 102 may transmit the encoded video data of some transformation blocks of a picture and not transmit the encoded video data of other transformation blocks of a picture. In some examples, the transmitting device 102 may transmit the higher bits of a certain amount of the transformation coefficients and not transmit the lower bits of the transformation coefficients.
[0098]
[0115] In some examples, the transmitting device 102 may also transmit values for one or more parameters to the receiving device 104. The parameter values may control how the receiving device 104 reconstructs the video data. For example, the parameters may include a transformation size parameter indicating the size of the transformation blocks in the encoded video data generated by the video encoder 210. In this example, the receiving device 104 may need to interpret the received encoded video data according to the same transformation block size in order to properly reconstruct the video data. In other examples, the parameters may include parameters indicating the amount of transformation coefficients, bit width parameters, and so on.
[0099]
[0116] Therefore, in the example of Figure 4, the transmitting device 102 can acquire video data from a video source. Based on a set of parameters, the transmitting device 102 can generate encoded video data for a first picture of the video data and encoded video data for a second picture of the video data. The transmitting device 102 can perform channel coding on the encoded video data of the first picture and the encoded video data of the second picture in order to generate error correction data for the first picture and error correction data for the second picture. The transmitting device 102 can transmit the encoded video data of the first picture, the error correction data for the first picture, and the error correction data for the second picture. In some examples, the transmitting device 102 can transmit parameter values to the receiving device 104.
[0100]
[0117] Figure 5 is a flowchart illustrating exemplary operation of a receiving device 104 using the technique of the present disclosure. In the example of Figure 5, the receiving device 104 may receive first encoded video data and first error correction data from the transmitting device 102 (500). The first encoded data represents one or more blocks of first pictures of video data. The first error correction data may provide error correction information relating to the blocks of first pictures.
[0101]
[0118] The receiving device 104 may use the first error correction data to generate first error-corrected encoded video data in order to perform an error correction operation on the first encoded video data (502). For example, the channel decoder 222 of the receiving device 104 may use the first error correction data to perform low-density parity check (LDPC) coding on the first encoded video data. In other examples, the channel decoder 222 may use the error correction data in other error correction algorithms such as forward error correction (FEC) or turbo coding. Performing an error correction operation on the first encoded video data may remove errors caused by noise in channel 230.
[0102]
[0119] The video decoder 224 of the receiving device 104 may perform a first reconstruction operation (504) to reconstruct a block of a first picture based on a first error-corrected encoded video data. The first reconstruction operation is controlled by the values of one or more parameters. For example, the video decoder 224 may perform an inverse transform on the transformed block of the first error-corrected encoded video data to obtain residual data. In addition, in this example, the video decoder 224 may generate prediction data using, for example, intra-prediction. In this example, the video decoder 224 may use the prediction data and residual data to reconstruct a block of a first picture.
[0103]
[0120] In some examples, video encoders 210 and 228 may use quantization parameters to generate encoded video data in order to quantize the transformation coefficients generated based on prediction data for the picture. When performing a reconstruction operation, video decoder 224 may use quantization parameters to dequantize the transformation coefficients of the error-corrected encoded video data. In some examples, transmitting device 102 and / or receiving device 104 may calculate quantization parameters based on the entropy ratio of quantized to unquantized transformation coefficients, for example, as described above.
[0104]
[0121] In some examples, the parameters include a transformation size parameter. As part of generating encoded video data, video encoders 210 and 228 may apply a forward transformation to the picture's sample region data (e.g., predicted sample data or residual data) having a transformation size indicated by the transformation size parameter. As part of performing a reconstruction operation, video decoder 224 may apply an inverse transformation to the transformation coefficients of the error-corrected encoded video data having a transformation size indicated by the transformation size parameter.
[0105]
[0122] In some examples, the parameters include a parameter indicating the amount of conversion coefficients. As part of generating encoded video data, video encoders 210 and 228 may include in the encoded video data a set of conversion coefficients, each containing the indicated amount of conversion coefficients. When performing a reconstruction operation, video decoder 224 may analyze the error-corrected encoded video data to obtain a set of conversion coefficients, each containing the indicated amount of conversion coefficients. Furthermore, in some examples, receiving device 104 may receive encoded video data and error-corrected data from transmitting device via a communication channel, and receiving device 104 may apply an optimization process to determine the number of conversion coefficients based on the signal-to-noise ratio of the data transmitted over the communication channel, for example, as described above.
[0106]
[0123] In some examples, the parameters include bit width parameters for multiple index values. Performing a reconstruction operation for each of the multiple index values may involve analyzing a first set of bits from the error-corrected encoded video data. The first set of bits may represent conversion coefficients having each index value, and the amount of bits in the first set of bits is equal to the bit width indicated by the bit width parameter for each index value. As part of generating the encoded video data, video encoders 210 and 228 may include a second set of bits in the encoded video data. The second set of bits may represent conversion coefficients having each index value, and the amount of bits in the second set of bits is equal to the bit width indicated by the bit width parameter for each index value. Video decoder 224 may analyze a third set of bits from the error-corrected encoded video data. The third set of bits may represent conversion coefficients having each index value, and the amount of bits in the third set of bits is equal to the bit width indicated by the bit width parameter for each index value.
[0107]
[0124] Other parameters may include one or more of the following: color space, conversion size, quantization parameters, the number of conversion coefficients in the first encoded video data, or the number of bits per conversion coefficient in the first encoded video data.
[0108]
[0125] The receiving device 104 may obtain second error correction data from the transmitting device 102 (506). The second error correction data provides error correction information relating to one or more blocks of a second picture of the video data.
[0109]
[0126] In addition, the picture estimation unit 226 of the receiving device 104 may estimate a second picture based on one or more previously reconstructed pictures, such as the first picture (508). The estimated second picture includes predictions of blocks of the second picture in video data based at least in part on blocks of the first picture. For example, the picture estimation unit 226 may generate prediction data using interpretation, a combination of intrapretation and interpretation, or other video coding tools, as described elsewhere in this disclosure.
[0110]
[0127] The video encoder 228 of the receiving device 104 may generate second encoded video data based on the estimated second picture (510). For example, the video encoder 228 of the receiving device 104 may generate residual data based on prediction data. For example, the video encoder 228 may perform an intra-prediction to generate second prediction data based on the estimated second picture. The video encoder 228 may then generate residual data by subtracting the second prediction data from the prediction data generated by the picture estimation unit 226. The video encoder 228 may then generate a transformation block by applying one or more forward transformations to the residual data. The video encoder 228 may perform the same process as the video encoder 210 of the transmitting device 102 and therefore may need to use the same parameters as the video encoder 210.
[0111]
[0128] The channel decoder 222 of the receiving device 104 may perform a channel decoding process to generate second error-corrected encoded video data based on second error-corrected data and second encoded video data (512). The channel decoder 222 of the receiving device 104 may perform the same process as the channel decoder 222 performed when generating the first error-corrected encoded video data to generate the second error-corrected encoded video data.
[0112]
[0129] The video decoder 224 of the receiving device 104 may perform a second video decoding process to reconstruct a second block of the picture based on the second error-corrected encoded video data (514). The second reconstruction operation is controlled by the value of a parameter. The video decoder 224 may perform the second reconstruction operation in the same manner as the first reconstruction operation. In this way, the receiving device 104 can reconstruct the video data of a picture (or block) without receiving all of the encoded video data of each picture (or block).
[0113]
[0130] As described above, the puncturing unit 214 of the transmitting device 102 may perform a bit puncturing operation on the error correction data generated by the channel encoder 212. Bit puncturing involves the transmitting device 102 selectively discarding some of the error correction data before transmitting the error correction data. The discarded bits are usually not the most important for performing error correction. The depuncturing unit 220 of the receiving device 104 may perform an inverse bit puncturing operation (i.e., a bit depuncturing operation) which reverses the bit puncturing operation performed by the puncturing unit 214. The puncturing unit 214 may perform the bit puncturing operation according to one or more sets of puncturing parameters. In different examples, the puncturing parameters may be predetermined, static, or semi-static.
[0114]
[0131] According to one or more techniques of this disclosure, the transmitting device 102 may perform a decimation procedure that can reduce memory bandwidth and enhance compression. For example, the video encoder 210 of the transmitting device 102 may divide the picture of video data into a grid of blocks (e.g., MBs, LCUs, etc.) and generate a transformed block for each block. The transmitting device 102 may need to store each transformed block intended for transmission to the receiving device 104 in memory (e.g., memory 116). The transmitting device 102 may then retrieve the stored transformed blocks from memory for channel coding and ultimately for transmission. These writing to and reading from memory can increase time and energy requirements. These time and energy requirements may be directly related to the amount of data written and read. Therefore, it may be advantageous to reduce the amount of data written to and read from memory.
[0115]
[0132] Performing a decimation process can reduce the amount of data in a translated block that is written to and read from memory. Performing a decimation process can also reduce the amount of data that is sent by the transmitting device 102 to the receiving device 104. In some examples, when the video encoder 210 is encoding the current block of the current picture, the video encoder 210 may generate a translated block for the current block. In addition, the channel encoder 212 of the transmitting device 102 may determine, based on a decimation pattern, whether the current block is subject to decimation. If the current block is subject to decimation (i.e., the translated block is a "non-anchor translated block"), the channel encoder 212 may reduce the number of bits in the non-anchor translated block before storing the translated block in memory. If the current block is not subject to decimation (i.e., the translated block is an "anchor block"), the channel encoder 212 does not reduce the number of bits in the anchor block. The channel encoder 212 may perform a decimation process after generating error correction data. Therefore, the error correction data generated by the channel encoder 212 for non-anchor conversion blocks (and potentially sent to the receiving device 104) may be based on the full set of bits of the conversion block, rather than a reduced number of bits.
[0116]
[0133] Figure 6 is a conceptual diagram showing an exemplary decimation pattern 600 using the technique of the present disclosure. The example in Figure 6 shows a grid of DCT blocks. A DCT block is a block of transformation coefficients generated by applying a DCT transformation to video data, such as residual data or sample data. In other examples, a DCT block may be a transformation block generated using other types of transformations. In Figure 6, the "X" marks in the decimation pattern 600 indicate the DCT blocks (i.e., non-anchor transformation blocks) that are subject to decimation. Thus, in the example in Figure 6, the decimation pattern 600 decimates the DCT blocks by half horizontally and vertically. In some examples, which may be referred to as “complete” decimation herein, the video encoder 210 may reduce the number of bits in the non-anchor transformation blocks to zero.
[0117]
[0134] Therefore, in some examples, the decimation pattern defines the pattern of anchored and unanchored converted blocks in the picture. The receiving device 104 may receive systematic bits of the anchored converted blocks, but may not receive systematic bits of the unanchored converted blocks. The systematic bits of the anchored converted blocks may represent the conversion coefficients within the anchored converted blocks. The systematic bits of the unanchored converted blocks may represent a reduced-bit-depth version of the original conversion coefficients within the unanchored converted blocks. The error correction data may include error correction data for the anchored converted blocks and error correction data for the unanchored converted blocks. The error correction data for the unanchored converted blocks is based on the original conversion coefficients within the unanchored converted blocks. As part of generating error-corrected encoded video data, the channel decoder 222 may use the error correction data for the anchored converted blocks to perform error correction on the systematic bits of the anchored converted blocks. The channel decoder 222 may also use the error correction data for the unanchored converted blocks to perform error correction on the portion of the encoded video data corresponding to the unanchored converted blocks. In some examples, the receiving device 104 may determine a decimation pattern and transmit the decimation pattern to the transmitting device 102.
[0118]
[0135] In some examples, the transmitting device 102 stores encoded bits (encoded video data and error correction data) in a cyclic buffer. The transmitting device 102 uses two parameters to select which bits in the cyclic buffer to transmit. The first parameter is the starting position, and the second parameter indicates the number of consecutive bits to transmit. The starting position can have various values only if it is instructed to support selective transmission and non-transmission of system bits. The starting position may be selected to skip the transmission of certain systematic bits (i.e., bits of encoded video data) without skipping the transmission of error correction data. Thus, decimation of non-anchor conversion blocks can be achieved simply by manipulating the first and second parameters so that the transmitting device 102 does not transmit bits of the non-anchor conversion block.
[0119]
[0136] In some examples, the channel encoder 212 applies a hybrid decimation technique that does not reduce the number of bits in any of the targeted "non-anchor" transformation blocks to zero, but reduces the number of bits in the transformation coefficients within the non-anchor transformation blocks. For example, in the example in Figure 6, the channel encoder 212 may reduce the number of bits in each transformation coefficient within the DCT block marked "X" by a predetermined number (e.g., 2, 4, 5, etc.). The channel encoder 212 does not reduce the number of bits in transformation coefficients that are not targeted by the decimation pattern.
[0120]
[0137] The channel decoder 222 of the receiving device 104 may receive the remaining reduced bits of the non-anchor conversion block and error correction data for the non-anchor conversion block. As part of the channel decoding process, the channel decoder 222 may perform an error correction process to restore the removed bits of the non-anchor conversion block using the error correction data for the non-anchor conversion block. This error correction process may be the same error correction process that the channel decoder 222 uses to correct errors caused by noise in channel 230. In summary, the channel encoder 212 generates error correction data because it will be needed to correct the unavoidable noise in channel 230, but this same error correction data is used to restore bits as if the noise in channel 230 had occurred in such a way that it corrupted the least significant bit of a specific conversion coefficient in a specific conversion block in a particular picture. Thus, the number of bits transmitted in channel 230 may be substantially reduced.
[0121]
[0138] In some examples, the channel encoder 212 may generate a correlation matrix based on the picture's transformation block set before performing any decimation process on any non-anchor transformation block in the set of transformation blocks. The correlation matrix contains values indicating the level of correlation between transformation coefficients at their corresponding locations within the transformation blocks. For example, the correlation matrix may contain correlation values for the DC transformation coefficients (i.e., the top-left transformation coefficients) of the set of transformation blocks. If the differences between the DC transformation coefficients are relatively small, the correlation values for the DC transformation coefficients may be relatively high. Conversely, if the differences between the DC transformation coefficients are relatively large, the correlation values for the DC transformation coefficients may be relatively small. Each correlation value can be between 0 and 1.
[0122]
[0139] In some examples, the channel encoder 212 may calculate the correlation value for the DC conversion coefficient using the following formula:
[0123]
number
[0124] In equation 5 above, l represents the interval between transformation blocks containing DC coefficients, and N represents the number of transformation blocks on which the calculation is performed. The function y is a transformation coefficient. The line above y represents its conjugate. The interval may be 1 when each consecutive transformation block is used, the interval may be 2 when alternating transformation blocks are used, and so on. The channel encoder 212 may calculate correlation values for the corresponding AC transformation coefficients (i.e., non-DC transformation coefficients) in the same manner. In this disclosure, the corresponding transformation coefficients occupy the same location within the transformation block. Therefore, by calculating the autocorrelation value for each transformation coefficient in the transformation block, the channel encoder 212 may generate a correlation matrix for the transformation block. Since the channel encoder 212 uses transformation coefficients from different transformation blocks when calculating the correlation values, the channel encoder 212 may repeat the process of generating a correlation matrix for each transformation block.
[0125]
[0140] The transmitting device 102 may transmit the correlation matrix along with the encoded video data and error correction data to the receiving device 104. The channel decoder 222 of the receiving device 104 may use the correlation matrix as part of the process to restore the non-anchor conversion blocks to their original bit widths. For example, continuing with the DC conversion coefficient example, after applying the error correction process, the channel decoder 222 may obtain the values of the non-anchor DC conversion coefficients (i.e., the DC conversion coefficients in the non-anchor conversion block).
[0126]
[0141] The error correction process may use correlation values to estimate non-anchor transformation coefficients. For example, if all even transformation blocks are anchor transformation blocks and non-even transformation blocks are non-anchor transformation blocks, the channel decoder 222 may estimate the value of the non-anchor transformation coefficient in transformation block n by the following formula.
[0127]
number
[0128] In equation (6) above, R yy [1] shows the correlation values in the correlation matrix for the transformation block with index 1 (i.e., non-even, non-anchor transformation blocks), y[n+1] shows the corresponding transformation coefficients for the anchor transformation block with index n+1, R yy [2] shows the correlation value in the correlation matrix for the transformation block with index 3, and y[n+3] shows that the corresponding transformation coefficient in the anchor transformation block is at index n+3, and so on. The number of transformation blocks used in equation (6) may be configurable. In this way, the estimated value of the non-anchor transformation coefficient can be considered as a weighted average of the corresponding transformation coefficients in the anchor blocks, weighted based on the correlation value at the corresponding location in the correlation matrix. In other words, the channel decoder 222 may interpolate the values of the non-anchor transformation coefficients according to the correlation matrix. The channel decoder 222 may output the calculated values of the transformation coefficients to the video decoder 224. The channel decoder 222 may perform this process for other transformation coefficients. Using the correlation matrix in this way may improve the quality of the reconstructed video data.
[0129]
[0142] In some examples, the video encoder 210 may reduce the bit width of each transformation coefficient in the targeted transformation block by the same amount. In other examples, the video encoder 210 may reduce the bit width of different transformation coefficients in the targeted transformation block by different amounts. In some examples, the amount by which the video encoder 210 reduces the bit width of a transformation coefficient is related to the distance of the transformation coefficient from the anchor transformation block. The anchor transformation block is a transformation block that is not targeted by the decimation pattern.
[0130]
[0143] Figure 7 is a flowchart illustrating exemplary operation of a transmitting device 102 for hybrid decimation of transform blocks using the technique of the present disclosure. In the example of Figure 7, the transmitting device 102 may acquire video data from a video source 120 (Figure 7) (700). The video encoder 210 of the transmitting device 102 may generate transform blocks based on the video data (702). For example, the video encoder 210 may generate a prediction block by performing an intra-prediction on a block of pictures in the video data. The video encoder 210 may use the prediction block to generate residual data. The video encoder 210 may generate a transform block by applying a transform, such as DCT, DST, or other transform, to the residual data. In another example, the video encoder 210 may generate a transform block by directly applying a transform to a block of video data.
[0131]
[0144] The channel encoder 212 may then determine, based on the decimation pattern, which of the transformation blocks are anchor transformation blocks (704). For example, in the example where the channel encoder 212 uses the decimation pattern 600 of Figure 6, the video encoder 210 may determine that every other transformation block in both the horizontal and vertical directions is an anchor transformation block. In other examples, the channel encoder 212 may use other decimation patterns to determine which of the transformation blocks are anchor transformation blocks. The channel encoder 212 may store the anchor transformation blocks in the memory of the transmitting device 102 (for example, memory 116 (Figure 1)) (706).
[0132]
[0145] The channel encoder 212 can compute a correlation matrix for a set of transformation blocks (708). Each set of transformation blocks includes one or more anchored transformation blocks and one or more unanchored transformation blocks. For example, each set of transformation blocks may correspond to a different row of transformation blocks in Figure 6. In another example, each set of transformation blocks may correspond to a group of 2x2 transformation blocks. The number of values in the correlation matrix for a set of transformation blocks is individually the same as the number of transformation coefficients in each of the transformation blocks. Each value in the correlation matrix corresponds to a different position in the transformation coefficient block. For example, the value at position (0,0) in the correlation matrix corresponds to the transformation coefficient at position (0,0) in each transformation coefficient block in the set of transformation blocks, and so on. The transformation coefficient at position (0,0) in the transformation coefficient block may be called the DC coefficient, and all other transformation coefficients may be called the AC coefficients.
[0133]
[0146] In addition, the channel encoder 212 may generate a bit-reduced non-anchor transformation matrix (710). A non-anchor transformation matrix is a transformation matrix other than the anchor transformation matrix. For example, referring to Figure 6, the transformation matrix marked with X may be a non-anchor transformation matrix. The transformation coefficients in the bit-reduced non-anchor transformation matrix may contain fewer bits than the transformation coefficients in the original version of the non-anchor transformation matrix. The transmitting device 102 may then transmit the anchor transformation block, the non-anchor transformation block, the bit-reduced value, the correlation matrix, and the error correction data (712).
[0134]
[0147] The channel encoder 212 may reduce the number of bits in the non-anchor conversion coefficients in one of several ways. For example, the channel encoder 212 may determine a bit reduction value for each conversion coefficient in the non-anchor conversion coefficient block. In this example, to calculate the bit reduction value, the channel encoder 212 may calculate the interpolated value of the conversion coefficients according to the correlation matrix. The channel encoder 212 may calculate the interpolated value using equation (6) above. The channel encoder 212 may then subtract the interpolated value of the conversion coefficient from the original value of the conversion coefficient to calculate a first strain value. The channel encoder 212 may then reduce the number of bits in the original value of the conversion coefficient by 1. The channel encoder 212 may then subtract the interpolated value from the bit-reduced original value of the conversion coefficient to calculate a second strain value. Based on the first and second strain values, the channel encoder 212 may determine whether the second strain value is acceptable. For example, the channel encoder 212 may calculate the mean or greatest squares error from the interpolated value. The channel encoder 212 can determine if the second strain value is acceptable by comparing it with a predetermined threshold value.
[0135]
[0148] If the second distortion value is acceptable, the channel encoder 212 may reduce the number of bits in the original value of the conversion coefficient and repeat this process. If the second distortion value is unacceptable, the channel encoder 212 may increase the number of bits in the original value of the conversion coefficient. The number of bits obtained by reducing the original value of the conversion coefficient by that amount is the bit reduction value.
[0136]
[0149] In some examples, the channel encoder 212 may determine a bit reduction value for an entire block of non-anchor conversion coefficients. In this example, to calculate the bit reduction value for the block of conversion coefficients, the channel encoder 212 may calculate the interpolated value of each conversion coefficient according to the correlation matrix, for example, as described above. The video encoder 210 may then subtract the interpolated value of the conversion coefficient from the original value of the conversion coefficient and use the resulting difference to calculate a first distortion value. For example, the video encoder 210 may calculate the first distortion value as the mean squared error. The video encoder 210 may then reduce the number of bits in each original value of the conversion coefficient by 1. The video encoder 210 may subtract the interpolated value from the bit-reduction original value of the conversion coefficient and use the resulting value to calculate a second distortion value (for example, using the mean squared error). Based on the first and second distortion values, the channel encoder 212 may determine whether the second distortion value is acceptable. If the second distortion value is acceptable, the video encoder 210 may reduce the number of bits in the original value of the conversion coefficient again and repeat this process. If the second distortion value is unacceptable, the video encoder 210 may increase the number of bits in the original value of the conversion coefficient. The number of bits obtained by reducing the original value of the conversion coefficient by that amount is the bit reduction value.
[0137]
[0150] Therefore, in some examples, the transmitting device 102 may acquire video data from a video source. The transmitting device 102 may generate transformation blocks based on the video data. The transmitting device 102 may determine which of the transformation blocks are anchor transformation blocks. The transmitting device 102 may calculate a correlation matrix for the set of transformation blocks. In addition, the transmitting device 102 may generate a bit-reduced non-anchor transformation matrix. The transmitting device 102 may send the anchor transformation blocks, the non-anchor transformation blocks, and the correlation matrix to the receiving device. In some examples, the transmitting device 102 may receive instructions for a decimation pattern from the receiving device 104.
[0138]
[0151] Figure 8 is a flowchart illustrating exemplary operation of a receiving device 104 for hybrid decimation of a transform block using the technique of the present disclosure. In the example of Figure 8, the receiving device 104 may receive an anchored transform block, a non-anchored transform block, one or more bit reduction values, and a correlation matrix for the non-anchored block (800).
[0139]
[0152] Furthermore, in the example in Figure 8, the channel decoder 222 of the receiving device 104 can calculate an interpolated value of the current non-anchor transformation coefficient (802). The current non-anchor transformation coefficient is one of the transformation coefficients in the non-anchor transformation block. The channel decoder 222 can calculate an interpolated value of the current non-anchor transformation coefficient based on a correlation matrix for the non-anchor block. The channel decoder 222 can calculate an interpolated value of the current non-anchor transformation coefficient in one of several ways. For example, in some examples, the channel decoder 222 may apply a machine learning model that takes one or more non-anchor transformation coefficients (including the current non-anchor transformation coefficient), one or more anchor transformation coefficients, and a correlation matrix for the non-anchor transformation coefficient block as input. In this example, the machine learning model may output an interpolated value of the current non-anchor transformation coefficient. In this example, the machine learning model may be implemented as a neural network model, a support vector machine, a regression model, or another type of machine learning model.
[0140]
[0153] In another example, the correlation matrix may include values showing the correlation between the current non-anchor transformation coefficients and each corresponding anchor transformation coefficient in one or more blocks of anchor transformation coefficients. The video decoder 224 may calculate the interpolated values of the current non-anchor transformation coefficients as follows:
[0141]
number
[0142] In the above equation, t int is the current non-anchor transformation coefficient, and ai c is the anchor transformation coefficient, i The current non-anchor transformation coefficient and a i This is a correlation value that shows the correlation between and , where n is the number of anchor transformation coefficients that are derived from the current non-anchor transformation coefficient. In this example, from c0 to c n The sum of the values can be 1.
[0143]
[0154] In addition, the video decoder 224 may calculate the reconstructed value of the non-anchor conversion coefficient (804). The video decoder 224 may calculate the reconstructed value of the non-anchor conversion coefficient based on the interpolated value of the current non-anchor conversion coefficient and the transmitted value of the non-anchor conversion coefficient. The transmitted value of the non-anchor conversion coefficient is included in the received non-anchor conversion block. In some examples, the video decoder 224 calculates the reconstructed value of the non-anchor conversion coefficient as the average of the interpolated value of the current non-anchor conversion coefficient and the transmitted value of the non-anchor conversion coefficient.
[0144]
[0155] The video decoder 224 may determine whether there are any remaining non-anchor conversion coefficients in the non-anchor conversion block (806). If there are one or more remaining non-anchor conversion coefficients in the non-anchor conversion block (the "yes" branch of 806), the video decoder 224 may repeat steps 802-806 with another of the non-anchor conversion coefficients. The video decoder 224 may continue doing so until there are no more remaining non-anchor conversion coefficients (the "no" branch of 806). In this way, the video decoder 224 may calculate the reconstructed value for each of the non-anchor conversion coefficients.
[0145]
[0156] In this way, the receiving device 104 can receive systematic bits of the anchor transformation block, systematic bits of the non-anchor transformation block, and a correlation matrix. The systematic bits of the anchor transformation block may represent the transformation coefficients within the anchor transformation block. The systematic bits of the non-anchor transformation block may represent a reduced-bit-depth version of the original transformation coefficients within the non-anchor transformation block. As part of performing a reconstruction operation, the receiving device 104 may calculate an interpolated value of each non-anchor transformation coefficient in the non-anchor transformation block based on the correlation matrix and the corresponding anchor transformation coefficient. The receiving device 104 may calculate a reconstructed value of the non-anchor transformation coefficient based on the interpolated value of the non-anchor transformation coefficient and the value of the non-anchor transformation coefficient in the error-corrected encoded video data.
[0146]
[0157] In some examples, the receiving device 104 may adaptively select a decimation pattern to be used to reduce or remove bits from a particular conversion block. In such examples, the receiving device 104 may communicate to the transmitting device 102 to return the selected decimation pattern. The transmitting device 102 may then use the selected decimation pattern in one or more pictures of the video data.
[0147]
[0158] Figure 9 is a conceptual diagram showing an exemplary decimation pattern 900 adaptively selected by the receiving device 104 using one or more techniques of the present disclosure. In contrast to the decimation pattern 600 in Figure 6, the non-anchor blocks in the decimation pattern 900 do not necessarily occur at regular intervals or periods.
[0148]
[0159] The receiving device 104 may determine a decimation pattern based on information about the previous picture in the video data. The previous picture may be decimated or undecimated. In some examples, the receiving device 104 may send a request to the transmitting device 102 for the undecimated version of the picture. The transmitting device 102 may, upon request, send the undecimated version of the picture to the receiving device 104. After receiving the undecimated version of the picture, the receiving device 104 may determine a decimation pattern based on the undecimated version of the picture. For example, the receiving device 104 may perform a rate-distortion optimization process that evaluates several possible decimation patterns to determine which of the decimation patterns results in the best combination of bitrate and distortion.
[0149]
[0160] In some examples, the transmitting device 102 and the receiving device 104 may continue to use the selected decimation pattern for a predetermined number of pictures, after which the transmitting device 102 and / or the receiving device 104 may adaptively select a different decimation pattern. In some examples, the transmitting device 102 may send a message to the receiving device 104 requesting that the transmitting device 102 select a different decimation pattern. In some examples, the receiving device 104 may determine that an event or condition has occurred that would favor the selection of a different decimation pattern. For example, the receiving device 104 may determine that it would be favorable to select a different decimation pattern when it determines that a scene change has occurred, when it determines that motion in the video data has exceeded one or more thresholds, or when it determines that other characteristics of the video data have changed.
[0150]
[0161] The receiving device 104 may signal the selected decimation pattern to the transmitting device 102 in one of several ways. For example, the receiving device 104 may signal the selected decimation pattern to the transmitting device 102 by indicating the difference from an existing decimation pattern, such as the decimation pattern currently in use. For example, in this example, the receiving device 104 may indicate the selected decimation pattern to the transmitting device 102 by specifying a change in downsampling or upsampling along a particular axis, specifying a change in a particular area of the picture, deleting a particular transformation block, or enabling a particular transformation block.
[0151]
[0162] In some cases, there may be a predefined mapping of index values to predetermined decimation patterns. In such cases, the receiving device 104 may select a decimation pattern from the predetermined decimation patterns and signal the index value of the selected decimation pattern to the transmitting device 102.
[0152]
[0163] Figure 10 is a block diagram showing exemplary components of a transmitting device and a receiving device according to the technique of the present disclosure. In the example of Figure 10, the transmitting device 102 may include the same components as those shown in Figure 2. However, in the example of Figure 10, the receiving device 104 may additionally include a reliability unit 1002. Unless otherwise stated, components named similarly in the transmitting device 102 and the receiving device 104 in Figures 2 and 10 perform the same function.
[0153]
[0164] In general, accurately predicting the most significant bits (MSBs) of a conversion coefficient can be easier than predicting the least significant bits (LSBs). This is because the MSBs are converted to longer geometric distances within the video picture region. In addition, blocks of a video picture with significant motion can be more difficult to predict accurately than blocks in areas with less motion.
[0154]
[0165] According to one or more techniques of this disclosure, the transmitting device 102 and the receiving device 104 may implement a system in which bit-level reliability values are used. The use of bit-level reliability values may enable the transmitting device 102 and the receiving device 104 to correctly weight a priori information. This may result in improved system performance (e.g., a reduction in the amount of data transmitted and / or improved video quality). For example, if "soft" information is used, decoding performance may be improved. In other words, if the receiving device 104 is "notified" of the reliability of each bit (what is the a priori probability that the bit value is "0" or "1"), the receiving device 104 can utilize this information and performance may be improved.
[0155]
[0166] In the example shown in Figure 10, the reliability unit 1002 of the receiving device 104 receives encoded video data from the video encoder 228 of the receiving device 104. The encoded video data may include conversion coefficients for the video data conversion block. In addition, in some examples, the reliability unit 1002 receives predictive quality information from the picture estimation unit 226 of the receiving device 104.
[0156]
[0167] In a given scaling process, for each bit position of the conversion coefficients in a conversion block, the prediction quality information includes a confidence value for that bit position. For example, the most significant bit of the conversion coefficient has a first confidence value, the second most significant bit has a second confidence value, the third most significant bit has a third confidence value, and so on. The confidence value of a bit position is a measure of how likely it is that the bit at that position has an incorrect value. For example, a bit at a given position may have an incorrect value if its predicted bit value is "0" but its actual value is "1", or if its predicted bit value is "1" but its actual value is "0".
[0157]
[0168] The picture estimation unit 226 may determine the reliability value of a bit position by collecting statistics on the rate of errors occurring in the bits at that bit position. For example, the picture estimation unit 226 may determine the probability that the most significant bit of the conversion coefficient contains an error, the second most significant bit of the conversion coefficient contains an error, the third most significant bit of the conversion coefficient contains an error, and so on. The picture estimation unit 226 may collect these statistics by counting the number of times the predicted bit (i.e., a bit in the estimated picture) was inaccurate out of the number of events tested. For example, the picture estimation unit 226 may estimate a picture, and the video encoder 228 may encode video data of the estimated picture. Subsequently, the video decoder 224 may decode the error-corrected encoded video data of the picture. The picture estimation unit 226 may compare the bits in the conversion coefficients of the encoded video data of the estimated picture with the error-corrected video data of the picture to determine whether a bit in the encoded video data of the estimated picture is incorrect.
[0158]
[0169] The reliability unit 1002 can convert probability values to LLR values. In some examples, the reliability unit 1002 can convert probability values to LLR values using the following formula.
[0159]
number
[0160] In the above formula, M is the LLR value, P error is the error probability value, and ln represents the natural logarithm function. In some examples, the LLR value is the confidence value.
[0161]
[0170] Figure 11 shows a diagram of exemplary error probabilities and corresponding log-likelihood ratio (LLR) absolute values using one or more techniques of the present disclosure. In the example in Figure 11, each transform block is represented using 150 bits. Graph 1100 plots the error probabilities for individual bit positions within the transform block. As can be seen from Graph 1100, bits at certain positions have a higher probability of error. Graph 1102 shows the error probabilities transformed into LLR absolute values. LLR absolute values can be scaled.
[0162]
[0171] In some examples, the receiving device 104 uses a dynamic scaling process. In the dynamic scaling process, the picture estimation unit 226 dynamically determines prediction quality information based on the video data. For example, some areas of the picture are more difficult to predict than areas that are easier to predict (e.g., areas where prediction accuracy is reduced). Examples of areas of the picture that are more difficult to predict may include areas with greater motion. For example, if the sum of the motion vectors in an area exceeds a threshold, the picture estimation unit 226 may determine that the area is a difficult-to-predict area. Thus, the picture estimation unit 226 may identify such areas and generate confidence values for the bits of the transformation coefficients of the transformation block, at least in part on whether the transformation block is inside or outside such areas. In some examples, the picture estimation unit 226 determines confidence values for the bits of the transformation coefficients based on general statistics on the error of bit position corrected based on whether the transformation block containing the transformation coefficients is in a difficult-to-predict area.
[0163]
[0172] The reliability unit 1002 may use reliability values to scale the bits of the conversion coefficients in the encoded video data generated by the video encoder 228. For example, bits of the encoded video data generated by the video encoder 228 may be considered "hard" bits and may have values of strictly 0 or strictly 1. The reliability unit 1002 may use the predictive quality information generated by the picture estimation unit 226 and the encoded video data generated by the video encoder 228 to determine "soft" bits that lie between 0 and 1. For example, if the value of a bit in the encoded video data is 1, the reliability unit 1002 may generate a "soft" value for that bit by multiplying the LLR absolute value for that bit by a negative 1 (i.e., -1). If the value of a bit in the encoded video data is 0, the reliability unit 1002 may generate a "soft" value for that bit by multiplying the LLR absolute value for that bit by a positive 1 (i.e., +1). Thus, the "soft" value or scaled value of a bit of a conversion coefficient may be M or -M.
[0164]
[0173] Therefore, in some examples, each bit can be converted to a scaled value in which the higher the confidence that the bit has a value of 0, the more positive the value, and the higher the confidence that the bit has a value of 1. The confidence unit 1002 provides the scaled value to the channel decoder 222 as priori information.
[0165]
[0174] The channel decoder 222 performs the channel decoding process using scaled values. For example, the reliability unit 1002 and the channel encoder 212 may encode the encoded video data into a codeword (e.g., a low-density parity check code). The bits of the codeword generated by the reliability unit 1002 may be scaled as described above. The bits of the codeword generated by the channel encoder 212 may be modified while passing through channel 230 so that the bits of the codeword can be received as values between -1 and 1. The channel decoder 222 may apply the LPDC decoding process to the codeword to correct errors in the codeword. The bit values in the corrected codeword are 0 or 1. The LPDC decoding process may then convert the codeword from the encoded video data to the original data. In other examples, other coding schemes may be used. In this example, the error-corrected data received by the channel decoder 222 may include cyclic redundancy check (CRC) data that is not used in the LPDC decoding process. In this way, the channel decoder 222 may determine the value of each bit of the conversion coefficient.
[0166]
[0175] In some examples, the channel encoder 212 may use predictive quality feedback to classify bits of the conversion coefficients before applying unequal protection channel coding, such as polar or spinal coding. Unequal protection channel coding involves assigning coding redundancy according to the importance of the information bits. For example, the channel encoder 212 may use reliability data to protect different bits according to the predictability by the receiving device 104.
[0167]
[0176] In some examples, the reliability unit 1002 sends predictive quality feedback to the transmitting device 102. The video encoder 210 may adjust one or more encoding parameters of the video encoding process that the video encoder 210 applies to the video data. For example, the transmitting device 102 may determine the compression ratio of a limited video encoding process based on the predictability (reliability) of the receiving device 104. For example, if the predictability is low, the transmitting device 102 may reduce the quality of the encoded video by reducing the number of conversion factors or the bit width per conversion factor so that less information is transmitted.
[0168]
[0177] In some examples, the transmitting device 102 may adjust one or more channel coding parameters used by the channel encoder 212 based on predictive quality feedback. For example, the channel encoder 212 may use unequal protection codes where protection depends on predictive quality feedback. In some examples, the channel encoder 212 may select different LDPC graphs depending on reliability. In some examples, the channel encoder 212 may change the encoding scheme for generating error-corrected data to increase the error correction capability for bit positions or areas of less reliable pictures, or to decrease the error correction capability for bit positions or areas of more reliable pictures.
[0169]
[0178] In some examples, the transmitting device 102 may update one or more bit puncturing parameters used by the puncturing unit 214 based on predictive quality feedback. For example, the puncturing unit 214 may modify the puncturing pattern to allow the transmission of more error correction data to less reliable bit locations and / or picture regions. Thus, the transmitting device 102 may avoid puncturing less reliable bits. In some examples, the puncturing unit 214 may modify the puncturing pattern to allow the transmission of less error correction data to more reliable bit locations and / or picture regions.
[0170]
[0179] In some examples, the predictive quality feedback that the reliability unit 1002 sends to the transmitting device 102 is applicable to the entire picture. In some examples, the reliability unit 1002 may send predictive quality feedback to the transmitting device 102 on a region-by-region basis. Each region may be a defined area within the picture. The reliability unit 1002 may send predictive quality feedback for some regions of the picture but not for others.
[0171]
[0180] In some examples, the reliability unit 1002 may periodically send predictive quality feedback to the transmitting device 102. For example, the reliability unit 1002 may send predictive quality feedback to the transmitting device 102 every N pictures, where N is an integer. In some examples, the reliability unit 1002 sends predictive quality feedback to the transmitting device 102 after a certain number of picture groups (GOPs) have been completed. In other examples, the reliability unit 1002 may send predictive quality feedback aperiodically, such as in response to certain conditions or events.
[0172]
[0181] In some examples, the predictive quality feedback may include predictive quality data based on one or more noise models, such as a Gaussian noise model or a Laplacian noise model. The noise model parameters may control one or more noise models. The reliability unit 1002 may transmit the noise model parameters to the transmitting device 102. The use of a noise model is an alternative for collecting error statistics. In this mode, the predictive error (statistics of the error between the estimated picture estimated by the picture estimation unit 226 and the actual picture reconstructed by the video decoder 224) may be modeled using a small number of parameters that describe the error distribution function. Since the noise model parameters may contain less data compared to bitwise statistics, the noise model parameters may be easier for the receiving device 104 to communicate to the transmitting device 102. The transmitting device 102 may use a noise model in the same way that the transmitting device 102 may use other types of predictive quality feedback.
[0173]
[0182] In some examples, instead of the reliability unit 1002 receiving predictive quality information from the picture estimation unit 226 of the receiving device 104, the video encoder 210 may generate the predictive quality information and send it to the reliability unit 1002. The video encoder 210 may determine the predictive quality information based on prior information about the video coding process. For example, the video encoder 210 may evaluate bitwise reliability (e.g., MSBs are more reliable than LSBs, lower frequency conversion coefficients are more reliable than higher frequency conversion coefficients) based on compression parameters. The video encoder 210 may perform a prediction process and evaluate statistics itself, and may have predetermined statistics for different light compression sets of parameters. In addition, the transmitting device 102 may use other sensors to evaluate instantaneous motion and adjust the reliability accordingly.
[0174]
[0183] Figure 12 is a flowchart illustrating exemplary operation of a transmitting device 102 using scaled bits according to the technique of the present disclosure. In the example of Figure 12, the transmitting device 102 may, for example, acquire video data from a video source 120 (1200). Furthermore, the transmitting device 102 may acquire predictive quality feedback (1202). In some examples, the transmitting device 102 may acquire predictive quality feedback from a receiving device 104. The predictive quality feedback includes bit reliability information. In some examples, the predictive quality feedback is expressed in terms of noise model parameters, such as parameters of a Gaussian noise model or a Laplacian noise model.
[0175]
[0184] The transmitting device 102 may adapt one or more of the video coding parameters, channel coding parameters, or bit puncturing parameters based on predictive quality feedback (1204). The video encoder 210 of the transmitting device 102 may perform a video coding process to generate coded video data (1206). The video coding process may be controlled by video coding parameters. For example, video coding parameters may control the number of conversion coefficients included in a conversion block, the number of bits included in the conversion coefficients, etc. In some examples, video coding parameters include quantization parameters, and the video encoder 210 may adapt the quantization parameters based on predictive quality feedback. For example, if the predictive quality feedback indicates low reliability, the quantization parameters may be reduced to lower the level of quantization. As part of performing the video coding process, the video encoder 210 may use quantization parameters to quantize the conversion coefficients of one or more picture conversion blocks.
[0176]
[0185] The channel encoder 212 of the transmitting device 102 may perform a channel coding process on scaled bits to generate channel-coded data (1208). The channel coding process may be controlled by channel coding parameters. For example, the channel coding parameters may control which LDPC graph the channel coding process uses to generate codewords, or the channel coding parameters may control the error correction capability, and so on. For example, the channel coding parameters may include an LDPC graph, and the channel encoder 212 may adapt an LDPC graph and use the LDPC graph to generate codewords for transmission to the receiving device.
[0177]
[0186] Furthermore, the puncturing unit 214 of the transmitting device 102 may perform bit puncturing on the error correction data generated by the channel encoder 212 (1210). The bit puncturing process may be controlled by bit puncturing parameters. For example, predictive quality feedback may indicate that a portion of the encoded video data is less reliable. Therefore, the transmitting device 102 may adjust the bit puncturing parameters to reduce bit puncturing on the error correction data for the less reliable portion of the encoded video data. The transmitting device 102 may transmit the channel-encoded data and the bit-punctured error correction data to the receiving device 104 (1212).
[0178]
[0187] Figure 13 is a flowchart illustrating exemplary operation of a receiving device 104 using scaled bits according to the technique of the present disclosure. In the example of Figure 13, the receiving device 104 may receive error correction data from the transmitting device (1300). The error correction data provides error correction information about the picture of the video data.
[0179]
[0188] The picture estimation unit 226 may generate prediction data for the picture (1302). The prediction data for the picture may include predictions of blocks of the picture based at least in part on one or more previously reconstructed pictures of the video data. For example, the picture estimation unit 226 may use interpretation and / or intrapretation to generate block predictions.
[0180]
[0189] Furthermore, the receiving device 104 may generate encoded video data based on predictive data for the picture (1304). For example, the video encoder 228 of the receiving device 104 may perform a video encoding process that generates encoded video data. The encoded video data includes a transformation block containing transformation coefficients.
[0181]
[0190] The receiving device 104 may scale the bits of the conversion coefficients of the conversion block based on a confidence value for bit positions (1306). In some examples, the receiving device 104 generates a confidence value. For example, the receiving device 104 may generate a confidence value based on statistics regarding the occurrence of errors at bit positions. In some examples, the receiving device 104 may generate a confidence value based on the reliability characteristics for individual regions of a picture of video data. In some examples, the receiving device 104 may generate a confidence value based on a noise model. Furthermore, in some examples, the receiving device 104 may send the confidence value to the transmitting device 102. In other examples, the receiving device 104 may receive a confidence value from the transmitting device 102.
[0182]
[0191] In addition, the channel decoder 222 of the receiving device 104 may use error correction data to generate error-corrected encoded video data in order to perform error correction operations on the scaled bits of the conversion coefficients of the conversion block (1308).
[0183]
[0192] The video decoder 224 of the channel decoder 222 can reconstruct the picture based on the error-corrected encoded video data (1310).
[0184]
[0193] This disclosure describes techniques that can reduce the complexity of video encoding in transmitting devices such as augmented reality (XR) headsets. A transmitting device may acquire multiview video data, which may include pictures from two or more viewpoints. For example, an XR headset may include two cameras for stereoscopic viewing of the scene the user is seeing. In this example, the multiview video data may include pictures from each of the cameras.
[0185]
[0194] Processing the content of multi-view video data can consume significant processing resources. For example, in augmented reality (AR) or mixed reality (MR) scenarios, considerable processing resources may be required to determine where virtual elements should be placed and how they should appear. Multi-view video data can help in processing virtual elements. For instance, the same virtual element may need to be darker when it should be placed in a shaded area of the scene and brighter when it should be placed in a sunny area of the scene. Another example is when the system needs to analyze the scene's content to determine whether a virtual element should be obscured by a physical element in the scene, such as a rock or a tree. Multi-view video data can be useful in determining the depth of objects in a scene. To keep the XR headset bright and conserve battery power, it is sometimes desirable to minimize the processing of video data performed on the XR headset. Therefore, processing video data on another device, such as the user's smartphone or another nearby device, can help reduce the demand on processing resources for the XR headset.
[0186]
[0195] While multiview video data can be very useful in certain situations, simply transmitting unencoded multiview video data can be impractical because the amount of data required to transmit multiple parallel streams of video data simultaneously can be very large. However, there is often considerable redundancy between the pictures of different views in multiview video data. For example, what a person sees with their left eye is often not so different from what their right eye sees. Therefore, video compression techniques have been developed to reduce this redundancy in order to reduce the amount of data required to transmit multiview video data.
[0187]
[0196] However, some of the techniques for multiview video coding require considerable computational resources. For example, a video encoder may determine the difference between two or more sets of simultaneous pictures to determine a depth map of a scene. The depth map is an array of values indicating the depth / distance from the camera to the objects shown in the pictures. In this example, one of the simultaneous pictures may be an anchor picture, and one or more of the simultaneous pictures may be non-anchor pictures. The video encoder may use the depth map to calculate the disparity vector for blocks in the non-anchor pictures. The disparity vector for a block indicates the lateral displacement between the block and the corresponding block in another simultaneous picture, such as an anchor picture. Generally, blocks representing deeper objects have smaller disparity vectors than blocks representing closer objects. The video encoder may use the disparity vectors of the blocks to determine predicted blocks, generate residual data based on the predicted blocks, apply transformations to the residual data, quantize the transformation coefficients of the resulting transformed blocks, and signal the quantized transformation coefficients. In this example, considerable computational resources may be involved in generating the depth map.
[0188]
[0197] In another example, illumination levels may differ between simultaneous pictures from different viewpoints. These illumination differences can impair the coding efficiency of multi-view video coding. Therefore, illumination compensation may be applied to non-anchor pictures to temporarily adjust their illumination levels according to an illumination compensation factor in order to better match the illumination levels of non-anchor pictures to those of anchor pictures during video coding. The original illumination levels of non-anchor pictures can be restored during video decoding. Determining the illumination compensation factor can consume computational resources.
[0189]
[0198] The techniques of this disclosure may transfer some of the processing related to multiview video coding from a transmitting device (e.g., an XR headset) to a receiving device (e.g., a mobile device). For example, the transmitting device may obtain a first set of multiview pictures of video data. The first set of multiview pictures includes a first picture and a second picture. The first picture is from a first viewpoint, and the second picture is from a second viewpoint. The transmitting device may transmit first coded video data to the receiving device. The first coded video data is based on the first set of multiview pictures. The transmitting device may receive a multiview coding queue from the receiving device. Furthermore, the transmitting device may obtain a second set of multiview pictures of video data. The second set of multiview pictures includes a third picture and a fourth picture. The third picture is from a first viewpoint, and the fourth picture is from a second viewpoint. The transmitting device may perform a multiview encoding process on a second set of multiview pictures based on a multiview encoding queue received from the receiving device in order to generate second encoded video data. The multiview encoding process reduces interview redundancy between the third and fourth pictures. The transmitting device may then transmit the second encoded video data to the receiving device.
[0190]
[0199] Similarly, a receiving device may receive first encoded video data from a transmitting device. The first encoded video data is based on a first set of multiview pictures of video data. The first set of multiview pictures may include a first picture and a second picture. The first picture is from a first viewpoint, and the second picture is from a second viewpoint. The receiving device may determine a multiview coding queue based on the first encoded video data. The receiving device may send the multiview coding queue to the transmitting device. In addition, the receiving device may receive second encoded video data from the transmitting device. The second encoded video data is based on a second set of multiview pictures including a third picture and a fourth picture. The second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third and fourth pictures based on the multiview coding queue.
[0191]
[0200] Since the receiving device determines the multiview coding queue and sends it to the transmitting device, the burden of determining the multiview coding queue can be shifted from the transmitting device to the receiving device. This can reduce the demand on resources at the transmitting device.
[0192]
[0201] Referring to Figure 2, the video encoder 210 may perform a multiview coding process based on a multiview coding queue acquired from the receiving device 104. For example, the multiview coding queue may include a depth map. In this example, the video encoder 210 may use the depth map to estimate the disparity vector for blocks of pictures in non-anchor views of the multiview video data. The video encoder 210 may use the disparity vector of the current block of the current picture to determine a predicted block for the current block based on a sample of the concurrently referenced picture. The concurrently referenced picture has the same picture order count (POC) value as the current picture. The video encoder 210 may determine residual data for the current block based on the original sample of the current block and the predicted block for the current block. The video encoder 210 may apply one or more transformations to the residual data to generate one or more transformation blocks. The video encoder 210 may quantize the transformation coefficients in the transformation blocks. The coded video data generated by the video encoder 210 may be based on the quantized transformation coefficients.
[0193]
[0202] In some examples, a multiview coding queue may include one or more illumination compensation factors. When coding the current picture of multiview video data, the video encoder 210 may modify each sample of the current picture based on one or more illumination compensation factors. In some examples, different illumination compensation factors may be applied to different regions of the current picture. Modifying the samples of the current picture in this way may better match the illumination level of the current picture with the illumination level of the concurrently referenced picture. After modifying the samples of the current picture, the video encoder 210 may perform a multiview coding process, as described in the previous paragraph, to code a block of the current picture.
[0194]
[0203] Furthermore, according to some examples of this disclosure, the video decoder 224 may perform a multiview decoding process. For example, the video decoder 224 may use the disparity vector of the blocks of the current picture to generate predicted blocks. The video decoder 224 may use the predicted block and residual data received from the channel decoder 222 to reconstruct the samples of the blocks of the current picture. In some examples where the video encoder 210 has applied illumination compensation to the picture, the video decoder 224 may use illumination parameters to invert the illumination compensation applied to the picture. In other examples, the video decoder 224 may apply other multiview decoding operations.
[0195]
[0204] Furthermore, according to one or more techniques of this disclosure, the video decoder 224 may determine a multiview coding queue based on the coded video data received from the transmitting device 102. For example, the video decoder 224 may determine a depth map, illumination compensation parameters, and other information that may be used in the multiview coding operation. The receiving device 104 may send the multiview coding queue back to the transmitting device 102 so that the transmitting device 102 can use the multiview coding queue to perform the multiview coding process on subsequent pictures.
[0196]
[0205] The picture estimation unit 226 may generate an estimate of the next picture in the video data. In some examples, the picture estimation unit 226 may estimate a picture based on one or more previously reconstructed reference pictures associated with different views. For example, the picture estimation unit 226 may extrapolate the content of a picture from a picture at the same moment using information such as a disparity vector or depth map from a picture for a previous moment. In another example, the picture estimation unit 226 may extrapolate a picture based on one or more pictures associated with the same view, regardless of pictures associated with other views, in much the same way as discussed elsewhere in this disclosure with respect to the picture estimation unit 226 for estimating a picture in single-view video data.
[0197]
[0206] The video encoder 228 may perform the same operations as the video encoder 210 for the next estimated picture. For example, the video encoder 228 may perform intraprediction to predict a block and use the predicted block and the corresponding predicted block of the picture generated by the picture estimation unit 226 to generate residual data. In some examples, the video encoder 228 may perform the same multiview coding process as the video encoder 210 using a multiview coding queue. The video encoder 228 may apply a transformation (e.g., DCT transformation) to the residual data to generate transformation coefficients. The video encoder 228 may apply quantization to the transformed coefficients.
[0198]
[0207] Figure 14 is a flowchart illustrating an exemplary exchange of data between a transmitting device 102 and a receiving device 104 for multiview processing using one or more techniques of the present disclosure. In the example of Figure 14, the transmitting device 102 may acquire a first set of multiview pictures (1400). Based on the first set of multiview pictures, the transmitting device 102 may transmit first encoded video data to the receiving device 104. In some examples, the transmitting device 102 performs light compression on the first set of multiview pictures to generate the first encoded video data. In other examples, the first encoded video data may include an unencoded version of the first set of multiview pictures.
[0199]
[0208] The receiving device 104 may perform multiview processing on a first set of multiview pictures (1402). For example, the receiving device 104 may decode the first set of multiview pictures if necessary. In addition, the receiving device 104 may determine a multiview coding queue, for example, as described elsewhere in this disclosure. The receiving device 104 may transmit the multiview coding queue to the transmitting device 102.
[0200]
[0209] Furthermore, in the example of Figure 14, the transmitting device 102 may acquire a second set of multiview pictures (1404). The transmitting device 102 may perform a multiview coding process on the second set of multiview pictures to generate a second coded video data (1406). The second coded video data may include coded anchor pictures and secondary (non-anchor) pictures. The receiving device 104 may perform multiview decoding on the second coded video data to reconstruct the second set of multiview pictures (1408). The receiving device 104 may also perform multiview processing on the second set of multiview pictures to determine an updated multiview coding queue (1410). The receiving device 104 may send the updated multiview coding queue to the transmitting device 102. The transmitting device 102 may use the updated multiview coding queue for multiview coding of subsequent sets of multiview pictures.
[0201]
[0210] Figure 15 is a flowchart illustrating exemplary operation of a transmitting device 102 for multiview processing using the technique of the present disclosure. In the example of Figure 5, the transmitting device 102 may acquire a first set of multiview pictures of video data (1500). The first set of multiview pictures includes a first picture and a second picture. The first picture is from a first viewpoint, and the second picture is from a second viewpoint. The communication interface 118 of the transmitting device 102 (Figure 1) may transmit first encoded video data to the receiving device 104 (1502). The first encoded video data is based on the first set of multiview pictures. The transmitting device 102 may receive a multiview encoding queue from the receiving device 104 (1504). In some examples, the multiview coding queue includes one or more motion data for relative shifts between blocks of first-viewpoint pictures and blocks of second-viewpoint pictures, brightness corrections between first-viewpoint and second-viewpoint pictures, inter-block shifts between anchor blocks and reconstructed blocks, or reference shifts. The transmitting device 102 may receive the multiview coding queue in one of several ways. For example, the transmitting device 102 may receive the multiview coding queue via uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, or other types of messages.
[0202]
[0211] Furthermore, the transmitting device 102 may acquire a second set of multiview pictures of the video data (1506). The second set of multiview pictures includes a third picture and a fourth picture. The third picture is from the first viewpoint, and the fourth picture is from the second viewpoint. The video encoder 210 may perform a multiview encoding process on the second set of multiview pictures based on the multiview encoding queue received from the receiving device in order to generate second encoded video data (1508). The multiview encoding process reduces interview redundancy between the third picture and the fourth picture. The transmitting device 102 may then transmit the second encoded video data to the receiving device 104 (1510).
[0203]
[0212] The operation in Figure 15 may be performed multiple times for subsequent sets of multiview pictures. For example, after sending a second encoded video data to a receiving device, the transmitting device 102 may receive an updated multiview encoding queue from the receiving device 104. The transmitting device 102 may obtain a third set of multiview pictures of the video data. The third set of multiview pictures may include a fifth and a sixth picture, where the fifth picture is from the first viewpoint and the sixth picture is from the second viewpoint. The video encoder 210 of the transmitting device 102 may encode the third set of multiview pictures based on the updated multiview encoding queue received from the receiving device in order to generate a third encoded video data. The transmitting device 102 may send the third encoded video data to the receiving device 104.
[0204]
[0213] Figure 16 is a flowchart illustrating exemplary operation of a receiving device 104 for multiview processing using the technique of the present disclosure. In the example of Figure 16, the receiving device 104 may receive first encoded video data from the transmitting device 102 (1600). For example, the receiving device 104 may receive first encoded video data via a communication interface 134 (Figure 1). The first encoded video data is based on a first set of multiview pictures of video data. The first set of multiview pictures may include a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint. The receiving device 104 may determine a multiview encoding queue based on the first encoded video data (1602).
[0205]
[0214] The receiving device 104 may transmit a multiview coded queue to the transmitting device 102 (1604). The receiving device 104 may transmit the multiview coded queue in one of several ways. For example, the receiving device 104 may transmit a decimation pattern instruction via an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, or another type of message.
[0206]
[0215] In addition, the receiving device 104 may obtain second encoded video data from the transmitting device 102 (1606). The second encoded video data is based on a second set of multiview pictures, including a third picture and a fourth picture. The second encoded video data is encoded using a multiview encoding process that reduces interview redundancy between the third picture and the fourth picture based on a multiview encoding queue. The video decoder 224 of the receiving device 104 may decode the second encoded video data.
[0207]
[0216] In some examples, the multiview coding queue includes a depth map indicating the depth of objects represented in the first and second pictures. The receiving device 104 may determine the depth map based on the first and second pictures as part of determining the multiview coding queue. In some examples, the multiview coding queue includes one or more illumination compensation coefficients, and the receiving device 104 may determine the illumination compensation coefficients based on the first and second pictures as part of determining the multiview coding queue.
[0208]
[0217] The process in Figure 16 may be repeated multiple times. For example, the receiving device 104 may determine a second multiview coding queue based on the second encoded video data. The receiving device 104 may send the second multiview coding queue to the transmitting device 102. Subsequently, the receiving device 104 may receive a third encoded video data from the transmitting device 102. The third encoded video data is based on a third set of multiview pictures including a fifth picture and a sixth picture, and the third encoded video data is encoded using a multiview coding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview coding queue.
[0209]
[0218] According to one or more techniques of this disclosure, the transmitting device 102 may receive a decimation pattern instruction from the receiving device 104. The decimation pattern instruction may indicate a decimation pattern. The receiving device 104 may determine the decimation pattern, as will be described in more detail elsewhere in this disclosure. The decimation pattern may be a non-transmitted pattern of encoded video data.
[0210]
[0219] The transmitting device 102 may receive decimation pattern instructions in one of several ways. For example, the transmitting device 102 may receive decimation pattern instructions via uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, sidelink control information (SCI), or other types of messages.
[0211]
[0220] The transmitting device 102 may apply a decimation pattern to the encoded video data generated by the video encoder 210, thereby generating decimated video data. For example, the decimation pattern may indicate a pattern that skips the transmission of the encoded video data for an entire picture. Thus, in this example, the transmitting device 102 (e.g., the channel encoder 212 of the transmitting device 102) may transmit the encoded video data for some pictures and not for others according to the indicated pattern. For example, the transmitting device 102 may skip the transmission of the encoded video data for every other picture. In another example, the transmitting device 102 may transmit the encoded video data for one picture and then not transmit the encoded video data for the next two or more pictures.
[0212]
[0221] In another example, a decimation pattern may represent a pattern that skips the transmission of encoded video data for a specified region within a picture. Certain regions of a series of pictures may not change much from picture to picture, even if they do change. For example, changes may occur in a more localized region of interest, while the background of a scene from a static viewpoint may not change significantly. Since regions outside the region of interest do not change much, such regions may be easier to predict accurately. Therefore, according to the technique of this disclosure, the receiving device 104 can identify regions outside the region of interest. Thus, the transmitting device 102 may transmit encoded video data for the region of interest, but not for other regions.
[0213]
[0222] In another example, the video data may be multi-view video data, and the decimation pattern may indicate a pattern that skips the transmission of encoded video data for pictures from a given view. For example, two views, such as a view that primarily shows distant objects, may have very similar content. Therefore, in this example, the transmitting device 102 may transmit encoded video data for one of the views, as indicated by the decimation pattern, and may not transmit encoded video data for one or more other views.
[0214]
[0223] In another example, a decimation pattern may indicate a pattern of bits that should be omitted from the syntax elements representing the conversion coefficients. For example, a decimation pattern may indicate that a certain number of lower bits should be omitted from the conversion coefficients. In some examples, a decimation pattern may indicate that certain conversion coefficients (e.g., high-frequency conversion coefficients) should be omitted.
[0215]
[0224] Figure 17 is a block diagram showing exemplary components of a transmitting and receiving device that perform decimation on encoded video data using the technique of the present disclosure. In the example of Figure 17, the transmitting device 102 also includes a video encoder 210, a channel encoder 212, a puncturing unit 214, and a transmitter decimation unit 1700. The receiving device 104 also includes a depuncturing unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, a video encoder 228, and a receiver decimation unit 1702. The video encoder 210, channel encoder 212, puncturing unit 214, depuncturing unit 220, channel decoder 222, video decoder 224, picture estimation unit 226, and video encoder 228 may operate in the same manner as described elsewhere in the present disclosure.
[0216]
[0225] However, in the example of Figure 17, the transmitter decimation unit 1700 may apply a decimation pattern to the encoded video data after the channel encoder 212 has generated error correction data for the encoded video data. The decimation pattern indicates a non-transmitted pattern of the encoded video data. For example, the transmitter decimation unit 1700 may not cause the transmitting device 102 to transmit encoded video data for a particular picture, a region of a picture, a pattern of blocks within a picture, a particular view, etc. The receiver decimation unit 1702 may determine a decimation pattern instruction based on the picture reconstructed by the video decoder 224. The receiver decimation unit 1702 may transmit a decimation pattern instruction indicating a decimation pattern to the transmitting device 102. The transmitter decimation unit 1700 may apply the decimation pattern indicated by the decimation pattern instruction.
[0217]
[0226] Figure 17 illustrates a DVC-based scheme, but the techniques of this disclosure for sending decimation pattern instructions from the receiving device 104 to the transmitting device 102 are not necessarily limited thereto. For example, in some examples, the picture estimation unit 226 and the video encoder 228 may be omitted.
[0218]
[0227] Figure 18 is a conceptual diagram illustrating an exemplary exchange of information, including decimation pattern instructions, using the technique of the present disclosure. In the example of Figure 18, the transmitting device 102 transmits a first set of pictures (e.g., picture n-n1, picture nn) n1+1 The receiving device 104 may transmit encoded video data for picture n). The transmitting device 102 may also transmit error correction data for the first set of pictures.
[0219]
[0228] The receiving device 104 may transmit a decimation pattern instruction indicating a decimation pattern determined based on a first set of encoded pictures, and the transmitting device 102 may receive such an instruction. In the example in Figure 18, the decimation pattern decimates the pictures according to a 1:2 ratio. In other words, video data encoded at a ratio of 2 pictures to 1 picture is transmitted.
[0220]
[0229] Therefore, the transmitting device 102 can transmit encoded video data for a second set of pictures to the receiving device 104. According to the decimation pattern indicated by the received decimation pattern instruction, the transmitting device 102 skips transmitting the encoded video data for every other picture in the second set of pictures. As shown in the example in Figure 18, the index values of the pictures in the second set of pictures (e.g., n+2, n+4, n+n2) are incremented by 2 instead of 1, as was the case for the first set of pictures.
[0221]
[0230] Subsequently, the receiving device 104 may determine, based on the second set of pictures, that a more appropriate decimation pattern is the 1:1 decimation pattern (i.e., the decimation pattern in which the transmitting device 102 transmits encoded video data for each picture). Thus, in the example of Figure 18, the receiving device 104 may transmit and the transmitting device 102 may receive a second decimation pattern instruction indicating the second decimation pattern. Subsequently, the transmitting device 102 may transmit encoded video data for a third set of pictures. According to the second decimation pattern, the transmitting device 102 does not skip transmitting encoded video data for any of the pictures in the third set of pictures. Thus, as shown in the example of Figure 18, the index values (e.g., n+n2+1, n+n2+2, etc.) increment by 1 instead of 2.
[0222]
[0231] Figure 19 is a flowchart illustrating exemplary operation of the transmitting device 102 in which the transmitting device 102 receives a decimation pattern instruction using the technique of the present disclosure. In the example of Figure 19, the video encoder 210 of the transmitting device 102 may encode a first set of pictures of video data to generate first encoded video data (1900). The transmitting device 102 may transmit the first encoded video data to the receiving device 104 (1902).
[0223]
[0232] Furthermore, the transmitting device 102 may receive a decimation pattern instruction from the receiving device 104 indicating a decimation pattern determined based on a first set of pictures (1904). The decimation pattern may be a pattern of not transmitting encoded video data. For example, in some examples, the decimation pattern indicates a pattern of skipping the transmission of encoded video data for an entire picture. In other words, the transmitting device 102 may not transmit encoded video data for a particular picture, but may transmit some or all of the encoded video data for other pictures. In some examples, the decimation pattern indicates a pattern of skipping the transmission of encoded video data for a particular region within a picture. For example, the decimation pattern may indicate that the transmitting device 102 should skip the transmission of encoded video data associated with a particular block of a picture, as shown, for example, in Figures 6 and 9. In some examples where the video data is multiview video data, the decimation pattern may indicate a pattern of skipping the transmission of encoded video data for a picture from a particular view. In such examples, a view may be associated with a sensor on the same user device (e.g., the same XR headset), or it may be associated with a sensor on a different user device (e.g., a different XR headset worn by a different user) or a different camera. In some examples, there may be different decimation patterns for different areas without pictures. For example, no decimation may be applied to the region of interest, while a decimation pattern that restricts the transmission of blocks, lower bits, or high-frequency conversion coefficients may be applied to areas of the picture outside the region of interest.
[0224]
[0233] The video encoder 210 may encode a second set of pictures of video data to generate second encoded video data (1906). In addition, the transmitter decimation unit 1700 may apply a decimation pattern to the second encoded video data to generate decimated video data (1908). The transmitting device 102 may transmit the decimated video data to the receiving device (1910).
[0225]
[0234] In some examples, the transmitter decimation unit 1700 may determine a decimation pattern. Thus, in the situation shown in Figure 19, the transmitting device 102 may encode a third set of pictures of video data to generate a third encoded video data, determine a second decimation pattern indicating a second pattern of the encoded video data that is not transmitted, and apply the second decimation pattern to the third encoded video data to generate a second decimated video data. The transmitting device 102 may transmit the second decimated video data to the receiving device 104. The transmitting device 102 may transmit a second decimation pattern instruction to the receiving device. The second decimation pattern instruction indicates that the second decimation pattern has been applied to the third encoded video data.
[0226]
[0235] The transmitter decimation unit 1700 can determine the decimation pattern in various ways. For example, the transmitter decimation unit 1700 can test various decimation patterns. When testing decimation patterns, the transmitter decimation unit 1700 can apply the decimation pattern to a picture and reconstruct the picture from error correction data for one or more previous original pictures of that picture and video data. The transmitter decimation unit 1700 can compare the reconstructed picture to the original picture to determine the level of distortion. The transmitter decimation unit 1700 can compare the levels of distortion associated with different decimation patterns to determine the decimation pattern.
[0227]
[0236] In some examples, the operation shown in Figure 19 is performed in a DVC (Digital Video Correction) scenario. Thus, the channel encoder 212 of the transmitting device 102 may generate first error correction data based on first encoded video data. The transmitting device 102 may transmit the first error correction data to the receiving device 104. The channel encoder 212 may generate second error correction data based on second encoded video data. The transmitting device 102 may transmit the second error correction data to the receiving device.
[0228]
[0237] Figure 20 is a flowchart illustrating exemplary operation of a receiving device 104, which transmits a decimation pattern instruction using the technique of the present disclosure. In the example of Figure 20, the receiving device 104 may receive first encoded video data from the transmitting device 102 (2000). The video decoder 224 may perform a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data (2002).
[0229]
[0238] In addition, the receiver decimation unit 1702 may determine a decimation pattern indicating a pattern of non-transmission of encoded video data based on a first set of pictures (2004). In some examples, the decimation pattern indicates a pattern of skipping the transmission of encoded video data for the entire picture. In some examples, the decimation pattern indicates a pattern of skipping the transmission of encoded video data for a specified region or block within a picture, as shown in the examples in Figures 6 and 9. In some examples, the video data is multiview video data, and the decimation pattern indicates a pattern of skipping the transmission of encoded video data for pictures from a specified view.
[0230]
[0239] The receiver decimation unit 1702 may determine the decimation pattern in one of several ways. For example, the receiver decimation unit 1702 may apply one or more trial decimation patterns to the error-corrected encoded video data generated by the channel decoder 222 for a first set of pictures in order to generate decimated encoded video data. The receiver decimation unit 1702 may then cause the channel decoder 222 to apply an error correction process to correct the decimated encoded video data based on the first error-corrected data in order to generate error-corrected trial video data. The receiver decimation unit 1702 may then cause the video decoder 224 to apply a decoding process to reconstruct the first set of pictures based on the error-corrected trial video data. The receiver decimation unit 1702 may determine whether a decimation pattern satisfies one or more criteria based on a comparison between a first set of pictures reconstructed based on error-corrected trial video data and a first set of pictures reconstructed based on the first error-corrected video data. For ease of explanation, the disclosure may refer to the pictures reconstructed based on error-corrected trial video data as “trial pictures” and the pictures reconstructed based on the first error-corrected video data as “baseline pictures.” The receiver decimation unit 1702 may repeat this procedure with multiple trial decimation patterns until the receiver decimation unit 1702 identifies a decimation pattern that satisfies the criteria.
[0231]
[0240] For example, the receiver decimation unit 1702 may compare each trial picture with a corresponding baseline picture to determine whether the trial picture meets a criterion. For example, the receiver decimation unit 1702 may determine that a trial picture meets a criterion if the sum of the differences between the trial picture and the corresponding baseline picture is less than a certain amount. If at least a given number of trial pictures exceed the threshold, the receiver decimation unit 1702 may select a decimation pattern associated with the trial pictures.
[0232]
[0241] In a more general example, the receiver decimation unit 1702 may apply a function to the trial picture and the corresponding baseline picture to generate a value. If the value is less than a threshold, the receiver decimation unit 1702 may select a decimation pattern associated with the trial picture.
[0233]
[0242] Furthermore, in some examples, the receiver decimation unit 1702 may, under certain conditions, cancel the use of a decimation pattern (for example, revert to a pattern in which all encoded video data is transmitted) or change to a less aggressive decimation pattern. For example, the receiver decimation unit 1702 may cancel the use of a decimation pattern if it determines that a given number of baseline pictures do not meet the criteria. For example, the receiver decimation unit 1702 may determine that a trial picture does not meet the criteria if the sum of the differences between the trial picture and the corresponding baseline picture is greater than a certain amount. If at least a given amount of trial pictures that do not meet the criteria exceeds a threshold, the receiver decimation unit 1702 may cancel the use of a decimation pattern or revert to a less aggressive decimation pattern. In a more general example, the receiver decimation unit 1702 may apply a function to the trial picture and the corresponding baseline picture to generate a value. If the value is greater than the second threshold, the receiver decimation unit 1702 may cancel the use of the decimation pattern or revert to a less aggressive decimation pattern.
[0234]
[0243] The receiver decimation unit 1702 may transmit a decimation pattern instruction to the transmitting device 102 indicating the determined decimation pattern (2006). The receiver decimation unit 1702 may transmit the decimation pattern instruction using an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, a sidelink control information (SCI) message, or another type of message.
[0235]
[0244] The receiving device 104 may receive decimated video data from the transmitting device 102 (2008). The decimated video data may include second encoded video data to which a decimation pattern has been applied. The second encoded video data is generated based on a second set of pictures of video data.
[0236]
[0245] The video decoder 224 may perform a decoding process to reconstruct a second set of pictures based on the second error-corrected encoded video data (2010). The video decoder 224 may perform the same decoding process as described elsewhere in this disclosure.
[0237]
[0246] In some examples, the receiving device 104 may receive and use a decimation pattern instruction from the transmitting device 102. Thus, in the example of Figure 20, the receiving device 104 may receive a second decimation pattern instruction indicating a second pattern of untransmitted encoded video data. The receiving device 104 may receive third error correction data and second decimated video data from the transmitting device 102. The second decimated video data may include third encoded video data to which the second decimation pattern has been applied. The third encoded video data may be generated based on a third set of pictures of the video data. The channel decoder 222 may apply an error correction process to generate third error-corrected encoded video data based on the third encoded video data and the third error correction data. The video decoder 224 may apply a decoding process to reconstruct the third set of pictures based on the third error-corrected encoded video data.
[0238]
[0247] In some examples, the process shown in Figure 20 may be performed in a DVC-based implementation. Thus, the receiving device 104 may receive first error correction data from the transmitting device 102. The receiving device 104 applies an error correction process to correct the first encoded video data based on the first error correction data in order to generate first error-corrected encoded video data. The receiving device 104 may perform a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data. In addition, the receiving device 104 may receive second error correction data from the transmitting device 102. The receiving device 104 may apply an error correction process to generate second error-corrected encoded video data based on the second encoded video data, the predicted encoded video data generated by the video encoder 228, and the second error correction data. The receiving device 104 may perform a decoding process to reconstruct a second set of pictures based on the second error-corrected encoded video data.
[0239]
[0248] During the video encoding process, a video encoder typically analyzes multiple encoding options and selects the best one. For example, a video encoder may analyze multiple ways of dividing a maximum coding unit (LCU) or macroblock into coding units (CUs) and / or prediction units (PUs). In another example, when a video encoder performs intra-prediction to generate prediction blocks for PUs, it may analyze multiple intra-prediction modes. In yet another example, when a video encoder performs inter-prediction to generate prediction blocks for PUs, it may analyze multiple reference pictures and motion vectors. Such analysis and selection can be resource-intensive. For example, to be efficient, a video encoder may need to process multiple options in parallel, which increases the hardware complexity and power requirements of the video encoder. Analysis and selection may also involve multiple requests to read and write data to memory, which further increases power requirements.
[0240]
[0249] According to one or more techniques of this disclosure, the majority of the process of analyzing and selecting encoding operations is transferred from a transmitting device (e.g., transmitting device 102) to a receiving device (e.g., receiving device 104). For example, the transmitting device may encode a first picture of video data in order to generate first encoded video data. The transmitting device may transmit the first encoded video data to the receiving device. The receiving device may receive the first encoded video data from the transmitting device and reconstruct the first picture based on the first encoded video data. In addition, the receiving device may estimate a second picture of the video data based on the first picture. The second picture may be a picture that follows the first picture in the decoding order. The receiving device may generate encoding selection data for the estimated second picture. The encoding selection data indicates the encoding selection used to encode the estimated second picture. The receiving device may transmit the encoding selection data for the second picture. The transmitting device may receive the encoding selection data for the second picture of the video data. The transmitting device may encode a second picture based on encoding selection data to generate second encoded video data. The transmitting device may transmit the second encoded video data to the receiving device. The receiving device may receive the second encoded video data from the transmitting device. The receiving device may reconstruct the second picture based on the second encoded video data. In this way, since the transmitting device receives encoding selection data from the receiving device, the transmitting device does not need to perform resource-intensive analysis and selection processes while encoding the second picture, because the analysis and selection processes have already been performed on the second picture at the receiving device. This can reduce the resource requirements of the transmitting device.
[0241]
[0250] Transmitting and receiving devices may communicate using low-frequency, low-power links over ultra-wideband (e.g., broadband) communication links. In some examples, transmitting and receiving devices may communicate using time division duplexing (TDD), sub-band non-overlapping full duplex (SBFD), or SFFD. The low latency associated with this type of communication may allow the transmitting device to receive encoded selection data quickly enough to continue transmitting encoded video data to meet a given picture rate.
[0242]
[0251] Figure 21 is a block diagram showing exemplary components of a transmitting device 102 and a receiving device 104 that transmits encoded selection data to the transmitting device using the technique of the present disclosure. In the example of Figure 21, the transmitting device 102 may include a video encoder 210, a channel encoder 212, and a puncturing unit 214. The receiving device 104 may include a depuncturing unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, and a video encoder 228.
[0243]
[0252] In the example in Figure 21, the video encoder 210 may encode pictures of video data. Unlike some of the examples provided above, the video encoder 210 may perform a complete video encoding process, which may include intra-prediction and inter-prediction. In some examples, the video encoder 210 may encode video data using video codecs such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, Essential Video Coding (EVC), and AV1. The channel encoder 212, puncturing unit 214, depuncturing unit 220, and channel decoder 222 may operate in the same manner as described elsewhere in this disclosure.
[0244]
[0253] Furthermore, in the example shown in Figure 21, the video decoder 224 of the receiving device 104 may perform a video decoding process on the error-corrected encoded video data generated by the channel decoder 222. The video decoder 224 may perform a complete video decoding process, including intra-prediction and inter-prediction. The video decoder 224 may use the same video codec as the video encoder 210.
[0245]
[0254] After the video decoder 224 reconstructs at least a portion of the picture in the video data, the picture estimation unit 226 may estimate the corresponding portion of the subsequent picture following the reconstructed picture. The picture estimation unit 226 may estimate the subsequent picture in the same manner as described elsewhere in this disclosure. Furthermore, the video encoder 228 may apply a video encoding process to the subsequent picture. In examples where the video encoder 210 and video decoder 224 use a video codec, the video encoder 228 may use the same codec.
[0246]
[0255] However, according to one or more techniques of the present disclosure, the receiving device 104 may transmit coding selection data 2100 to the transmitting device 102. The coding selection data 2100 indicates the coding selection used to encode the estimated subsequent picture. For example, the coding selection data may include motion parameters for a block in the estimated subsequent picture. The motion parameters for a block may include motion vectors, reference picture indicators, merge candidate indices, affine motion parameters, and other data used to determine the predicted block for a block in one or more reference pictures. Thus, in this example, the video encoder 228 may encode the block using interprediction, and the coding selection data 2100 may include motion parameters indicating how the video encoder 228 encoded the block using interprediction.
[0247]
[0256] In some examples, the coding selection data may include intra-prediction parameters for blocks in the estimated subsequent picture. The intra-prediction parameters may include data indicating the intra-prediction mode (e.g., planar mode, DC mode, directional prediction mode, etc.) used by the video encoder 228 for intra-prediction of the blocks. In some examples, the coding selection data may include other information, such as information describing how the video encoder 228 divided the estimated subsequent picture into blocks, whether residual prediction is used, whether and how intra-block copying (IBC) is used, and whether any particular filter is used.
[0248]
[0257] The video encoder 210 may use coding selection data 2100 when encoding actual (unpredicted) subsequent pictures. That is, instead of exploring different possibilities during the video encoding process, the video encoder 210 may use the video coding selection indicated by the coding selection data 2100. For example, the coding selection data 2100 may indicate that a particular block of a subsequent picture will be encoded in a particular intra-prediction mode. Thus, in this example, when encoding a subsequent picture, the video encoder 210 may encode a particular block using a particular intra-prediction mode without analyzing different potential intra-prediction modes to select a particular intra-prediction mode. In another example, the coding selection data 2100 may indicate motion vectors and reference pictures for a particular block of a subsequent picture. Thus, in this example, when encoding a subsequent picture, the video encoder 210 may use motion vectors to determine the predicted block in the reference picture without analyzing potential reference pictures and motion vectors. The video encoder 210 may use predicted blocks to encode a particular block.
[0249]
[0258] The transmitting device 102 can handle the encoded video data for subsequent pictures in the same manner as other pictures. Additionally, the receiving device 104 can handle the encoded video data for subsequent pictures in the same manner as other encoded video data. Thus, after the video decoder 224 reconstructs at least a part of a subsequent picture, the picture estimation unit 226 may predict the corresponding part of the picture following the subsequent picture, and the video encoder 228 may encode the video data of the estimated subsequent picture and transmit encoded selection data for the estimated subsequent picture, and this cycle can be repeated. In this way, a part of the burden of encoding video data can be shifted from the video encoder 210 of the transmitting device 102 to the video encoder 228 of the receiving device 104. This can reduce the resource requirements of the transmitting device 102.
[0250]
[0259] The process described with respect to FIG. 21 can be adapted for use with DVC techniques. For example, the transmitting device 102 may apply a decimation pattern to the encoded video data (e.g., according to any of the examples provided elsewhere in this disclosure) such that the transmitting device 102 transmits only some of the encoded video data but still transmits error correction data for the decimated video data. The channel decoder 222 of the receiving device 104 may receive error correction data for a particular picture from the transmitting device (e.g., by the de-puncturing unit 220). In this example, the channel decoder 222 may apply an error correction process to generate error-corrected encoded video data based on the error correction data for a particular picture and the encoded video data for the particular picture generated by the video encoder 228 of the receiving device 104. The video decoder 224 may decode the error-corrected encoded video data to reconstruct particular features.
[0251]
[0260] In some cases, the transmitting device 102 needs to transmit the encoded video data according to a schedule. For example, the transmitting device 102 may need to transmit the encoded video data to the receiving device 104 according to a rate of a predetermined number of pictures per minute to support a specific application. Therefore, a situation may occur where the transmitting device 102 does not receive the encoding selection data for a picture in a timely manner for encoding and transmitting the encoded video data for the picture. Thus, in some examples, based on the determination that the encoding selection data for a picture is not received from the receiving device 104 before the expiration of the time limit, the video encoder 210 may encode the picture without using the encoding selection data for the picture. The video encoder 210 may use a limited video encoding process to encode the picture. Further, in some examples, the encoded video data for a picture may include the encoding selection data generated by the video encoder 210. In some examples, the encoded video data for a picture may include data indicating that the encoded video data for the picture was not generated based on the encoding selection data generated by the receiving device 104. The time limit may depend on or be defined based on the capabilities of the transmitting device 102.
[0252]
[0261] The encoding selection data and the time limit may be defined for each image segment (e.g., slice, region, etc.). Therefore, in the present disclosure, the discussion of the encoding selection data, the encoded video data, or other types of data for a picture may only apply to individual segments of the picture.
[0253]
[0262] Figure 22 is a communication diagram illustrating exemplary data exchange between a transmitting device 102 and a receiving device 104, including the transmission and reception of encoding selection data using the technique of the present disclosure. In the example of Figure 22, the transmitting device 102 transmits encoded video data for picture n-1 to the receiving device 104. The receiving device 104 may reconstruct picture n-1 based on the encoded video data for picture n-1. In addition, the receiving device 104 may estimate and encode picture n based on picture n-1. The receiving device 104 may transmit encoding selection data for picture n to the transmitting device 102. The transmitting device 102 may encode picture n based on the encoding selection data for picture n and transmit the resulting encoded video data for picture n to the receiving device 104. This process may be repeated multiple times. Therefore, in the example of Figure 22, the receiving device 104 may reconstruct picture n based on the encoded video data for picture n, estimate picture n+1 based on picture n and / or one or more other previously reconstructed pictures, encode picture n+1, and send the encoding selection data for picture n+1 to the transmitting device 102.
[0254]
[0263] Figure 23 is a flowchart illustrating exemplary operation of the transmitting device 102 as it receives encoding selection data according to the technique of the present disclosure. In the example of Figure 23, the video encoder 210 encodes a first picture of video data to generate first encoded video data (2300). In some examples, if the transmitting device 102 has not received encoding selection data for the first picture, the transmitting device 102 may perform a limited video encoding process for the first picture. A limited video encoding process may use coding tools that are relatively less computationally intensive than a full video encoding process. For example, a limited video encoding process may use intra-prediction but not inter-prediction.
[0255]
[0264] The transmitting device 102 may transmit the first encoded video data to the receiving device 104 (2302). In some examples, the transmitting device 102 may apply a channeling encoding process to the first encoded video data to generate error correction data for the first encoded video data. The transmitting device 102 may transmit the first encoded video data and the error correction data to the receiving device 104.
[0256]
[0265] Next, the transmitting device 102 may receive coding selection data for a second picture of the video data from the receiving device 104 (2304). The coding selection data may indicate the coding selection used to encode an estimate of the second picture. The second picture is after the first picture in the decoding order. In some examples, the second picture may be before or after the first picture in the output codeder. In some examples, the coding selection data is entropy coded. Thus, in such examples, the transmitting device 102 may entropy decode the coding selection data. For example, the transmitting device 102 may apply CABAC decoding, Golomb-Rice decoding, or another type of entropy decoding to the coding selection data. In some examples, the coding selection data is channel coded. Thus, the transmitting device 102 may apply error correction operations to the coding selection data based on error correction data for the coding selection data.
[0257]
[0266] The video encoder 210 of the transmitting device 102 may encode a second picture based on encoding selection data to generate second encoded video data (2306). For example, the encoding selection data may include data indicating how to partition a particular macroblock into CUs. In this example, the video encoder 210 may partition the macroblock into CUs in the manner indicated by the encoding selection data. In another example, the encoding selection data may indicate an intra-prediction mode for a block (e.g., CU or PU), and the video encoder 210 may use the indicated intra-prediction mode to encode the block. Thus, in this example, the encoding selection data received from the receiving device 104 may include intra-prediction parameters for blocks of the second picture, and the video encoder 210 may perform intra-prediction based on the intra-prediction parameters for blocks of the second picture to generate predicted blocks as part of encoding the second picture. The second encoded video data may include encoded video data based on predicted blocks.
[0258]
[0267] In another example, the encoding selection data received from the receiving device 104 includes motion parameters for a block of the second picture, and the transmitting device 102 may perform motion compensation based on the motion parameters for the block of the second picture to generate a predicted block as part of encoding the second picture. The second encoded video data includes encoded video data based on the predicted block.
[0259]
[0268] The transmitting device 102 may transmit second encoded video data to the receiving device (2308). In some examples, the second encoded video data does not include encoding selection data indicating the encoding selection used by the transmitting device 102 when encoding the second picture, or the encoding selection used by the receiving device 104 when encoding an estimate of the second picture. Since the receiving device 104 generates encoding selection data and therefore already has encoding selection data, it may not be necessary for the second encoded video data to include encoding selection data.
[0260]
[0269] In some examples, the operation shown in Figure 23 may be used in conjunction with the DVC technique. For example, transmitting device 102 may receive encoding selection data for a third picture of video data. The encoding selection data for the third picture may indicate the encoding selection used to encode an estimate of the third picture. Video encoder 210 may encode the third picture based on the encoding selection data for the third picture in order to generate the third encoded video data. Channel encoder 212 may apply a channel encoding process to generate error correction data for the third encoded video data. Transmitting device 102 may transmit the error correction data for the third encoded video data to the receiving device without transmitting at least a portion of the third encoded video data.
[0261]
[0270] Figure 24 is a flowchart illustrating exemplary operation of a receiving device 104, in which the receiving device 104 transmits encoded selection data in accordance with the technique of the present disclosure. In the example of Figure 24, the receiving device 104 may receive first encoded video data from the transmitting device (2400).
[0262]
[0271] The video decoder 224 of the receiving device 104 may reconstruct a first picture of the video data based on the first encoded video data (2402). The picture estimation unit 226 of the receiving device 104 may estimate a second picture of the video data based on the first picture (2404). The second picture may be a picture that comes after the first picture in the decoding order.
[0263]
[0272] The video encoder 228 of the receiving device 104 may generate coding selection data for the estimated second picture (2406). The coding selection data indicates the coding selection used to encode the estimated second picture. For example, as part of encoding the estimated second picture, the video encoder 228 may perform motion compensation based on motion parameters for blocks of the second picture to generate predicted blocks. In this example, the coding selection data may include motion parameters for one or more processor blocks of the second picture. In some examples, as part of encoding the second picture, the video encoder 228 may perform intra-prediction based on intra-prediction parameters for blocks of the second picture to generate predicted blocks. In this example, the coding selection data may include intra-prediction parameters for blocks of the second picture.
[0264]
[0273] The receiving device 104 may transmit coded selection data for a second picture to the transmitting device 102 (2408). In some examples, the receiving device 104 may apply entropy coding (e.g., CABAC coding, Golomm-Rice coding, etc.) to the coded selection data for the second picture before transmitting it. In some examples, the receiving device 104 may perform a channel coding process on the coded selection data to generate error correction data for the coded selection data. The receiving device 104 may transmit the coded selection data and the error correction data for the coded selection data to the transmitting device 102. In some examples, the communication interface 134 of the receiving device 104 (Figure 1) may modulate the coded selection data at a lower modulation order compared to other data transmissions in the data link between the receiving device 104 and the transmitting device 102 (e.g., wireless sidelink channel 112, wireless uplink / downlink channel, etc.). This may increase the likelihood that the transmitting device 102 will correctly receive the coded selection data.
[0265]
[0274] Next, the receiving device 104 may receive second encoded video data from the transmitting device (2410). The video decoder 224 may reconstruct a second picture based on the second encoded video data (2412). In some examples, the second encoded video data does not include encoding selection data. The video decoder 224 may apply a decoding process that includes using encoding selection data to reconstruct a second picture based on the second encoded video data.
[0266]
[0275] The process in Figure 24 can be used in conjunction with the DVC technique. For example, the picture estimation unit 226 may estimate a third picture of video data based on one or more of the first or second pictures. The video encoder 228 may encode the estimated third picture to generate third encoded video data. The receiving device 104 may transmit third encoding selection data to the transmitting device 102. The third encoding selection data may indicate the encoding selection used to encode the estimated third picture. The receiving device 104 may then receive error correction data for the third picture. The channel decoder 222 may apply an error correction process to generate error-corrected encoded video data for the third picture based on the error correction data for the third picture and the third encoded video data. The video decoder 224 may apply a decoding process to reconstruct the third picture based on the error-corrected encoded video data for the third picture. In some examples where the transmitting device 102 and receiving device 104 use the DVC technique, the error-corrected video data for the third picture does not include the third encoding selection data. However, the video decoder 224 may apply a decoding process using the third encoding selection data generated by the video encoder 228 of the receiving device 104 to reconstruct the third picture based on the error-corrected encoded video data for the third picture.
[0267]
[0276] Figure 25 is a conceptual diagram illustrating an exemplary hierarchy of encoded video data using the techniques of this disclosure. More specifically, Figure 25 shows the hierarchy of encoded video data generated using the H.264 / AVC video coding standard. As illustrated in the example in Figure 25, the network abstraction layer (NAL) is the highest level of hierarchy. At the network abstraction layer, data is organized into NAL units. In some examples, NAL units are assigned to different packets or coding blocks for transmission. Network abstraction layer NAL units may include sequence parameter sets (SPSs) and picture parameter sets (PPSs) containing high-level syntax. Network abstraction layer NAL units may also include video coding layer (VCL) NAL units. VCL NAL units may include slice NAL units containing slice-level data. A slice may be a series of macroblocks within a picture. A slice may include instantaneous decoder refresh (IDR) slices and regular slices. Decoding of an IDR slice is independent of other slices. Regular slices may have dependencies on other slices.
[0268]
[0277] Each slice NAL unit may contain a slice header and slice data. The slice header of a slice NAL unit contains information for decoding the slice data of the slice NAL unit. The slice data of a slice NAL unit contains a set of macroblocks (MBs). Skip instructions may be scattered between MBs. Each MB contains encoded video data for a particular block of the slice. Furthermore, as shown in Figure 25, an MB may contain a type indicator, prediction information, coding block pattern, quantization parameters (QP), and encoded residual data. If an MB is encoded using intra-prediction, the prediction data may indicate one or more intra-modes used to encode the MB. If an MB is encoded using inter-prediction, the prediction data may indicate one or more reference pictures and one or more motion vectors. The encoded residual data for an MB may include encoded residual data for the rumor block in the MB, encoded residual data for the Cb block in the MB, and encoded residual data for the Cr block in the MB. Generally, the encoded residual data is the largest portion of the encoded video data.
[0269]
[0278] According to the techniques of this disclosure, anything in the hierarchy of the macroblock layer, excluding the encoded residual data, can be encoded selective data. Thus, in some examples, the receiving device 104 may transmit type data, prediction data, coded block pattern, and QP to the transmitting device 102 for each MB of the estimated picture. Furthermore, in some examples, the transmitting device 102 may transmit only the encoded residual data of the MB to the receiving device 104, and not the type data, prediction data, coded block pattern, or QP of the MB. In some examples, the encoded selective data transmitted by the receiving device 104 may include slice header data, SPS data, and PPS data. The transmitting device 102 and the receiving device 104 may exchange information or may be pre-configured with information indicating encoder and decoder capabilities.
[0270]
[0279] In some examples, the transmitting device 102 may transmit data in addition to encoded residual data for several pictures, several MBs, or several slices. The transmitting device 102 may signal information (e.g., bits) at the network abstraction layer (e.g., as a picture-level control field) that may indicate whether a DVC-based technique is used or whether normal compression should be used for a particular picture.
[0271]
[0280] Figure 26 is a block diagram showing exemplary alternative components of the transmitting device 102 using one or more techniques of the present disclosure. In the example of Figure 26, the transmitting device 102 performs digital and analog coding on the video data. The transmitting device 102 transmits the digitally coded video data and the analogously coded video data to the receiving device 104 via channel 230.
[0272]
[0281] In the example of FIG. 26, it includes a video encoder 2600, a residual generation unit 2602, an analog encoder 2604, a reliability sorting unit 2606, an interleaving unit 2608, a channel encoder 2610, and a puncturing unit 2612. The video encoder 2600 can obtain video data and operate in substantially the same manner as the video encoder 210 of FIG. 2. As shown in the example of FIG. 26A, the video encoder 2600 can receive the value of an encoded parameter (e.g., an encoding selection parameter) sent by the receiving device 104. In some examples, the video encoder 2600 can send the value of the encoded parameter and / or the encoding selection data to the receiving device 104. In this way, the video encoder 2600 of the transmitting device 102, the video encoder of the receiving device 104, and the video decoder of the receiving device 104 can operate based on the same value of the encoded parameter.
[0273]
[0282] The video encoder 2600 can also output prediction data to the residual generation unit 2602. In addition, the video encoder 2600 can apply a higher level of quantization than the video encoder 210. The residual generation unit 2602 can generate residual data based on the prediction data and the video data.
[0274]
[0283] The analog encoder 2604 can perform analog coding operations on residual data. Exemplary details of analog coding operations can be found in U.S. Patent No. 11,553,184, entitled "Hybrid Digital-Analog Modulation for Transmission of Video Data," filed December 29, 2020; U.S. Patent No. 11,431,962, entitled "Analog Modulated Video Transmission with Variable Symbol Rate," filed December 29, 2020; and U.S. Patent No. 11,457,224, entitled "Interlaced Coefficients in Hybrid Digital-Analog Modulation for Transmission of Video Data," filed December 29, 2020.
[0275]
[0284] For example, in some cases, the analog encoder 2604 may generate coefficients based on residual data. For example, the analog encoder 2604 may generate coefficients by binarizing the residual data. The analog encoder 2604 may quantize the coefficients. In other examples where coefficients are generated based on video data, the analog encoder 2604 may perform more, fewer, or different steps. For example, in some cases, the analog encoder 2604 does not perform the quantization step. In yet another example, the analog encoder 2604 does not perform the step of binarizing the residual data.
[0276]
[0285] Furthermore, the analog encoder 2604 can generate coefficient vectors. Each coefficient vector contains n of the coefficients. The analog encoder 2604 can generate coefficient vectors in one of several ways. For example, in one example, the analog encoder 2604 can generate coefficient vectors as groups of n consecutive coefficients according to a coefficient coding order. Various coefficient coding orders can be used, such as raster scan order, zigzag scan order, inverse raster scan order, and vertical scan order. In some examples, the coefficient vectors may contain one or more negative coefficients and one or more positive coefficients (i.e., signed coefficients). In some examples, the coefficient vectors may contain only non-negative coefficients (i.e., unsigned coefficients).
[0277]
[0286] For each of the coefficient vectors, the analog encoder 2604 can determine the amplitude value for the coefficient vector based on the mapping pattern. For each of the allowed coefficient vectors among the multiple allowed coefficient vectors, the mapping pattern maps each allowed coefficient vector to each amplitude value among the multiple amplitude values. Each amplitude value is adjacent in n-dimensional space to at least one other amplitude value among the multiple amplitude values adjacent to it on the monotonic number line of amplitude values.
[0278]
[0287] In some examples, the analog encoder 2604 may determine a position in n-dimensional space in order to determine the amplitude value for a coefficient vector. The coordinates of the position in n-dimensional space are based on the coefficients of the coefficient vector, and the mapping pattern maps different positions in n-dimensional space to different amplitude values among multiple amplitude values. The analog encoder 2604 may determine the amplitude value for a coefficient vector as the amplitude value corresponding to the determined position in n-dimensional space.
[0279]
[0288] The analog encoder 2604 can modulate an analog signal based on amplitude values for a coefficient vector. For example, the analog encoder 2604 can determine an analog symbol based on a pair of amplitude values. The analog symbol may correspond to the phase shift and power of a point in the IQ plane having coordinates indicated by the amplitude value pair. Based on the determined phase shift and power, the analog encoder 2604 can modulate the analog signal at the moment of symbol sampling. The modem of the transmitting device 102 (e.g., the communication interface 118) may be configured to output an analog signal.
[0280]
[0289] Furthermore, in the example of Figure 26A, the reliability sorting unit 2606 may acquire encoded video data generated by the video encoder 2600. The reliability sorting unit 2606 may acquire reliability-side information from the video encoder 2600. In some examples, the reliability sorting unit 2606 may receive values for channel and compression state feedback (CCSF) parameters. The values of the CCSF parameters may provide information about the state of channel 230 (e.g., signal-to-noise ratio, latency time, network bandwidth congestion, etc.). In some examples, the values of the CCSF parameters may provide information about the reliability and quality of the prediction. For example, the values of the CCSF parameters may include a decimation pattern indicator. In some examples, the values of the CCSF parameters may allow the channel encoder 2610 to determine the decimation pattern.
[0281]
[0290] The interleaving unit 2608 can perform an interleaving process that can ensure that trusted bits and untrusted bits are equally distributed across code blocks. For example, encoded video data may be divided into code blocks. The channel encoder 2610 may generate a separate set of error correction data for each code block. Before the channel encoder 2610 generates the error correction data, the interleaving unit 2608 may interleave the encoded video data across code blocks according to a predetermined interleaving pattern. For example, encoded video data representing different adjacent pixels may be interleaved into different code blocks. The deinterleaving process performed by the receiving device 104 reverses the interleaving process after the channel decoding process is applied. Therefore, if one of the code blocks is corrupted during transmission, pixels decoded from the corrupted code block may be spatially distributed within the picture among pixels decoded from the uncorrupted code block.
[0282]
[0291] The channel encoder 2610 of the transmitting device 102 may perform a channel coding process on the video data acquired from the interleaving unit 2608. The channel encoder 2610 may perform a channel coding process according to any of the examples provided with respect to the channel encoder 212 (Figure 2). The puncturing unit 2612 may perform a bit puncturing operation on the error correction data generated by the channel encoder 2610. The puncturing unit 2612 may perform a bit puncturing operation on the error correction data according to any of the examples provided with respect to the puncturing unit 214 (Figure 2). The transmitting device 102 may transmit the coded video data and error correction data (e.g., bit-punctured error correction data) to the receiving device 104 via channel 230.
[0283]
[0292] Figure 27 is a block diagram showing exemplary alternative components of the receiving device 104 according to one or more techniques of the present disclosure. The version of the receiving device 104 shown in Figure 27 may be compatible with the version of the transmitting device 102 shown in Figure 26. In the example of Figure 27, the receiving device 104 includes an analog decoder 2700, a depunching unit 2702, a channel decoder 2704, a deinterleaving unit 2706, a video decoder 2708, a reconstruction unit 2710, a picture estimation unit 2712, a video encoder 2714, a reliability unit 2716, and a feedback unit 2718.
[0284]
[0293] The analog decoder 2700 can acquire analog-encoded video data. The analog decoder 2700 can perform analog decoding operations to reconstruct residual data. Exemplary details of analog decoding operations can be found in U.S. Patents 11,553,184, 11,431,962, and 11,457,224.
[0285]
[0294] For example, in some cases, the analog decoder 2700 may determine amplitude values for multiple coefficient vectors based on an analog signal. For example, the analog decoder 2700 may determine the instantaneous phase shift and power of the symbol sampling of the analog signal. The analog decoder 2700 may determine a point in the IQ plane represented by the determined phase shift and power. The analog decoder 2700 may then determine a pair of amplitude values as the coordinates of the point in the IQ plane.
[0286]
[0295] For each coefficient vector, the analog decoder 2700 may determine the coefficients in the coefficient vector based on the amplitude value and mapping pattern for the coefficient vector. For each of the allowed coefficient vectors among the multiple allowed coefficient vectors, the mapping pattern may map each allowed coefficient vector to each amplitude value among the multiple amplitude values. Each amplitude value is adjacent in n-dimensional space to at least one other amplitude value among the multiple amplitude values adjacent to it on the monotonic number line of amplitude values. Each coefficient vector may contain n coefficients. The value n can be 2 or greater. In some examples, the analog decoder 2700 may determine the coefficients in the coefficient vector as coordinates of positions in n-dimensional space corresponding to amplitude values. The mapping pattern maps different positions in n-dimensional space to different amplitude values among the multiple amplitude values. In some examples, the coefficient vector may contain one or more negative coefficients and one or more positive coefficients. In other examples, the coefficient vector may contain only non-negative coefficients.
[0287]
[0296] In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain a sign value, which indicates the positive / negative sign of the coefficient in the coefficient vector. In such examples, the analog decoder 2700 may determine the absolute value of the coefficient in the coefficient vector based on the amplitude value and mapping pattern for the coefficient vector. The analog decoder 2700 may, at least in part, reconstruct the coefficient in the coefficient vector by applying the sign value to the absolute value of the coefficient in the coefficient vector. In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain data representing a shift value. In such examples, the shift value indicates the coefficient with the most negative value among the coefficients in the coefficient vector. In addition, in such examples, the analog decoder 2700 may determine the midpoint of the coefficient in the coefficient vector based on the amplitude value and mapping pattern for the coefficient vector. The analog decoder 2700 may, at least in part, reconstruct the coefficient in the coefficient vector by adding the shift value to each of the midpoints of the coefficient in the coefficient vector.
[0288]
[0297] Furthermore, the analog decoder 2700 can generate residual data based on the coefficients in the coefficient vector. For example, in one example, the analog decoder 2700 can inverse quantize the coefficients of the coefficient vector. In this example, the analog decoder 2700 can perform a de-binarization process to convert the coefficients into digital sample values. For example, the analog decoder 2700 can apply an inverse DCT to the coefficients to convert them into digital sample values. In this way, the analog decoder 2700 can generate digital residual sample values.
[0289]
[0298] The depunching unit 2702 may obtain encoded video data and bit-punctured error-corrected data. The depunching unit 2702 may apply a depunching process to the bit-punctured error-corrected data in order to reconstruct the error-corrected data. The depunching unit 2702 may apply a depunching process according to any of the examples provided elsewhere in this disclosure with respect to the depunching unit 220 in Figure 2.
[0290]
[0299] The channel decoder 2704 may perform a channel decoding process to correct the encoded video data (for example, encoded video data received via channel 230, or encoded video data generated by the video encoder 2714 and, in some examples, corrected by the reliability unit 714) based on error correction data. The channel decoder 2704 may perform a channel decoding process according to any of the examples provided elsewhere in this disclosure with respect to the channel decoder 222 in Figure 2.
[0291]
[0300] The deinterleaving unit 2706 may perform a deinterleaving operation on the error-corrected encoded video data generated by the channel decoder 2704. For example, the deinterleaving process may be the reverse of the interleaving process performed by the interleaving unit 2608 of the transmitting device 102. For example, the deinterleaving process may perform the deinterleaving process according to the interleaving pattern used by the interleaving unit 2608.
[0292]
[0301] The video decoder 2708 may acquire encoded video data (for example, deinterleaved encoded video data generated by the deinterleaving unit 2706). The video decoder 2708 may perform a video decoding process on the encoded video data to reconstruct the picture of the video data. The video decoding process performed by the video decoder 2708 may be the same as that described in any of the examples provided elsewhere in this disclosure with respect to the video decoder 224. The reconstruction unit 2710 of the receiving device 104 may add the residual data generated by the analog decoder 2700 to the corresponding samples of the reconstructed video data generated by the video decoder 2708, thereby completely reconstructing the picture of the video data.
[0293]
[0302] Furthermore, in the example of Figure 27, the picture estimation unit 2712 may estimate one or more pictures based on a previously reconstructed picture. As previously mentioned in this disclosure, the description of a picture may apply to a segment of a picture, such as a slice. The picture estimation unit 2712 may estimate a picture according to any of the examples provided elsewhere in this disclosure with respect to the picture estimation unit 226. The video encoder 2714 may perform a video encoding process on the estimated picture. As part of performing the video encoding process, the video encoder 2714 may determine encoding selection data, such as encoding selection data 2100, as discussed previously. The receiving device 104 may transmit the encoding selection data to the transmitting device 102. In some examples, the video encoder 2714 may also transmit encoding parameters, such as those discussed with respect to Figures 4 and 5, to the transmitting device 102, and the video encoder 2600 of the transmitting device 102 may perform a limited encoding process in the same manner as the video encoder 2714 of the receiving device 104. In some examples, the video encoder 2714 may determine a multiview coding queue and send it to the transmitting device 102.
[0294]
[0303] Reliability unit 2716 may operate in much the same way as reliability unit 1002 (Figure 10). Based on the output of reliability unit 1002, feedback unit 2718 may send predictive quality feedback (e.g., CCSF parameters) to transmission device 102.
[0295]
[0304] In some examples of this disclosure, a video encoder 2600 of a transmitting device 102 may generate prediction data for a set of pictures, and a residual generation unit 2602 may generate residual data based on the first prediction data and the first set of pictures. The video encoder 2600 may apply a transform to the prediction data to generate a transform block, quantize the transform coefficients of the transform block, and apply entropy coding to the syntax elements representing the quantized transform coefficients to generate entropy-coded syntax elements. The first encoded video data may include the entropy-coded syntax elements. A channel encoder 2610 may perform a channel coding process to generate error correction data for the encoded video data, which includes the entropy-coded syntax elements. An analog encoder 2604 may perform analog modulation on the residual data to generate analog-modulated residual data. A communication interface of the transmitting device 102 may transmit the analog-modulated residual data, the error correction data, and the encoded video data.
[0296]
[0305] The communication interface of the receiving device 104 can receive analog-modulated residual data and error-corrected data and decimated video data from the transmitting device. The decimated video data may include encoded video data to which a decimation pattern has been applied. The encoded video data is generated based on a set of pictures of video data. The channel decoder 2704 may apply an error correction process to generate error-corrected encoded video data based on the encoded video data and error-corrected data. The error-corrected encoded video data includes entropy-coded syntax elements representing quantized transformation coefficients. The video decoder 2708 may perform a decoding process to reconstruct a second set of pictures based on the error-corrected encoded video data. As part of applying the decoding process to reconstruct the set of pictures, the video decoder 2708 may apply entropy decoding to the syntax elements to obtain quantized transformation coefficients, dequantize the quantized transformation coefficients to generate inversely quantized transformation coefficients, and apply an inverse transform to the inversely quantized transformation coefficients to generate prediction data. The analog decoder 2700 can demodulate the analog-modulated residual data to obtain the residual data. The reconstruction unit 2710 can reconstruct the set of pictures based on the prediction data and the residual data.
[0297]
[0306] The following is a non-exclusive list of the techniques used in this disclosure to define the provisions of one or more of these provisions.
[0298]
[0307] Clause 1A. A method for decoding video data, comprising: receiving error correction data from a transmitting device in a receiving device, which provides error correction information and is generated based on encoded video data of one or more blocks of picture in the video data; generating prediction data for picture, which includes predictions of blocks of picture based at least in part on blocks of picture that have been previously reconstructed in one or more video data, using one or more coding tools not used to generate encoded video data of one or more blocks in the receiving device; generating encoded video data based on the prediction data for picture in the receiving device; generating error-corrected encoded video data using the error correction data in order to perform an error correction operation on the encoded video data in the receiving device; and performing a reconstruction operation in the receiving device for reconstructing blocks of picture based on the error-corrected encoded video data, wherein the reconstruction operation is controlled by the values of one or more parameters.
[0299]
[0308] Clause 2A. The method of Clause 1A, further comprising receiving a parameter value from a transmitting device in a receiving device.
[0300]
[0309] Clause 3A. The method of Clause 1A, further comprising determining the value of a parameter in a receiving device without receiving the value of the parameter from a transmitting device.
[0301]
[0310] Clause 4A. Any method of Clauses 1A to 3A, wherein the parameter includes one or more quantization parameters, generating encoded video data includes using the quantization parameters to quantize transformation coefficients generated based on prediction data for the picture, and performing a reconstruction operation includes using the quantization parameters to dequantize the transformation coefficients of the error-corrected encoded video data.
[0302]
[0311] Clause 5A. The method of Clause 4A, further comprising calculating the quantization parameter based on the entropy ratio of the quantized transformation coefficients to the unquantized transformation coefficients.
[0303]
[0312] Clause 6A. Any method of Clauses 1A to 5A, wherein the parameter includes a transformation size parameter, generating encoded video data includes applying a forward transformation to sample region data for a picture having a transformation size indicated by the transformation size parameter, and performing a reconstruction operation includes applying an inverse transformation to the transformation coefficients of the error-corrected encoded video data having a transformation size indicated by the transformation size parameter.
[0304]
[0313] Clause 7A. Any method of Clauses 1A to 6A, wherein the parameter includes a parameter indicating the amount of transformation coefficients, generating encoded video data includes including a set of transformation coefficients in the encoded video data that includes the indicated amount of transformation coefficients, and performing a reconstruction operation includes analyzing a set of transformation coefficients from the error-corrected encoded video data that includes the indicated amount of transformation coefficients.
[0305]
[0314] The method of the provisions provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of provisions of the provision of
[0306]
[0315] Clause 9A. Any method of Clauses 1A to 8A, wherein the parameter includes a bit width parameter for a plurality of index values, and for each of the plurality of index values, the reconstruction operation includes analyzing a first set of bits from error-corrected encoded video data, wherein the first set of bits represents a conversion coefficient having each index value, and the amount of bits in the first set of bits is equal to the bit width indicated by the bit width parameter for each index value, thereby generating encoded video data, which includes including a second set of bits in the encoded video data, wherein the second set of bits represents a conversion coefficient having each index value, and the amount of bits in the second set of bits is equal to the bit width indicated by the bit width parameter for each index value, and analyzing a third set of bits from error-corrected encoded video data, wherein the third set of bits represents a conversion coefficient having each index value, and the amount of bits in the third set of bits is equal to the bit width indicated by the bit width parameter for each index value.
[0307]
[0316] Any method of the provisions of 1A to 9A, further comprising performing a bit depuncturing operation on error-corrected data before generating error-corrected encoded video data.
[0308]
[0317] Clause 11A. Any method of Clauses 1A to 10A, wherein the parameter includes one or more of the following: color space, conversion size, quantization parameter, number of conversion coefficients in the first encoded video data, or number of bits per conversion coefficient in the first encoded video data.
[0309]
[0318] Clause 12A. A method of any of Clauses 1A to 11A, wherein the decimation pattern defines a pattern of anchored and non-anchored converted blocks in a picture, and the method further comprises, in a receiving device, receiving systematic bits of anchored converted blocks rather than systematic bits of non-anchored converted blocks, the systematic bits of anchored converted blocks representing conversion coefficients in the anchored converted blocks, the systematic bits of non-anchored converted blocks representing bit-depth reduced versions of the original conversion coefficients in the non-anchored converted blocks, and the error correction data comprises error correction data for anchored converted blocks and error correction data for non-anchored converted blocks, and the method of generating error-corrected encoded video data based on the original conversion coefficients in the non-anchored converted blocks comprises using error correction data for anchored converted blocks to perform error correction on the systematic bits of anchored converted blocks and using error correction data for non-anchored converted blocks to perform error correction on the portion of encoded video data corresponding to the non-anchored converted blocks.
[0310]
[0319] Clause 13A. The method of Clause 12A, further comprising determining a decimation pattern in a receiving device and sending the decimation pattern to a transmitting device in a receiving device.
[0311]
[0320] Clause 14A. Any method of Clauses 1A to 13A, wherein the decimation pattern defines a pattern of anchored and non-anchored transformed blocks in a picture, and the transformed coefficients in the non-anchored transformed blocks have a bit depth that is reduced relative to that of the anchored transformed blocks, and the receiving device receives the systematic bits of the anchored transformed blocks, the systematic bits of the non-anchored transformed blocks, and the correlation matrix, wherein the systematic bits of the anchored transformed blocks represent the transformed coefficients in the anchored transformed blocks, and the systematic bits of the non-anchored transformed blocks represent a bit-depth reduced version of the original transformed coefficients in the non-anchored transformed blocks, and the receiving device performs a reconstruction operation which includes, for each non-anchored transformed coefficient in the non-anchored transformed block, the receiving device calculates an interpolated value of the non-anchored transformed coefficient based on the correlation matrix and the corresponding anchored transformed coefficient, and the receiving device calculates a reconstructed value of the non-anchored transformed coefficient based on the interpolated value of the non-anchored transformed coefficient and the value of the non-anchored transformed coefficient in the error-corrected encoded video data.
[0312]
[0321] Clause 15A. A method for encoding video data, comprising: a transmitting device acquiring video data from a video source; a transmitting device generating encoded video data of a first picture and encoded video data of a second picture of the video data based on a set of parameters; a transmitting device performing channel coding on the encoded video data of the first picture and the encoded video data of the second picture in order to generate error correction data for the first picture and error correction data for the second picture; and a transmitting device transmitting the encoded video data of the first picture, the error correction data for the first picture, and the error correction data for the second picture.
[0313]
[0322] Clause 16A. The method of Clause 15A, further comprising transmitting a parameter value to a receiving device in a transmitting device.
[0314]
[0323] The method of Clause 16A, wherein the parameter includes one or more quantization parameters, and generating encoded video data involves using the quantization parameters to quantize the transformation coefficients of a first picture and a transformation coefficient of a second picture.
[0315]
[0324] Clause 18A. The method of generating encoded video data, wherein the parameter includes a transformation size parameter, comprises applying a forward transformation to blocks of residual data for a first picture and blocks of residual data for a second picture, wherein the forward transformation has a transformation size indicated by the transformation size parameter, in any of the methods of Clauses 16A to 17A.
[0316]
[0325] Clause 19A. Any method of Clauses 16A to 18A, wherein the parameter includes a parameter indicating the amount of a conversion coefficient, and generating encoded video data of a first picture and encoded video data of a second picture includes including a set of conversion coefficients in the encoded video data of the first picture and encoded video data of the second picture, the amount of the conversion coefficient indicated.
[0317]
[0326] Clause 20A. Any method of Clauses 16A to 10A, wherein the parameter includes one or more of the following: color space, conversion size, quantization parameter, number of conversion coefficients in the encoded video data, or number of bits per conversion coefficient in the encoded video data.
[0318]
[0327] Clause 21A. A method for encoding video data, comprising: a transmitting device obtaining video data from a video source; a transmitting device generating transformation blocks based on the video data; a transmitting device determining which of the transformation blocks are anchor transformation blocks; a transmitting device calculating a correlation matrix for the set of transformation blocks; a transmitting device generating a bit-reduced non-anchor transformation matrix; and a transmitting device transmitting the anchor transformation blocks, the non-anchor transformation blocks, and the correlation matrix to a receiving device.
[0319]
[0328] Clause 22A. The method of Clause 21A, further comprising the transmitting device receiving instructions for a decimation pattern from the receiving device.
[0320]
[0329] Clause 23A. A device comprising a memory configured to store video data, a communication interface, and one or more processes implemented in the circuit and coupled to the memory, wherein one or more processors are configured to perform any of the methods of Clauses 1A to 22A.
[0321]
[0330] Clause 24A. A device including means for performing any of the methods described in Clauses 1A through 22A.
[0322]
[0331] Clause 25A. A computer-readable data storage medium storing instructions, wherein, when the instructions are executed, causes a device to perform any of the methods described in Clauses 1A through 22A.
[0323]
[0332] Clause 1B. A device for processing video data, comprising: a memory configured to store video data; a communication interface configured to receive error correction data from a transmitting device, which provides error correction information relating to pictures of video data; and one or more processes implemented in the circuit and coupled to the memory, wherein one or more processors are configured to generate prediction data for pictures, which includes predictions of blocks of pictures based at least in part on one or more previously reconstructed pictures of video data; generate encoded video data, which includes encoded blocks including transform coefficients, based on the prediction data for pictures; scale the bits of the transform coefficients of the transform blocks based on confidence values for bit positions; generate error-corrected encoded video data using the error correction data to perform error correction operations on the scaled bits of the transform coefficients of the transform blocks; and reconstruct a picture based on the error-corrected encoded video data.
[0324]
[0333] Clause 2B. The device of Clause 1B, further configured with one or more processors to generate a reliability value in the receiving device.
[0325]
[0334] Clause 3B. A device of Clause 2B in which one or more processors are configured to generate a confidence value based on statistics regarding the occurrence of errors at bit locations.
[0326]
[0335] Clause 4B. A device according to either Clause 2B or 3B, in which one or more processors are configured to generate reliability values based on reliability characteristics for individual regions of a picture of video data.
[0327]
[0336] Clause 5B. A device according to any of Clauses 2B through 4B, in which one or more processors are configured to generate confidence values based on a noise model.
[0328]
[0337] Clause 6B. Any device under Clauses 1B through 5B, whose communication interface is further configured to send reliability values to a transmitting device.
[0329]
[0338] Clause 7B. Any device of Clauses 1B through 5B, whose communication interface is further configured to receive reliability values from a transmitting device.
[0330]
[0339] Clause 8B. A device for processing video data, the method comprising: a memory configured to store video data; one or more processes implemented in a circuit and coupled to the memory, wherein one or more processors are configured to acquire video data, acquire predictive quality feedback, which is based on the reliability of estimated pictures produced by a receiving device, adapt one or more video coding parameters or channel coding parameters based on the predictive quality feedback, and to perform a video coding process, controlled by video coding parameters, on the encoded video data to generate encoded video data based on one or more pictures of the acquired video data; and one or more processors configured to perform a channel coding process, controlled by channel coding parameters, on the encoded video data to generate channel coded data; and a communication interface configured to transmit channel coded data to a receiving device.
[0331]
[0340] A device of Clause 9B, wherein the video coding parameters include quantization parameters, and one or more processors are configured to adapt the quantization parameters as part of adapting the video coding parameters, and one or more processors are configured to use the quantization parameters to quantize the conversion coefficients of one or more picture conversion blocks as part of performing the video coding process.
[0332]
[0341] A device according to any of the clauses 8B to 9B, wherein the channel coding parameters include a low-density parity check (LDPC) graph, one or more processors are configured to adapt the LDPC graph as part of adapting the channel coding parameters, and one or more processors are configured to use the LDPC graph to generate codewords to be included in the channel-coded data as part of performing the channel coding process.
[0333]
[0342] Clause 11B. A device according to any of Clauses 8B to 10B, wherein the channel-encoded data includes error-corrected data, one or more processors are further configured to adapt one or more bit-puncturing parameters based on predictive quality feedback, one or more processors are configured to perform a bit-puncturing process on the error-corrected data, and the bit-puncturing process is controlled by one or more bit-puncturing parameters.
[0334]
[0343] Clause 12B. A method for processing video data, comprising: receiving error correction data from a transmitting device in a receiving device, which provides error correction information relating to pictures in the video data; generating prediction data for pictures in a receiving device, which includes predictions for blocks of pictures based at least in part on pictures previously reconstructed in one or more video data; generating encoded video data in a receiving device based on the prediction data for pictures, which includes encoded video data including a transformation block containing transformation coefficients; scaling the bits of the transformation coefficients of the transformation block in a receiving device based on a confidence value for bit positions; generating error-corrected encoded video data in a receiving device using the error correction data to perform error correction operations on the scaled bits of the transformation coefficients of the transformation block; and reconstructing a picture in a receiving device based on the error-corrected encoded video data.
[0335]
[0344] Clause 13B. The method of Clause 12B, further comprising generating a reliability value in a receiving device.
[0336]
[0345] Clause 14B. The method of Clause 13B, wherein generating a reliability value includes generating a reliability value in a receiving device based on statistics regarding the occurrence of errors in bit positions.
[0337]
[0346] Clause 15B. Any method of Clause 13B or 14B, wherein generating a reliability value in a receiving device includes generating a reliability value based on the reliability characteristics of individual areas of a picture of video data.
[0338]
[0347] Clause 16B. Generating a reliability value in any of the methods of Clauses 13B to 15B, which includes generating a reliability value based on a noise model in a receiving device.
[0339]
[0348] Clause 17B. Any method of Clauses 12B to 16B, further comprising sending a confidence value to the transmitting device at the receiving device.
[0340]
[0349] Clause 18B. Any method of Clauses 12B to 16B, further comprising receiving a confidence value from a transmitting device in a receiving device.
[0341]
[0350] Clause 19B. A method for processing video data, comprising: acquiring video data; acquiring predictive quality feedback, which is based on the reliability of estimated pictures produced by a receiving device; adapting one or more video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process, controlled by video coding parameters, to generate coded video data based on one or more pictures of the acquired video data; performing a channel coding process, controlled by channel coding parameters, on the coded video data to generate channel coded data; and transmitting the channel coded data to a receiving device.
[0342]
[0351] Clause 20B. The method of Clause 19B, wherein the video coding parameters include quantization parameters, adapting the video coding parameters includes adapting the quantization parameters, and performing the video coding process includes using the quantization parameters to quantize the conversion coefficients of one or more picture conversion blocks.
[0343]
[0352] Clause 21B. Any method of Clauses 19B to 20B, wherein the channel coding parameters include a low-density parity check (LDPC) graph, adapting the channel coding parameters includes adapting an LPDC graph, and performing the channel coding process includes using the LDPC graph to generate codewords contained in the channel-coded data.
[0344]
[0353] Clause 22B. Any method of Clauses 19B to 21B, wherein channel-encoded data includes error-corrected data, and the method further comprises adapting one or more bit-puncturing parameters based on predictive quality feedback, and performing a bit-puncturing process on the error-corrected data, wherein the bit-puncturing process is controlled by one or more bit-puncturing parameters.
[0345]
[0354] Clause 23B. A device including means for carrying out any of the methods described in Clauses 12B through 22B.
[0346]
[0355] Clause 24B. A computer-readable data storage medium that stores instructions, when executed, causing a device to perform any of the methods described in Clauses 12B through 22B.
[0347]
[0356] Clause 1C. A device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, wherein one or more processors acquire a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint, transmits first encoded video data to a receiving device, the first encoded video data being based on the first set of multiview pictures, and receives a multiview encoding queue from the receiving device. A device configured to obtain a second set of multiview pictures of video data, the second set of multiview pictures including a third picture and a fourth picture, where the third picture is from a first viewpoint and the fourth picture is from a second viewpoint, and to perform a multiview coding process on the second set of multiview pictures, where the multiview coding process reduces interview redundancy between the third picture and the fourth picture, in order to generate second encoded video data based on a multiview coding queue received from a receiving device, and to transmit the second encoded video data to the receiving device.
[0348]
[0357] Clause 2C. A device of Clause 1C, further configured to send a second encoded video data to a receiving device, receive an updated multiview coding queue from the receiving device, obtain a third set of multiview pictures of the video data, the third set of multiview pictures includes a fifth picture and a sixth picture, where the fifth picture is from a first viewpoint and the sixth picture is from a second viewpoint, encode the third set of multiview pictures based on the updated multiview coding queue received from the receiving device to generate a third encoded video data, and send the third encoded video data to the receiving device.
[0349]
[0358] A device of any of the clauses 1C to 2C, in which the multiview coding queue includes one or more motion data for relative shifts between blocks of the first picture and blocks of the second picture, brightness corrections between the first picture and the second picture, interblock shifts between anchor blocks and reconstructed blocks, or reference shifts.
[0350]
[0359] Clause 4C. A device of any of Clauses 1C through 3C, where the device is an augmented reality (XR) headset and one or more processors are configured to receive virtual element data generated from a receiving device based on a first set and a second set of multiview pictures and to output the virtual element data for display in an XR scene.
[0351]
[0360] Clause 5C. A device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, wherein one or more processors receive from a transmitting device first encoded video data, the first encoded video data being based on a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint, and the first A device configured to determine a multiview coding queue based on encoded video data, send the multiview coding queue to a transmitting device, and retrieve from the transmitting device second encoded video data, the second encoded video data being based on a second set of multiview pictures including a third picture and a fourth picture, and the second encoded video data being coded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.
[0352]
[0361] The device of Clause 5C, wherein one or more processors are further configured to decode second encoded video data.
[0353]
[0362] Any device of Clause 7C, wherein the multiview coding queue is a first multiview coding queue, and one or more processors determine a second multiview coding queue based on second coded video data, transmit the second multiview coding queue to a transmitting device, and is further configured to receive from the transmitting device third coded video data, wherein the third coded video data is based on a third set of multiview pictures including a fifth picture and a sixth picture, and the third coded video data is coded using a multiview coding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview coding queue.
[0354]
[0363] A device according to any of the clauses 5C through 7C, wherein the multiview coding queue includes a depth map indicating the depth of objects represented in a first picture and a second picture, and one or more processors are configured to determine the depth map based on the first picture and the second picture as part of determining the multiview coding queue.
[0355]
[0364] A device according to Clause 9C, wherein the multiview coding queue includes one or more illumination compensation coefficients, and one or more processors are configured to determine the illumination compensation coefficients based on a first picture and a second picture as part of determining the multiview coding queue.
[0356]
[0365] Clause 10C. A device of any of Clauses 5C through 9C, wherein the transmitting device is an augmented reality (XR) headset, and one or more processors are further configured to process a second set of pictures to generate virtual element data and transmit the virtual element data to the XR headset.
[0357]
[0366] Clause 11C. A method for processing video data, comprising: obtaining a first set of multiview pictures of video data, wherein the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; transmitting to a receiving device a first encoded video data, wherein the first encoded video data is based on the first set of multiview pictures; receiving a multiview encoding queue from the receiving device; and obtaining a second set of multiview pictures of video data. A method comprising: obtaining a second set of multiview pictures in which the second set of multiview pictures includes a third picture and a fourth picture, where the third picture is from a first viewpoint and the fourth picture is from a second viewpoint; performing a multiview coding process on the second set of multiview pictures, where the multiview coding process reduces interview redundancy between the third picture and the fourth picture, in order to generate second encoded video data based on a multiview coding queue received from a receiving device; and transmitting the second encoded video data to a receiving device.
[0358]
[0367] The method of Clause 12C, further comprising: receiving an updated multiview coding queue from a receiving device after transmitting second encoded video data to a receiving device; obtaining a third set of multiview pictures of video data, wherein the third set of multiview pictures includes a fifth picture and a sixth picture, where the fifth picture is from a first viewpoint and the sixth picture is from a second viewpoint; coding the third set of multiview pictures based on the updated multiview coding queue received from the receiving device to generate third encoded video data; and transmitting the third encoded video data to a receiving device.
[0359]
[0368] Clause 13C. Any method of Clauses 11C to 12C, wherein the multiview coding queue includes one or more motion data for relative shifts between blocks of the first picture and blocks of the second picture, brightness corrections between the first picture and the second picture, interblock shifts between anchor blocks and reconstructed blocks, or reference shifts.
[0360]
[0369] Any method of Clause 11C to 13C further includes receiving virtual element data generated based on a first set and a second set of multiview pictures from a receiving device and outputting virtual element data for display in an augmented reality (XR) scene.
[0361]
[0370] Clause 15C. A method for processing video data, comprising: obtaining from a transmitting device first encoded video data, wherein the first encoded video data is based on a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint; determining a multiview coding queue based on the first encoded video data; transmitting the multiview coding queue to the transmitting device; and obtaining from the transmitting device second encoded video data, wherein the second encoded video data is based on a second set of multiview pictures comprising a third picture and a fourth picture, the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.
[0362]
[0371] The method of Clause 15C, further comprising decoding the second encoded video data.
[0363]
[0372] Any method of Clause 17C, wherein the multiview coding queue is a first multiview coding queue, and the method further comprises determining a second multiview coding queue based on second coded video data, transmitting the second multiview coding queue to a transmitting device, and retrieving from the transmitting device third coded video data, wherein the third coded video data is based on a third set of multiview pictures including a fifth picture and a sixth picture, and the third coded video data is coded using a multiview coding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview coding queue.
[0364]
[0373] Clause 18C. Any method of Clauses 15C to 17C, wherein the multiview coding queue includes a depth map indicating the depth of objects represented in the first and second pictures, and determining the multiview coding queue includes determining the depth map based on the first and second pictures.
[0365]
[0374] Clause 19C. The methods of Clauses 15C to 18C, wherein the multiview coding queue includes one or more illumination compensation coefficients, and determining the multiview coding queue includes determining the illumination compensation coefficients based on a first picture and a second picture.
[0366]
[0375] Clause 20C. Any method of Clauses 15C to 19C, wherein the transmitting device is an augmented reality (XR) headset, and the method further comprises processing a second set of pictures to generate virtual element data and transmitting the virtual element data to the XR headset.
[0367]
[0376] Clause 21C. Means for acquiring a first set of multiview pictures of video data, wherein the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; means for transmitting to a receiving device first encoded video data, wherein the first encoded video data is based on the first set of multiview pictures; means for receiving a multiview encoding queue from a receiving device; and a second set of multiview pictures of video data, wherein the multiview pictures A device comprising: means for acquiring a second set of multiview pictures, wherein the second set of pictures includes a third picture and a fourth picture, where the third picture is from a first viewpoint and the fourth picture is from a second viewpoint; means for performing a multiview coding process on the second set of multiview pictures to generate second coded video data based on a multiview coding queue received from a receiving device, wherein the multiview coding process reduces interview redundancy between the third picture and the fourth picture; and means for transmitting the second coded video data to a receiving device.
[0368]
[0377] Clause 22C. A device comprising: means for obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint; means for determining a multiview coding queue based on the first encoded video data; means for transmitting the multiview coding queue to the transmitting device; and means for obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multiview pictures comprising a third picture and a fourth picture, the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.
[0369]
[0378] Clause 1D. A device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, wherein one or more processors are configured to encode a first set of pictures of video data to generate first encoded video data, transmit the first encoded video data to a receiving device, receive from the receiving device a decimation pattern instruction indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data, encode a second set of pictures of video data to generate second encoded video data, apply the decimation pattern to the second encoded video data to generate decimated video data, and transmit the decimated video data to a receiving device.
[0370]
[0379] The device of Clause 1D, wherein one or more processors are configured to generate first error correction data based on first encoded video data and transmit the first error correction data to a receiving device, and to generate second error correction data based on second encoded video data and transmit the second error correction data to a receiving device.
[0371]
[0380] Clause 3D: A device that displays a decimation pattern that skips the transmission of encoded video data of the complete picture, either Clause 1D or 2D.
[0372]
[0381] Clause 4D: A device in any of Clauses 1D through 3D where the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specific area within the picture.
[0373]
[0382] Clause 5D: Any device under Clauses 1D through 4D, where the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a particular view.
[0374]
[0383] A device according to any of the clauses 1D to 5D, wherein a decimation pattern instruction is a first decimation pattern instruction, a non-transmit pattern of encoded video data is a first non-transmit pattern of encoded video data, decimated video data is the first decimated video data, and one or more processors encode a third set of pictures of video data to generate third encoded video data, determine a second decimation pattern indicating a second non-transmit pattern of encoded video data, apply the second decimation pattern to the third encoded video data to generate the second decimated video data, transmit the second decimated video data to a receiving device, and further configured to transmit to the receiving device a second decimation pattern instruction, the second decimation pattern instruction indicating that the second decimation pattern has been applied to the third encoded video data.
[0375]
[0384] Clause 7D. A device according to any of Clauses 1D to 6D, wherein one or more processors are configured to generate first prediction data for a first set of pictures as part of encoding a first set of pictures, generate residual data based on the first prediction data and the first set of pictures, apply transformation to the first prediction data to generate a transformation block, quantize the transformation coefficients of the transformation block, and apply entropy coding to syntax elements representing the quantized transformation coefficients to generate first entropy-coded syntax elements, the first encoded video data comprising the first entropy-coded syntax elements, and one or more processors are further configured to perform analog modulation on the residual data to generate first analog-modulated residual data, the device further includes a communication interface configured to transmit the first analog-modulated residual data and the first encoded video data.
[0376]
[0385] Clause 8D. A device according to any of Clauses 1D through 7D, wherein the device is an augmented reality (XR) headset and includes a display system, one or more processors further configured to receive virtual element data from a receiving device, and the display system configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0377]
[0386] Clause 9D. A device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, wherein one or more processors receive first encoded video data from a transmitting device, perform a decoding process to reconstruct a first set of pictures based on the first encoded video data, determine a decimation pattern indicating a non-transmitted pattern of the encoded video data based on the first set of pictures, transmit a decimation pattern instruction to the transmitting device indicating the determined decimation pattern, receive decimated video data from the transmitting device, wherein the decimated video data comprises second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data, and perform a decoding process to reconstruct a second set of pictures based on the second encoded video data.
[0378]
[0387] The device of Clause 10D, wherein one or more processors are configured to receive first error correction data from a transmitting device, apply an error correction process to correct the first encoded video data based on the first error correction data in order to generate first error-corrected encoded video data, and one or more processors are configured to perform a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data, and one or more processors are further configured to receive second error correction data from a transmitting device, apply an error correction process to generate second error-corrected encoded video data based on the second encoded video data and the second error correction data, and one or more processors are configured to perform a decoding process to reconstruct a second set of pictures based on the second error-corrected encoded video data.
[0379]
[0388] Clause 11D. Any device under Clauses 9D through 10D whose decimation pattern indicates a pattern that skips the transmission of encoded video data of the complete picture.
[0380]
[0389] A device according to any of the clauses 9D to 11D, wherein one or more processors are configured to apply a decimation pattern to first encoded video data in order to generate decimated encoded video data as part of determining a decimation pattern; to apply an error correction process in order to correct the decimated encoded video data based on first error correction data in order to generate error-corrected trial video data; to apply a decoding process in order to reconstruct a first set of pictures based on the error-corrected trial video data; and to determine whether the decimation pattern meets a criterion based on a comparison between the first set of pictures reconstructed based on the error-corrected trial video data and the first set of pictures reconstructed based on the first video data.
[0381]
[0390] A device under any of the clauses 9D through 12D in which the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specified area within a picture.
[0382]
[0391] Clause 14D. Any device under Clauses 9D through 13D, where the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a specified view.
[0383]
[0392] Clause 15D. A device of any of Clauses 9D to 14D, wherein a decimation pattern instruction is a first decimation pattern instruction, a non-transmit pattern for encoded video data is a first non-transmit pattern for encoded video data, and decimated video data is the first decimated video data, and one or more processors receive a second decimation pattern instruction indicating a second non-transmit pattern for encoded video data, and receive a second decimated video data from a transmitting device, the second decimated video data comprising a third encoded video data to which the second decimation pattern is applied, and the third encoded video data is generated based on a third set of pictures of video data, and is further configured to apply a decoding process to reconstruct the third set of pictures based on the third encoded video data.
[0384]
[0393] Clause 16D. A device according to any of Clauses 9D to 15D, further comprising a communication interface configured to receive analog-modulated residual data, wherein the second encoded video data comprises entropy-coded syntax elements representing quantized transformation coefficients, and one or more processors are configured to apply entropy decoding to the syntax elements to obtain quantized transformation coefficients as part of applying a decoding process to reconstruct a second set of pictures, inverse quantization of the quantized transformation coefficients to generate inverse quantized transformation coefficients, inverse transformation to the inverse quantized transformation coefficients to generate prediction data, demodulate the analog-modulated residual data to obtain residual data, and reconstruct a second set of pictures based on the prediction data and residual data.
[0385]
[0394] Clause 17D. A device of any of Clauses 9D through 16D, wherein one or more processors are further configured to process a second set of pictures to generate virtual element data, and the transmitting device is an XR headset configured to display one or more virtual elements in an augmented reality (XR) scene based on the virtual element data.
[0386]
[0395] A method comprising: encoding a first set of pictures of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving a decimation pattern instruction from the receiving device indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; encoding a second set of pictures of video data to generate second encoded video data; applying the decimation pattern to the second encoded video data to generate decimated video data; and transmitting the decimated video data to a receiving device.
[0387]
[0396] The method of Clause 18D, further comprising generating first error correction data based on first encoded video data, transmitting the first error correction data to a receiving device, generating second error correction data based on second encoded video data, and transmitting the second error correction data to a receiving device.
[0388]
[0397] Clause 20D. Any method of Clauses 18D through 19D, where the decimation pattern indicates a pattern that skips the transmission of encoded video data of the complete picture.
[0389]
[0398] Clause 21D. Any method of Clauses 18D through 20D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specific area within the picture.
[0390]
[0399] Clause 22D. Any method of Clauses 18D through 21D, wherein the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a particular view.
[0391]
[0400] Clause 23D. Any method of Clauses 18D to 22D, wherein Decimation Pattern Indicator is a first Decimation Pattern Indicator, Non-Transmit Pattern of Encoded Video Data is a first Non-Transmit Pattern of Encoded Video Data, and Decimated Video Data is the first Decimated Video Data, and the method further comprises encoding a third set of pictures of video data to generate a third encoded video data; determining a second Decimation Pattern indicating a second Non-Transmit Pattern of Encoded Video Data; applying the second Decimation Pattern to the third encoded video data to generate the second decimated video data; transmitting the second decimated video data to a receiving device; and transmitting to the receiving device a second Decimation Pattern Indicator, the second Decimation Pattern Indicator indicating that the second Decimation Pattern has been applied to the third encoded video data.
[0392]
[0401] Clause 24D. Any method of Clauses 18D to 23D, wherein encoding a first set of pictures includes generating first prediction data for the first set of pictures, generating residual data based on the first prediction data and the first set of pictures, transforming the first prediction data to generate a transform block, quantizing the transform coefficients of the transform block, and applying entropy coding to syntax elements representing the quantized transform coefficients to generate a first entropy-coded syntax element, wherein the first encoded video data includes the first entropy-coded syntax element, and the method further includes performing analog modulation on the residual data to generate first analog-modulated residual data, and transmitting the first analog-modulated residual data and the first encoded video data.
[0393]
[0402] Any method of the provisions
[0394]
[0403] Clause 26D. A method comprising: receiving first encoded video data from a transmitting device; applying a decoding process to reconstruct a first set of pictures based on the first encoded video data; determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; transmitting a decimation pattern instruction to the transmitting device indicating the determined decimation pattern; receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.
[0395]
[0404] The method of Clause 27D, further comprising receiving first error correction data from a transmitting device, applying an error correction process to correct the first encoded video data based on the first error correction data in order to generate first error-corrected encoded video data, and performing a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data, the method further comprising receiving second error correction data from a transmitting device, applying an error correction process to generate second error-corrected encoded video data based on the second encoded video data and the second error correction data, and performing a decoding process to reconstruct a second set of pictures based on the second error-corrected encoded video data, the method of Clause 26D.
[0396]
[0405] Clause 28D. Any method of Clauses 26D through 27D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data of the complete picture.
[0397]
[0406] Clause 29D. Any method of Clauses 26D to 28D, wherein determining a decimation pattern includes: applying a decimation pattern to first encoded video data to generate decimated encoded video data; applying an error correction process to correct the decimated encoded video data based on first error correction data to generate error-corrected trial video data; applying a decoding process to reconstruct a first set of pictures based on the error-corrected trial video data; and determining whether the decimation pattern meets a criterion based on a comparison of the first set of pictures reconstructed based on the error-corrected trial video data with the first set of pictures reconstructed based on the first video data.
[0398]
[0407] Clause 30D. Any method of Clauses 26D through 29D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specified area within the picture.
[0399]
[0408] Clause 31D. Any method of Clauses 26D through 30D, wherein the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a specified view.
[0400]
[0409] Clause 32D. Any method of Clauses 26D to 31D, wherein a decimation pattern instruction is a first decimation pattern instruction, a non-transmitted pattern of encoded video data is a first non-transmitted pattern of encoded video data, and decimated video data is the first decimated video data, the method further comprising receiving a second decimation pattern instruction indicating a second non-transmitted pattern of encoded video data; receiving from a transmitting device a second decimated video data, the second decimated video data comprising a third encoded video data to which the second decimation pattern is applied, and the third encoded video data is generated based on a third set of pictures of video data; and applying a decoding process to reconstruct the third set of pictures based on the third encoded video data.
[0401]
[0410] Clause 33D. Any method of Clauses 26D to 32D, further comprising receiving analog-modulated residual data, wherein the second encoded video data includes entropy-coded syntax elements representing quantized transformation coefficients, and applying a decoding process to reconstruct a second set of pictures, comprising: applying entropy decoding to the syntax elements to obtain quantized transformation coefficients; dequantizing the quantized transformation coefficients to generate dequantized transformation coefficients; applying an inverse transform to the dequantized transformation coefficients to generate prediction data; demodulating the analog-modulated residual data to obtain residual data; and reconstructing a second set of pictures based on the prediction data and the residual data.
[0402]
[0411] Any method of the provisions of 34D, further comprising processing a second set of pictures to generate virtual element data, wherein the transmitting device is an XR headset configured to display one or more virtual elements in an augmented reality (XR) scene based on the virtual element data.
[0403]
[0412] A device comprising: means for encoding a first set of pictures of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving a decimation pattern instruction from the receiving device indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; means for encoding a second set of pictures of video data to generate second encoded video data; means for applying the decimation pattern to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to a receiving device.
[0404]
[0413] Clause 36D. A device comprising: means for receiving first encoded video data from a transmitting device; means for performing a decoding process to reconstruct a first set of pictures based on the first encoded video data; means for determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; means for transmitting a decimation pattern instruction to the transmitting device indicating the determined decimation pattern; and means for receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and means for performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.
[0405]
[0414] Clause 1E. A device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, wherein one or more processors are configured to encode a first picture of video data to generate first encoded video data, transmit the first encoded video data to a receiving device, receive from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in the decoding order, encode the second picture based on the encoded selection data for the second picture to generate second encoded video data, and transmit the second encoded video data to a receiving device.
[0406]
[0415] Clause 2E. The device of Clause 1E, wherein the encoding selection data received from the receiving device includes motion parameters for a block of a second picture, and one or more processors are configured to perform motion compensation based on the motion parameters for a block of a second picture in order to generate a prediction block as part of encoding the second picture, and the second encoded video data includes encoded video data based on the prediction block.
[0407]
[0416] Clause 3E. A device according to any of Clauses 1E to 2E, wherein the encoding selection data received from the receiving device includes intra-prediction parameters for a block of a second picture, and one or more processors are configured to perform intra-prediction based on the intra-prediction parameters for a block of a second picture in order to generate a prediction block as part of encoding the second picture, and the second encoded video data includes encoded video data based on the prediction block.
[0408]
[0417] Clause 4E. Any device under Clauses 1E through 3E where the second encoded video data does not contain the encoded selection data.
[0409]
[0418] Clause 5E. A device according to any of Clauses 1E through 4E, in which one or more processors are configured to entropy decode the encoded selection data for the second picture pair.
[0410]
[0419] A device according to any of the clauses 1E to 5E, wherein one or more processors are further configured to generate first error correction data based on first encoded video data and to transmit the first encoded video data and the first error correction data to a receiving device.
[0411]
[0420] Any device of Clause 1E to 6E, wherein one or more processors are further configured to encode a third picture without using encoding selection data for the third picture, based on the determination that encoding selection data for the third picture will not be received from the receiving device before the expiration of the time limit.
[0412]
[0421] A device according to any of the clauses 1E to 7E, wherein one or more processors receive encoding selection data for a third picture of video data, wherein the encoding selection data for the third picture indicates an encoding selection used to encode an estimate of the third picture; and apply a channel coding process to encode the third picture based on the encoding selection data for the third picture in order to generate a third encoded video data; and transmit the error correction data for the third encoded video data to a receiving device without transmitting at least a portion of the third encoded video data.
[0413]
[0422] Clause 9E. A device according to any of Clauses 1E through 8E, wherein the device is an augmented reality (XR) headset and includes a display system, one or more processors further configured to receive virtual element data from a receiving device, and the display system configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0414]
[0423] Clause 10E. A device comprising a memory configured to store video data and one or more processors implemented in a circuit and coupled to the memory, wherein one or more processors are configured to receive first encoded video data from a transmitting device, reconstruct a first picture of the video data based on the first encoded video data, estimate a second picture of the video data based on the first picture, wherein the second picture is a picture that follows the first picture in the decoding order, generate encoding selection data for the estimated second picture, wherein the encoding selection data indicates the encoding selection used to encode the estimated second picture, transmit the encoding selection data for the second picture to the transmitting device, receive second encoded video data from the transmitting device, and reconstruct a second picture based on the second encoded video data.
[0415]
[0424] Clause 11E. A device of Clause 10E in which one or more processors are configured to perform motion compensation based on motion parameters for blocks of a second picture in order to generate predicted blocks as part of encoding a predicted second picture, wherein the encoded selection data includes motion parameters for blocks of a second picture, and the second encoded video data includes encoded video data based on predicted blocks.
[0416]
[0425] A device according to any of the clauses 10E to 11E, wherein one or more processors are configured to perform intra-prediction based on intra-prediction parameters for blocks of a second picture in order to generate predictive blocks as part of encoding a second picture, and the encoded selection data includes intra-prediction parameters for blocks of a second picture, and the second encoded video data includes encoded video data based on the predictive blocks.
[0417]
[0426] A device under any of the clauses 10E through 12E, wherein the second encoded video data does not include encoding selection data, and one or more processors are configured to use the encoding selection data to reconstruct a second picture based on the second encoded video data as part of applying a decoding process.
[0418]
[0427] A device of any of the provisions of 10E to 13E, in which one or more processors are configured to entropically encode the encoding selection data for the second picture before transmitting the encoding selection data for the second picture.
[0419]
[0428] Any device of any of the clauses 10E to 14E, further configured to have one or more processors that estimate a third picture of video data based on one or more of the first or second pictures, perform an encoding process to encode the estimated third picture in order to generate third encoded video data, wherein third encoding selection data indicates the encoding selection used to encode the estimated third picture, transmit the third encoding selection data to a transmitting device, receive error correction data for the third picture from the transmitting device, apply an error correction process to generate error-corrected encoded video data for the third picture based on the error correction data for the third picture and the third encoded video data, and apply a decoding process to reconstruct the third picture based on the error-corrected encoded video data for the third picture.
[0420]
[0429] Clause 16E. A device of Clause 15E in which error-corrected encoded video data for a third picture does not include third encoding selection data, and one or more processors are configured to use third encoding selection data to reconstruct the third picture based on the error-corrected encoded video data for the third picture as part of applying a decoding process.
[0421]
[0430] A device according to any of the clauses 10E to 16E, wherein one or more processors are configured to apply a channel coding process to the coded selection data for the second picture in order to generate error correction data for the coded selection data for the second picture, and to transmit the error correction data for the coded selection data for the second picture to the transmitting device.
[0422]
[0431] Clause 18E. A device according to any of Clauses 10E to 17E, which includes a communication interface configured to modulate coded selection data at a lower modulation order compared to other data transmissions in the data link between the device and the transmitting device.
[0423]
[0432] A device of any of the clauses 10E through 18E, wherein one or more processors are further configured to process a second set of pictures to generate virtual element data, and the transmitting device is an XR headset configured to display one or more virtual elements in an augmented reality (XR) scene based on the virtual element data.
[0424]
[0433] Clause 20E. A method for processing video data, the method comprising: encoding a first picture of video data to generate a first encoded video data; transmitting the first encoded video data to a receiving device; receiving from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in the decoding order; encoding a second picture based on the encoded selection data for the second picture to generate a second encoded video data; and transmitting the second encoded video data to a receiving device.
[0425]
[0434] Clause 21E. The method of Clause 20E, wherein the encoding selection data received from the receiving device includes motion parameters for a block of a second picture, and encoding the second picture includes performing motion compensation based on the motion parameters for a block of a second picture to generate a predicted block, and the second encoded video data includes encoded video data based on the predicted block.
[0426]
[0435] Clause 22E. Encoding selection data received from a receiving device includes intra-prediction parameters for a block of a second picture, and encoding the second picture includes performing intra-prediction based on the intra-prediction parameters for a block of a second picture to generate a prediction block, and the second encoded video data includes encoded video data based on the prediction block, in any way of Clauses 20E to 21E.
[0427]
[0436] Clause 23E. The second encoded video data does not include the encoded selection data, in any manner described in Clauses 20E through 22E.
[0428]
[0437] Clause 24E. Any method of Clauses 20E to 23E further comprises entropy decoding of the encoded selected data for the second picture pair.
[0429]
[0438] Any method of the provisions of 20E to 24E, further comprising generating first error correction data based on first encoded video data and transmitting the first encoded video data and the first error correction data to a receiving device.
[0430]
[0439] Any method of the provisions of 20E to 25E, further comprising encoding the third picture without using the encoding selection data for the third picture, based on the determination that the encoding selection data for the third picture is not received from the receiving device before the expiration of the time limit.
[0431]
[0440] Any method of the provisions of 20E to 26E, further comprising: receiving coding selection data for a third picture of video data, wherein the coding selection data for the third picture indicates the coding selection used to encode an estimate of the third picture; encoding a third picture based on the coding selection data for the third picture to generate a third encoded video data; applying a channel coding process to generate error correction data for the third encoded video data; and transmitting the error correction data for the third encoded video data to a receiving device without transmitting at least a portion of the third encoded video data.
[0432]
[0441] Clause 28E. Any method of Clauses 20E to 27E, wherein the device is an augmented reality (XR) headset and includes a display system, and the method further includes receiving virtual element data from a receiving device and displaying one or more virtual elements in an XR scene on the display system based on the virtual element data.
[0433]
[0442] Clause 29E. A method for processing video data, comprising: receiving first encoded video data from a transmitting device; reconstructing a first picture of the video data based on the first encoded video data; estimating a second picture of the video data based on the first picture, wherein the second picture is a picture that follows the first picture in the decoding order; generating encoding selection data for the estimated second picture, wherein the encoding selection data indicates the encoding selection used to encode the estimated second picture; transmitting the encoding selection data for the second picture to a transmitting device; receiving second encoded video data from a transmitting device; and reconstructing a second picture based on the second encoded video data.
[0434]
[0443] Clause 30E. The method of Clause 29E, wherein encoding an estimated second picture comprises performing motion compensation based on motion parameters for blocks of the second picture to generate predicted blocks, the encoded selection data comprises motion parameters for blocks of the second picture, and the second encoded video data comprises encoded video data based on predicted blocks.
[0435]
[0444] Clause 31E. Encoding the second picture comprises performing an intra-prediction based on intra-prediction parameters for the second picture block to generate a predictive block, wherein the encoded selection data comprises the intra-prediction parameters for the second picture block, and the second encoded video data comprises the encoded video data based on the predictive block, in any way of Clauses 29E to 30E.
[0436]
[0445] Clause 32E. Any method of Clauses 29E to 31E, wherein the second encoded video data does not include encoded selection data, and the decoding process involves using the encoded selection data to reconstruct the second picture based on the second encoded video data.
[0437]
[0446] Clause 33E. Entropy encoding of the encoding selection data for the second picture is performed before transmitting the encoding selection data for the second picture, in any manner of Clauses 29E to 32E.
[0438]
[0447] Any method of the provisions of 29E to 33E, further comprising: estimating a third picture of video data based on one or more of the first or second pictures; performing an encoding process to encode the estimated third picture in order to generate a third encoded video data, wherein the third encoding selection data indicates the encoding selection used to encode the estimated third picture; transmitting the third encoding selection data to a transmitting device; receiving error correction data for the third picture from the transmitting device; applying an error correction process to generate error-corrected encoded video data for the third picture based on the error correction data for the third picture and t...
Claims
1. A device that processes video data, A memory configured to store video data, A communication interface is configured to receive error correction data from a transmitting device, which provides error correction information relating to the picture in the video data. The circuit is implemented and comprises one or more processors coupled to the memory, wherein the one or more processors Predictive data for the picture, which includes predictions of blocks of the picture based at least in part on pictures reconstructed before one or more of the video data, Based on the prediction data for the picture, encoded video data is generated, which includes a transformation block containing transformation coefficients. The bits of the conversion coefficients of the conversion block are scaled based on a reliability value for the bit position. To perform an error correction operation on the scaled bits of the conversion coefficients of the conversion block, error-corrected encoded video data is generated using the error-corrected data. A device configured to reconstruct the picture based on the error-corrected encoded video data.
2. The device according to claim 1, wherein the one or more processors are further configured to generate the reliability value.
3. The device according to claim 2, wherein one or more processors are configured to generate the reliability value based on statistics regarding the occurrence of errors at the bit positions.
4. The device according to claim 2, wherein the one or more processors are configured to generate the reliability value based on the reliability characteristics of individual regions of the video data picture.
5. The device according to claim 2, wherein one or more processors are configured to generate the reliability value based on a noise model.
6. The device according to claim 1, wherein the communication interface is further configured to send the reliability value to the transmitting device.
7. The device according to claim 1, wherein the communication interface is further configured to receive the reliability value from the transmitting device.
8. A device that processes video data, A memory configured to store the aforementioned video data, One or more processors implemented in the circuit and coupled to the memory, Acquire video data, Predictive quality feedback is obtained, which is based on the reliability of the estimated picture generated by the receiving device. Based on the aforementioned predictive quality feedback, one or more of the video coding parameters or channel coding parameters are adapted. Based on one or more pictures of the acquired video data, a video encoding process is performed, which is controlled by the video encoding parameters, in order to generate encoded video data. To generate channel-coded data, one or more processors are configured to perform a channel coding process on the encoded video data, the channel coding process being controlled by the channel coding parameters, A device including a communication interface configured to transmit the channel-encoded data to the receiving device.
9. The aforementioned video coding parameters include quantization parameters, The one or more processors are configured to adapt the quantization parameters as part of adapting the video encoding parameters, The device according to claim 8, wherein one or more processors are configured to use the quantization parameters to quantize the conversion coefficients of the conversion blocks of one or more pictures as part of performing the video encoding process.
10. The channel coding parameters include a low-density parity check (LDPC) graph. The one or more processors are configured to adapt the LPDC graph as part of adapting the channel coding parameters, The device according to claim 8, wherein one or more processors are configured to use the LDPC graph to generate codewords to be included in the channel-encoded data as part of performing the channel coding process.
11. The channel-encoded data includes error-corrected data, The one or more processors are further configured to adapt one or more bit puncturing parameters based on the predictive quality feedback, The device according to claim 8, wherein one or more processors are configured to perform a bit puncturing process on the error correction data, and the bit puncturing process is controlled by one or more bit puncturing parameters.
12. A method for processing video data, The receiving device obtains error correction data from the transmitting device, which provides error correction information relating to the picture in the video data. The receiving device generates prediction data for the picture, which includes predictions of blocks of the picture based at least in part on pictures reconstructed before one or more of the video data; The receiving device generates encoded video data, which includes a conversion block containing conversion coefficients, based on the prediction data for the picture. In the receiving device, the bits of the conversion coefficient of the conversion block are scaled based on a reliability value for the bit position. The receiving device generates error-corrected encoded video data using the error correction data in order to perform an error correction operation on the scaled bits of the conversion coefficients of the conversion block, A method comprising, in the receiving device, reconstructing the picture based on the error-corrected encoded video data.
13. The method according to claim 12, further comprising generating the reliability value in the receiving device.
14. The method according to claim 13, wherein generating the reliability value includes generating the reliability value in the receiving device based on statistics regarding the occurrence of errors at the bit positions.
15. The method according to claim 13, wherein generating the reliability value includes, in the receiving device, generating the reliability value based on the reliability characteristics for individual regions of the picture of the video data.
16. The method according to claim 13, wherein generating the reliability value includes generating the reliability value based on a noise model in the receiving device.
17. The method according to claim 12, further comprising sending the reliability value to the transmitting device in the receiving device.
18. The method according to claim 12, further comprising the receiving device receiving the reliability value from the transmitting device.
19. A method for processing video data, Acquiring video data, Obtaining predictive quality feedback, which is based on the reliability of the estimated picture generated by the receiving device, Adapting one or more of the video coding parameters or channel coding parameters based on the aforementioned predictive quality feedback, To generate encoded video data based on one or more pictures of the acquired video data, a video encoding process is performed, which is controlled by the video encoding parameters. To generate channel-encoded data, a channel coding process is performed on the encoded video data, which is controlled by the channel coding parameters. A method comprising transmitting the channel-encoded data to the receiving device.
20. The aforementioned video coding parameters include quantization parameters, Applying the aforementioned video coding parameters includes applying the aforementioned quantization parameters. The method according to claim 19, wherein performing the video encoding process includes using the quantization parameter to quantize the conversion coefficients of the one or more picture conversion blocks.
21. The channel coding parameters include a low-density parity check (LDPC) graph. Applying the channel coding parameters includes applying the LPDC graph, The method according to claim 19, wherein performing the channel coding process includes using the LDPC graph to generate codewords contained in the channel coded data.
22. The channel-encoded data includes error-corrected data, The method described above is Adapting one or more bit puncturing parameters based on the aforementioned predictive quality feedback, The method according to claim 19, further comprising performing a bit puncturing process on the error correction data, wherein the bit puncturing process is controlled by the one or more bit puncturing parameters.