Companion device-assisted multi-view video coding

Distributed video coding shifts encoding work from resource-constrained transmitting devices to receiving devices, using error correction and decimation patterns to reduce complexity and power consumption, ensuring efficient high-quality video transmission.

JP2026514474APending Publication Date: 2026-05-11QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-03-27
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Modern video encoding processes are resource-intensive and require complex processors, high-speed memory, and significant energy consumption, which is not suitable for devices with limited resources, especially in wireless communication systems like 5G and 6G, particularly when transmitting high-quality video data over short distances.

Method used

A distributed video coding (DVC) process is employed where the transmitting device performs a limited video coding process and sends error correction data, allowing the receiving device to estimate and fully encode pictures using more complex tools, reducing the need to transmit all encoded video data, and applying decimation patterns to further reduce data transmission.

Benefits of technology

This approach reduces the complexity and resource consumption at the transmitting device while maintaining low latency and power efficiency, enabling high-quality video reconstruction at the receiving device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514474000001_ABST
    Figure 2026514474000001_ABST
Patent Text Reader

Abstract

The device is configured to acquire a first set of multiview pictures, the first set of multiview pictures comprising a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; to transmit first encoded video data, the first encoded video data based on the first set of multiview pictures, to a receiving device; to receive a multiview encoding queue from the receiving device; to acquire a second set of multiview pictures of video data, the second set of multiview pictures comprising a third picture and a fourth picture, where the third picture is from a first viewpoint and the fourth picture is from a second viewpoint; to perform a multiview encoding process on the second set of multiview pictures to generate second encoded video data based on the multiview encoding queue; and to transmit second encoded video data to a receiving device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Patent Application No. 18 / 306,136, filed on 24 April 2023, which is incorporated herein by reference in its entirety.

[0002] This disclosure relates to video coding and video decoding. [Background technology]

[0003] Virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies are rapidly gaining popularity and are expected to be widely adopted in non-gaming applications such as healthcare, education, social media, retail, and many others. VR, AR, and MR are sometimes collectively referred to as Extended Reality (XR). Due to this growing popularity, there is increasing demand for XR devices, such as XR goggles, which feature high-quality 3D graphics, higher video resolution, and low latency response. [Overview of the project]

[0004] This disclosure describes techniques for processing video data in a transmitting device and a receiving device. The transmitting device may be an XR device or other type of device. The receiving device may be a user equipment (UE) device such as a smartphone or tablet. The transmitting device may perform a limited video coding process to generate encoded video data. The transmitting device may apply channel coding to the encoded video data to generate error correction data. The transmitting device may send the error correction data and at least a portion of the encoded video data to the receiving device. The receiving device may estimate the video data based on one or more previously reconstructed pictures. The receiving device may then encode the estimated video data. The receiving device may use one or more coding tools to encode the estimated video data that was not used by the transmitting device when performing a limited video coding process on the video data. The receiving device may use the error correction data and estimated video data to play back the portion of the encoded video data that was not sent by the transmitting device. This process may avoid the need to send the portion of the encoded video data.

[0005] In one example, the Disclosure describes a method for decoding video data, the method comprising: in a receiving device, obtaining error correction data from a transmitting device, wherein the error correction data provides error correction information and is generated based on encoded video data of one or more blocks of a picture in the video data; in a receiving device, generating prediction data for a picture, wherein the prediction data for a picture includes predictions of blocks of a picture based at least partially on blocks of a picture previously reconstructed in one or more video data, using one or more coding tools not used to generate encoded video data of one or more blocks; in a receiving device, generating encoded video data based on the prediction data for a picture; in a receiving device, generating error-corrected encoded video data using the error correction data in order to perform error correction processing on the encoded video data; and in a receiving device, performing a reconstruction operation to reconstruct blocks of an image based on the error-corrected encoded video data, wherein the reconstruction operation is controlled by the values ​​of one or more parameters.

[0006] In another example, the Disclosure describes a method for encoding video data, the method comprising: a transmitting device acquiring video data from a video source; the transmitting device generating encoded video data of a first picture and encoded video data of a second picture of the video data based on a set of parameters; the transmitting device performing channel coding on the encoded video data of the first picture and the encoded video data of the second picture in order to generate error correction data for the first picture and error correction data for the second picture; and the transmitting device transmitting the encoded video data of the first picture, the error correction data for the first picture, and the error correction data for the second picture.

[0007] In another example, the Disclosure describes a method for encoding video data, the method comprising: in a transmitting device, obtaining video data from a video source; in a transmitting device, generating transformation blocks based on the video data; in a transmitting device, determining which of the transformation blocks are anchor transformation blocks; in a transmitting device, calculating a correlation matrix for the set of transformation blocks; in a transmitting device, generating a bit-reduced non-anchor transformation matrix; and in a transmitting device, transmitting the anchor transformation blocks, the non-anchor transformation blocks, and the correlation matrix to a receiving device.

[0008] In another example, the present disclosure describes a device comprising a memory configured to store video data, a communication interface, and one or more processors implemented in a circuit and coupled to the memory, wherein one or more processors are configured to perform the method according to any one of claims 1 to 22.

[0009] In another example, the Disclosure describes a device for processing video data, comprising: a memory configured to store video data; a communication interface configured to receive error correction data from a transmitting device, wherein the error correction data provides error correction information relating to pictures of video data; and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors are configured to generate prediction data for pictures, wherein the prediction data for pictures includes predictions of blocks of pictures based at least in part on one or more previously reconstructed pictures of video data; to generate encoded video data encoded based on the prediction data for pictures, wherein the encoded video data includes a transformation block containing transformation coefficients; to scale the bits of the transformation coefficients of the transformation block based on confidence values ​​for bit positions; to generate error-corrected encoded video data using the error correction data to perform error correction operations on the scaled bits of the transformation coefficients of the transformation block; and to reconstruct a picture based on the error-corrected encoded video data.

[0010] In another example, the Disclosure describes a device for processing video data, comprising: a memory configured to store video data; one or more processes implemented in the circuit and coupled to the memory; one or more processors configured to acquire video data, acquire predictive quality feedback, wherein the predictive quality feedback is based on the reliability of estimated pictures generated by a receiving device; apply one or more video coding parameters or channel coding parameters based on the predictive quality feedback; execute a video coding process, wherein the video coding process is controlled by video coding parameters, to generate encoded video data based on one or more pictures of the acquired video data; and execute a channel coding process, wherein the channel coding process is controlled by channel coding parameters, to generate channel coded data; and a communication interface configured to transmit the channel coded data to a receiving device.

[0011] In another example, the Disclosure describes a method for processing video data, the method comprising: a receiving device obtaining error correction data from a transmitting device, wherein the error correction data provides error correction information relating to a picture of the video data; a receiving device generating prediction data for a picture, wherein the prediction data for a picture includes predictions of blocks of the picture, at least in part on a previously reconstructed picture of one or more video data; a receiving device generating encoded video data encoded based on the prediction data for a picture, wherein the encoded video data includes a transformation block containing transformation coefficients; a receiving device scaling the bits of the transformation coefficients of the transformation block based on confidence values ​​for bit positions; a receiving device generating error-corrected encoded video data using the error correction data to perform error correction operations on the scaled bits of the transformation coefficients of the transformation block; and a receiving device reconstructing a picture based on the error-corrected encoded video data.

[0012] In another example, the Disclosure describes a method for processing video data, which includes acquiring video data; acquiring predictive quality feedback, wherein the predictive quality feedback is based on the reliability of estimated pictures produced by a receiving device; adapting one or more video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process, controlled by video coding parameters, to generate encoded video data based on one or more pictures of the acquired video data; performing a channel coding process, controlled by channel coding parameters, on the encoded video data to generate channel coded data; and transmitting the channel coded data to a receiving device.

[0013] In another example, the Disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in the circuit and coupled to the memory, the one or more processors taking a first set of multiview pictures of video data, wherein the first set of multiview pictures comprises a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint, and taking a first encoded video data, wherein the first encoded video data is based on the first set of multiview pictures, and taking a multiview picture from the receiving device The system is configured to receive a new encoding queue, obtain a second set of multiview pictures of video data, the second set of multiview pictures includes a third picture and a fourth picture, the third picture is from a first viewpoint and the fourth picture is from a second viewpoint, and to perform a multiview encoding process on the second set of multiview pictures to generate second encoded video data based on the multiview encoding queue received from the receiving device, the multiview encoding process reduces interview redundancy between the third picture and the fourth picture, and transmit the second encoded video data to the receiving device.

[0014] In another example, the disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in the circuit and coupled to the memory, the one or more processors receiving first encoded video data from a transmitting device, the first encoded video data being based on a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint, and the first encoded video data being received The system is configured to obtain, determine a multiview coding queue based on the first encoded video data, send the multiview coding queue to the transmitting device, and obtain second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multiview pictures including a third picture and a fourth picture, and the second encoded video data is coded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.

[0015] In another example, the present disclosure describes a method for processing video data. The method includes obtaining a first set of multi-view pictures of video data, where the first set of multi-view pictures includes a first picture and a second picture, the first picture is from a first viewpoint, and the second picture is from a second viewpoint; transmitting first encoded video data, which is based on the first set of multi-view pictures, to a receiving device; receiving a multi-view encoding queue from the receiving device; obtaining a second set of multi-view pictures of video data, where the second set of multi-view pictures includes a third picture and a fourth picture, the third picture is from the first viewpoint, and the fourth picture is from the second viewpoint; performing a multi-view encoding process on the second set of multi-view pictures to generate second encoded video data based on the multi-view encoding queue received from the receiving device, where the multi-view encoding process reduces inter-view redundancy between the third picture and the fourth picture; and transmitting the second encoded video data to the receiving device.

[0016] In another example, the Disclosure describes a method for processing video data, the method comprising: obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of the video data, the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; determining a multiview coding queue based on the first encoded video data; transmitting the multiview coding queue to the transmitting device; and obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multiview pictures including a third picture and a fourth picture, the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.

[0017] In another example, the present disclosure describes a device, the device comprising means for obtaining a first set of multi-view pictures of video data, the first set of multi-view pictures including a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint; means for transmitting first encoded video data, the first encoded video data being based on the first set of multi-view pictures, to a receiving device; means for receiving a multi-view encoding queue from the receiving device; means for obtaining a second set of multi-view pictures of video data, the second set of multi-view pictures including a third picture and a fourth picture, the third picture being from the first viewpoint and the fourth picture being from the second viewpoint; means for performing a multi-view encoding process on the second set of multi-view pictures to generate second encoded video data based on the multi-view encoding queue received from the receiving device, the multi-view encoding process reducing inter-view redundancy between the third picture and the fourth picture; and means for transmitting the second encoded video data to the receiving device.

[0018] In another example, the Disclosure describes a device comprising means for acquiring first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of the video data, the first set of multiview pictures includes a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint; means for determining a multiview coding queue based on the first encoded video data; means for transmitting the multiview coding queue to the transmitting device; and means for acquiring second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multiview pictures including a third picture and a fourth picture, the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.

[0019] In another example, the Disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in the circuit and coupled to the memory, wherein the one or more processors encode a first set of pictures of video data to generate first encoded video data, transmit the first encoded video data to a receiving device, receive from the receiving device a decimation pattern indication indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data, encode a second set of pictures of video data to generate second encoded video data, apply the decimation pattern to the second encoded video data to generate decimated video data, and transmit the decimated video data to a receiving device.

[0020] In another example, the Disclosure describes a device comprising a memory configured to store video data and one or more processors circuit-implemented and coupled to the memory, the one or more processors receiving first encoded video data from a transmitting device, performing a decoding process to reconstruct a first set of pictures based on the first encoded video data, determining a decimation pattern indicating a non-transmitted pattern of the encoded video data based on the first set of pictures, and transmitting a decimation pattern indication indicating the determined decimation pattern to the transmitting device, receiving decimated video data from the transmitting device, wherein the decimated video data comprises second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data, and is configured to perform a decoding process to reconstruct a second set of pictures based on the second encoded video data.

[0021] In another example, the Disclosure describes a method which includes encoding a first set of pictures of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving a decimation pattern indication from the receiving device which indicates a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; encoding a second set of pictures of video data to generate second encoded video data; applying the decimation pattern to the second encoded video data to generate decimated video data; and transmitting the decimated video data to a receiving device.

[0022] In another example, the Disclosure describes a method which includes receiving first encoded video data from a transmitting device; applying a decoding process to reconstruct a first set of pictures based on the first encoded video data; determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; transmitting a decimation pattern indication indicating the determined decimation pattern to the transmitting device; receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.

[0023] In another example, the Disclosure describes a device comprising: means for encoding a first set of pictures of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving a decimation pattern indication from the receiving device, which indicates a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; means for encoding a second set of pictures of video data to generate second encoded video data; means for applying the decimation pattern to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to a receiving device.

[0024] In another example, the Disclosure describes a device comprising means for receiving first encoded video data from a transmitting device; means for applying a decoding process to reconstruct a first set of pictures based on the first encoded video data; means for determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; means for transmitting a decimation pattern indication indicating the determined decimation pattern to the transmitting device; means for receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and means for performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.

[0025] In another example, the Disclosure describes a device comprising a memory configured to store video data and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors encode a first picture of video data to generate first encoded video data, transmit the first encoded video data to a receiving device, receive from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, the second picture follows the first picture in decoding order, encodes the second picture based on the encoded selection data for the second picture to generate second encoded video data, and transmits the second encoded video data to the receiving device.

[0026] In another example, the Disclosure describes a device comprising a memory configured to store video data and one or more processors circuit-implemented and coupled to the memory, the one or more processors being configured to receive first encoded video data from a transmitting device, reconstruct a first picture of the video data based on the first encoded video data, estimate a second picture of the video data based on the first picture, wherein the second picture occurs after the first picture in the decoding order, generate encoding selection data for the second picture, wherein the encoding selection data for the second picture indicates the encoding selection used to encode the second picture, transmit the encoding selection data for the second picture to the transmitting device, receive second encoded video data from the transmitting device, and reconstruct a second picture based on the second encoded video data.

[0027] In another example, the Disclosure describes a method for processing video data, the method comprising: encoding a first picture of video data to generate a first encoded video data; transmitting the first encoded video data to a receiving device; receiving from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in decoding order; encoding a second picture based on the encoded selection data for the second picture to generate a second encoded video data; and transmitting the second encoded video data to a receiving device.

[0028] In another example, the Disclosure describes a method for processing video data, the method comprising: receiving a first encoded video data from a transmitting device; reconstructing a first picture of the video data based on the first encoded video data; estimating a second picture of the video data based on the first picture, wherein the second picture occurs after the first picture in the decoding order; generating encoding selection data for the second picture, wherein the encoding selection data for the second picture indicates the encoding selection used to encode the second picture; transmitting the encoding selection data for the second picture to a transmitting device; receiving a second encoded video data from the transmitting device; and reconstructing a second picture based on the second encoded video data.

[0029] In another example, the present disclosure describes a device comprising: means for encoding a first picture of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in decoding order; means for encoding a second picture based on the encoding selection data for the second picture to generate second encoded video data; and means for transmitting the second encoded video data to a receiving device.

[0030] In another example, the Disclosure describes a device comprising: means for receiving a first encoded video data from a transmitting device; means for reconstructing a first picture of the video data based on the first encoded video data; means for estimating a second picture of the video data based on the first picture, wherein the second picture is a picture that occurs after the first picture in the decoding order; means for generating encoding selection data for the second picture, wherein the encoding selection data for the second picture indicates an encoding selection used to encode the second picture; means for transmitting the encoding selection data for the second picture to a transmitting device; means for receiving a second encoded video data from a transmitting device; and means for reconstructing a second picture based on the second encoded video data.

[0031] Details of one or more examples are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]

[0032] [Figure 1] Block diagram illustrating an exemplary system using the techniques of this disclosure. [Figure 2] This is a block diagram showing exemplary components of a transmitting device and a receiving device according to the techniques of the present disclosure. [Figure 3A] This is a conceptual diagram illustrating an exemplary channel coding process using the techniques of this disclosure. [Figure 3B] Block diagram showing an exemplary channel decoding process using the techniques of the present disclosure. [Figure 4] This flowchart shows an exemplary operation of a transmitting device using the technique of this disclosure. [Figure 5] This flowchart shows an exemplary operation of a receiving device using the technique of this disclosure. [Figure 6]This is a conceptual diagram illustrating an exemplary decimation pattern using the technique of this disclosure. [Figure 7] This flowchart shows exemplary operation of a transmission device for hybrid decimation of a conversion block using the technique of the present disclosure. [Figure 8] This flowchart shows exemplary operation of a receiving device for hybrid decimation of a conversion block using the technique of the present disclosure. [Figure 9] This is a conceptual diagram showing exemplary decimation patterns adaptively selected by a receiving device using one or more of the techniques of the present disclosure. [Figure 10] This is a block diagram showing exemplary components of a transmitting device and a receiving device according to the techniques of the present disclosure. [Figure 11] This disclosure provides a chart illustrating exemplary error probabilities and corresponding absolute log-likelihood ratios (LLRs) using one or more of the techniques described herein. [Figure 12] This flowchart shows exemplary operation of a transmitting device using scaled bits according to the technique of the present disclosure. [Figure 13] This flowchart shows exemplary operation of a receiving device using scaled bits according to the technique of the present disclosure. [Figure 14] This flowchart illustrates an exemplary exchange of data between a transmitting device and a receiving device involved in multiview processing using one or more techniques of the present disclosure. [Figure 15] This flowchart shows an exemplary operation of a transmission device for multiview processing using the technique of the present disclosure. [Figure 16] This flowchart shows an exemplary operation of a receiving device for multiview processing using the technique of the present disclosure. [Figure 17] This block diagram shows exemplary components of a transmitting and receiving device that perform decimation on encoded video data using the technique of the present disclosure. [Figure 18]This is a conceptual diagram illustrating an exemplary exchange of information, including decimation pattern indication, using the techniques of the present disclosure. [Figure 19] This flowchart shows an exemplary operation of a transmitting device in which the transmitting device receives a decimation pattern indication using the technique of the present disclosure. [Figure 20] This flowchart shows exemplary operation of a receiving device in which the receiving device transmits a decimation pattern indication using the technique of the present disclosure. [Figure 21] This block diagram shows exemplary components of a transmitting device and a receiving device that transmits encoded selection data to the transmitting device, according to the technique of the present disclosure. [Figure 22] This is a communication diagram illustrating an exemplary exchange of data between a transmitting device and a receiving device, including the transmission and reception of coded selection data using the techniques of the present disclosure. [Figure 23] This flowchart shows an exemplary operation of a transmitting device in which it receives encoded selection data using the technique of the present disclosure. [Figure 24] This flowchart shows exemplary operation of a receiving device in which the receiving device transmits encoded selection data using the technique of the present disclosure. [Figure 25] This is a conceptual diagram illustrating an exemplary hierarchy of encoded video data using the techniques of this disclosure. [Figure 26] This block diagram shows exemplary alternative components for a transmitting device using one or more techniques of the present disclosure. [Figure 27] A block diagram showing exemplary alternative components of a receiving device using one or more techniques of the present disclosure. [Modes for carrying out the invention]

[0033] Modern video encoding processes can significantly reduce the amount of data required to represent video data, but such processes are generally resource-intensive and can involve a lot of memory activity. Therefore, modern video encoding processes can require complex processors, high-speed memory, and consume considerable energy. However, in some modern and future planned wireless communication systems, such as 5G and 6G wireless communication systems, wireless transmission bandwidth can be less constrained, especially when communicating over short distances, such as between devices on a person's body.

[0034] This disclosure describes a technique that can reduce the complexity of video coding at a transmitting device by using error correction performed as part of channel decoding using error correction data. The transmitting device may perform a limited video coding process that generates coded video data. The limited video coding process typically uses coding tools such as intra-prediction, which are relatively resource-intensive. Because the video coding process uses less complex coding tools, the resulting coded video data may be larger than the coded video data coded using more complex and resource-intensive coding tools. Error correction data is based on coded video data. The transmitting device may send error correction data to the receiving device. The transmitting device may not need to send all of the coded video data for one or more pictures to the receiving device.

[0035] The receiving device may estimate the picture in the video data based on one or more previously reconstructed pictures. In some cases, to estimate the picture, the receiving device may extrapolate the block content from previously reconstructed pictures. The receiving device may then perform a full video coding process on the estimated picture to generate estimated encoded video data for the picture. When performing a full video coding process, the receiving device may use more complex coding tools, such as interpretation, than the limited video coding process performed by the transmitting device. The receiving device may perform a channel decoding process to generate error-corrected encoded video data based on the estimated encoded video data for the picture and error-corrected data for the picture. In some situations, the channel decoding process may generate error-corrected encoded video data based on the error-corrected data for the picture and a combination of the estimated encoded video data for the picture and the encoded video data for the picture sent by the transmitting device. The receiving device may reconstruct the picture based on the error-corrected encoded video data. In this way, the receiving device may be able to reconstruct each picture of the video data even if the transmitting device did not transmit all of the encoded video data of the picture.

[0036] As further described in this disclosure, various techniques can be applied, such as the application of decimation patterns, to specify which transform blocks of lightly encoded video data are not signaled or have reduced bit depth. Furthermore, in some examples of this disclosure, confidence values ​​may be determined for bit positions, and the bits of the transform coefficients of the transform blocks may be scaled using these confidence values, and the scaled values ​​may be used in channel coding and channel decoding.

[0037] As further described in this disclosure, a receiving device may determine a decimation pattern based on a first set of pictures. The decimation pattern is a non-transmitted pattern of encoded video data. The receiving device may send a decimation pattern indication to a transmitting device showing the determined decimation pattern. The transmitting device receives the decimation pattern indication from the receiving device and may apply the indicated decimation pattern to the encoded video data to generate decimated video data. The transmitting device may transmit the decimated video data to the receiving device. In this way, the technique of this disclosure can further reduce resource consumption in the transmitting device while still avoiding the transmission of excessive amounts of data. This can further increase coding efficiency.

[0038] Figure 1 is a block diagram illustrating an exemplary system 100 using the techniques of the present disclosure. In the example of Figure 1, system 100 comprises a transmitting device 102, a receiving device 104, and a base station 106. The transmitting device 102 may be a device configured to run and include an Extended Reality (XR) device (e.g., an XR headset), a mobile device, a wearable device, a sensor device, an Internet of Things (IoT) device, an intermediate networking device, or another type of device. In some examples, the transmitting device 102 may be included in a robot or a vehicle. The receiving device 104 may be a computing device such as a mobile device (e.g., a cell phone or tablet computer), a personal computer, a vehicle-based computing device, a wireless base station, a wearable computing device, an intermediate networking device, a dedicated device, an Internet of Things (IoT) device, or another type of device. In some examples, the receiving device 104 may be a device that the user of the transmitting device 102 may have in addition to the transmitting device 102.

[0039] The transmitting device 102 and the receiving device 104 can communicate with the base station 106. In some examples, the transmitting device 102 and the receiving device 104 can communicate with the base station 106 using a fifth-generation (5G) wireless communication protocol, a sixth-generation (6G) wireless communication protocol, a WiFi protocol, a Bluetooth protocol, or another type of wireless communication protocol. The base station 106 can transmit data from the network 115 to the transmitting device 102 and the receiving device 104 via wireless downlink channels 108A and 108B (collectively, "wireless downlink channels 108"). The base station 106 can receive data from the transmitting device 102 and the receiving device 104 for transmission to other devices connected to the network 115 via wireless uplink channels 110A and 110B (collectively, "wireless uplink channels 110"). The transmitting device 102 and the receiving device 104 can communicate directly with each other via a wireless sidelink channel 112. In other examples, the transmitting device 102 and the receiving device 104 can communicate via other types of channels. In other examples, the transmitting device 102 and the receiving device 104 may communicate via other types of channels.

[0040] In the example in Figure 1, the transmitting device 102 comprises one or more processors 114, memory 116, a communication interface 118, a video source 120, and a display system 122. The receiving device 104 comprises one or more processors 130, memory 132, and a communication interface 134. Processors 114 and 130 may comprise circuits configured to perform various information processing tasks, including the execution of computer-readable instructions. Processors 114 and 130 may include microprocessors, digital signal processors, and other types of circuits. Memories 116 and 132 may be configured to store data such as computer-readable instructions, video data, and other types of data. The communication interfaces 118 and 134 may be configured to transmit and receive data, for example, via a wireless downlink channel 108, a wireless uplink channel 110, and a wireless sidelink channel 112.

[0041] Generally, video source 120 represents a source of video data (e.g., raw, unencoded video data). Video source 120 may include one or more video capture devices such as video cameras, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, video source 120 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video.

[0042] In an example where the transmitting device 102 is an XR device that presents MR and AR images to the user, video data from the video source 120 may need to be analyzed so that the display system 122 of the transmitting device 102 can display virtual elements in the correct locations. Processing video data in this way can require considerable computing resources. In other words, a powerful processor and a significant amount of energy may be used when processing video data. Since the transmitting device 102 may be designed to be worn on the user's head, it may be important to minimize the weight and power consumption of the transmitting device 102 while supporting high-quality, low-latency video.

[0043] Furthermore, in some examples, the transmitting device 102 may be an XR headset and may be configured to process a picture of video data to generate virtual element data. The receiving device 104 may be configured to transmit virtual element data (and the transmitting device 102 may be configured to receive virtual element data). The transmitting device 102 may include a display system 122 configured to display one or more virtual elements in an XR scene based on the virtual element data.

[0044] Therefore, it may be desirable to offload the processing of video data to a device other than the transmitting device 102, such as the receiving device 104. The receiving device 104 may have more resources than the transmitting device 102, either permanently or temporarily. For example, the receiving device 104 may have a larger battery and a relatively more powerful processor. However, for the receiving device 104 to process the video data, the transmitting device 102 may need to transmit the video data to the receiving device 104 via the wireless sidelink channel 112. Since a very large number of bits may be required to represent unencoded high-quality video data, it will take a considerable amount of time and energy for the transmitting device 102 to transmit unencoded high-quality video data to the receiving device 104. The time required for transmission may undermine the goal of providing low-latency video to the user. The energy required for transmission may undermine the goal of minimizing power consumption. Encoding video data using video coding specifications such as H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), or H.266 / Vulnerable Video Coding (VVC) can significantly reduce the amount of data required to represent the video data. However, the encoding process itself may introduce its own latency and power consumption requirements.

[0045] This disclosure describes techniques that can address these problems. According to the techniques of this disclosure, the transmitting device 102 and the receiving device 104 may use a distributed video coding (DVC) process. The DVC process reduces the amount of coding work performed by the transmitting device 102 and shifts some of the coding work to the receiving device 104. The receiving device 104 may have more resources (e.g., computing power, access to power, etc.) than the transmitting device 102 and can therefore be better equipped to perform the coding work. In some examples, the DVC process may be used to load balance computing tasks across devices. For example, the system may determine that it may be more efficient overall for the receiving device 104 to perform certain video-related computing tasks than for the transmitting device 102.

[0046] In addition to the video encoding process, the transmitting device 102 may perform a channel encoding process to prepare the encoded video data for transmission to the receiving device 104. The channel encoding process may generate error correction data for sequences of data within the encoded video data. Typically, the receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, according to the techniques of this disclosure, the transmitting device 102 may send error correction data for some encoded video data, but the error correction data may not send the corresponding encoded video data. The receiving device 104 may estimate one or more subsequent pictures. The receiving device 104 may perform a video encoding process on the subsequent pictures to generate estimated encoded video data. The receiving device may use the estimated encoded video data and the received error correction data to generate error-corrected encoded video data. The receiving device 104 may then decode the error-corrected encoded video data to reconstruct the video data that the transmitting device 102 did not send.

[0047] Therefore, in some examples, the receiving device 104 may receive first encoded video data and first error correction data from the transmitting device 102. The first encoded video data may represent one or more blocks of first pictures of video data. The first error correction data may provide error correction information regarding the blocks of first pictures. The receiving device 104 may use the first error correction data to generate first error-corrected encoded video data in order to perform an error correction operation on the first encoded video data. In addition, the receiving device 104 may perform a first reconstruction operation to reconstruct the blocks of first pictures based on the first encoded video data. The first reconstruction operation may be controlled by the values ​​of one or more parameters.

[0048] Furthermore, the receiving device 104 may obtain second error correction data from the transmitting device 102. The second error correction data may provide error correction information for one or more blocks of the second picture of the video data. The receiving device 104 may generate prediction data for the second picture. The prediction data for the second picture may include predictions of blocks of the second picture of the video data, at least partially based on blocks of one or more previously reconstructed pictures, such as the first picture. The receiving device 104 may use one or more coding tools to generate prediction data that was not used to generate encoded video data for the second picture. Based on the predictions of blocks of the second picture, the receiving device 104 may generate second encoded video data. The receiving device 104 may use the second error correction data to generate second error-corrected encoded video data in order to perform error correction operations on the second encoded video data. The receiving device 104 may perform a second reconstruction operation in which it reconstructs a second block of picture based on the second error-corrected encoded video data. The second reconstruction operation is controlled by the value of a parameter.

[0049] Furthermore, according to one or more techniques of this disclosure, a receiving device 104 may receive a decimation pattern indication from the receiving device 104. The receiving device 104 may determine a decimation pattern indication determined based on a previously reconstructed picture. The decimation pattern indication may indicate a pattern of non-transmission of encoded video data. For example, the decimation pattern may indicate a pattern of skipping the transmission of encoded video data for a full picture. In some examples, the decimation pattern indicates a pattern of skipping the transmission of encoded video data for a specified area within the picture. In some examples where the video data is multi-view video data, the decimation pattern may indicate a pattern of skipping the transmission of encoded video data for a picture from a specified view.

[0050] The transmitting device 102 may perform a video encoding process on the picture of the video data. The video encoding process may compress the picture less than "heavy" or more complex compression operations, such as those described in the H.264, H.265, and H.266 video coding standards. In addition to the video encoding process, the transmitting device 102 may perform a channel encoding process to prepare the encoded video data for transmission to the receiving device 104. The channel encoding process may generate error correction data for the sequence of data within the encoded video data. Typically, the receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, the receiving device 104 may also use the error correction data to recover information that was not intentionally transmitted to the receiving device. Therefore, the transmitting device 102 may apply a decimation pattern to the second encoded video data to generate decimated video data. The transmitting device 102 may transmit error correction data (generated based on undecimated encoded video data) and decimated video data to the receiving device 104.

[0051] The receiving device 104 may obtain first encoded video data and first error correction data from the transmitting device 102. The first encoded data may represent one or more blocks of first pictures of video data. The first error correction data may provide error correction information regarding the blocks of first pictures. The receiving device 104 may use the first error correction data to generate first error-corrected encoded video data in order to perform an error correction operation on the first encoded video data. In addition, the receiving device 104 may perform a first reconstruction operation to reconstruct the blocks of first pictures based on the first encoded video data. The first reconstruction operation may be controlled by the values ​​of one or more parameters.

[0052] Furthermore, the receiving device 104 may obtain first error correction data and first encoded video data from the transmitting device 102. The receiving device 104 may apply an error correction process to correct the first encoded video data based on the first error correction data in order to generate first error-corrected encoded video data. The receiving device 104 may also apply a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data. Based on the first set of pictures, the receiving device 104 may determine a decimation pattern indicating a pattern of non-transmission of the encoded video data. The receiving device 104 may transmit a decimation pattern indication indicating the determined decimation pattern to the transmitting device 102. The receiving device 104 may receive second error correction data and decimated video data from the transmitting device 102. The decimated video data may include second encoded video data to which the decimation pattern has been applied. The second encoded video data is generated based on a second set of pictures in the video data. The receiving device 104 may apply an error correction process to correct the second encoded video data based on the second error correction data in order to generate the second error-corrected encoded video data. The receiving device 104 may apply a decoding process to reconstruct the second set of pictures based on the second error-corrected encoded video data.

[0053] Figure 2 is a block diagram showing exemplary components of a transmitting device and a receiving device according to the technique of the present disclosure. System 200 comprises a transmitting device 102 and a receiving device 104. Transmitting device 102 is configured to transmit encoded video data to receiving device 104. In the example of Figure 2, transmitting device 102 comprises a video encoder 210, a channel encoder 212, and a puncturing unit 214. Receiver device 104 comprises a depuncturing unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, and a video encoder 228. In other examples, transmitting device 102 and receiver device 104 may include more, fewer, or different units. The processor 114 of transmitting device 102 (Figure 1) may implement the video encoder 210, the channel encoder 212, and the puncturing unit 214. The processor 130 of the receiving device 104 may implement a depuncturing unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, and a video encoder 228. A communication interface 118 (Figure 1) can transmit and receive data on behalf of the transmitting device 102. A communication interface 134 (Figure 1) can transmit and receive data on behalf of the receiving device 104.

[0054] The video encoder 210 of the transmitting device 102 may receive video data from a video source (e.g., video source 120 (Figure 1)). The video data may include, for example, raw, unencoded video pictures from video source 120. In some examples, the memory of the transmitting device 102 (e.g., memory 116 (Figure 1)) may store the video data. The video encoder 210 may perform a video encoding process on the video data to produce encoded video data. The video encoding process may be "limited" in the sense that it may be relatively fast and consume fewer resources than more robust video compression processes such as H.264 / AVC, H.265 / HEVC, or H.266 / VVC. The video encoding process may not reduce the number of bits representing the video data to the same extent as a more robust or full video encoding process.

[0055] The video encoder 210 may perform a limited video encoding process in one of several ways. For example, in some cases, the video encoder 210 may perform a prediction process, such as an intra-prediction process, for each picture of the video data to generate prediction data. In addition, the video encoder 210 may generate residual data based on the prediction data. For example, the video encoder 210 may subtract a sample of the prediction data from the corresponding sample of the original picture to determine the sample of the residual data. The sample may be a value representing a color value (such as a Y, Cb, or Cr value in the YCbCr color region, or a red, green, or blue value in the RGB color region).

[0056] The video encoder 210 may apply a transformation, such as a discrete cosine transform (DCT), to the residual data to generate a transformation block containing transformation coefficients. Furthermore, the video encoder 210 may quantize the transformation coefficients. The video encoder 210 may apply entropy coding, such as context-adaptive binary arithmetic coding (CABAC) coding or exponential Golomb-Rice coding, to the syntax elements representing the quantized transformation coefficients. The encoded video data may include entropy-coded syntax elements. In some examples, the video encoder 210 directly applies transformation and / or quantization to the video data without first using intra-prediction. In some examples where the video encoder 210 does not apply entropy coding, the encoded video data includes syntax elements representing quantized transformation coefficients, unquantized transformation coefficients, or residual data.

[0057] In cases where the video encoder 210 does not use interpicture prediction, fewer memory read requests may be required compared to a more robust video compression process that may need to read data about previously coded pictures from memory. Such memory read requests may be relatively time-intensive and energy-intensive.

[0058] In some examples where the video data is multiview video data, the video encoder 210 may perform multiview video coding to generate prediction data. For example, the video encoder 210 may use interview prediction to generate prediction data for blocks of non-anchor pictures (e.g., macroblocks, coding units, etc.). In some cases, interview prediction may involve determining a disparity vector for the block, which indicates the lateral displacement between the block and the corresponding block in the picture of one or more reference views.

[0059] The channel encoder 212 of the transmitting device 102 may apply a channel coding process to the encoded video data. The channel coding process prepares the encoded video data for transmission over a wireless communication channel, such as channel 230. Channel 230 may be wireless sidelink channel 112 (Figure 1) or another communication channel. The channel coded video data may include error correction data. The channel encoder 212 may generate error correction data in various ways. For example, the channel encoder 212 may generate error correction data as a convolutional code or a turbo code. The error correction data may help the receiving device 104 determine whether the received coded video data has been modified during transmission over channel 230 and may help the receiving device 104 correct such modification. A more detailed explanation of channel coding and channel decoding is provided below with respect to Figure 3.

[0060] Furthermore, in the example shown in Figure 2, the puncturing unit 214 of the transmitting device 202 may apply a bit puncturing process to the error correction data to generate bit-punctured error correction data. The bit puncturing process may reduce the number of bits in the error correction data. For example, the puncturing unit 214 may perform an operation to remove bits from the error correction data according to a puncturing pattern.

[0061] The transmitting device 102 may transmit data such as encoded video data and error correction data (e.g., bit-punctured error correction data) to the receiving device 104 via channel 230. Channel 230 may introduce noise into the transmitted data. In some examples, channel 230 is a multipath channel, and the data transmitted in channel 230 may vary over time. The receiving device 104 may receive the noise-corrected data. The receiving device 104 may store the noise-corrected data at least temporarily in a memory such as memory 132 (Figure 1).

[0062] The depunching unit 220 may perform a depunching operation on the received bit-punctured error-corrected data to reconstruct the error-corrected data. The depunching operation may replace punctured symbols with neutral values ​​as indicated by the puncture pattern. The depunching operation may generate erase bits indicating the presence of neutral symbols in the error-corrected data.

[0063] The channel decoder 222 can apply a channel decoding process to generate error-corrected encoded video data based on error-corrected data and encoded video data such as encoded video data received from the transmitting device and / or encoded video data generated by the receiving device 104. For example, the channel decoder 222 can modify the values ​​of the bit-encoded video data according to one of various error correction schemes, such as low-density parity-check (LDPC) coding or forward error correction (FEC).

[0064] The video decoder 224 may perform a video decoding process to reconstruct a picture based on the error-corrected encoded video data. For example, the video decoder 224 may apply an entropy decoding process to the bits of the error-corrected encoded video data to obtain quantized transformation coefficients. The video decoder 224 may apply an inverse quantization operation to the quantized transformation coefficients, apply an inverse transform to the inversely quantized transformation coefficients to generate residual data, generate prediction data to generate prediction data, and use the prediction data and residual data to reconstruct a picture of the video data. The video decoder 224 may generate prediction data in the same manner as the video encoder 210.

[0065] The picture estimation unit 226 can generate an estimate of the next picture in the video data. For example, the picture estimation unit 226 can extrapolate the next picture from two or more previously reconstructed pictures. For example, in this example, the picture estimation unit 226 can divide the first previously reconstructed picture into blocks. For each block of the first previously reconstructed picture, the picture estimation unit 226 can determine one or more corresponding blocks for the block in one or more additional previously reconstructed pictures. The corresponding blocks for a block may be the best available match for the block. Based on one or more corresponding blocks for a block, the picture estimation unit 226 can generate a prediction for the block. The picture estimation unit 226 can use unidirectional or bidirectional prediction to generate the prediction. Thus, by generating a prediction for each block of the next picture, the picture estimation unit 226 can generate an estimate of the next picture. In some examples, the picture estimation unit 226 generates the next picture by applying global motion to the previously reconstructed picture.

[0066] In some examples, the picture estimation unit 226 may re-encode the current picture decoded by the video decoder 224. The next picture in the video data may be the picture that follows the picture in the same decoding order that the video decoder 224 just decoded. In this example, the picture estimation unit 226 may perform intra-prediction or inter-prediction on blocks of the current picture. When performing inter-prediction on blocks, the picture estimation unit 226 may determine one or more motion vectors for the block. For example, the picture estimation unit 226 may determine that a particular block of the current picture has a motion vector of magnitude m relative to a reference block in a reference picture that has a picture order count (POC) distance p1 from the current picture. In this example, the current picture and the next picture may have a POC distance of p2. The picture estimation unit 226 may determine a scale factor s as p2 / p1. The picture estimation unit 226 may then scale the motion vector of a particular block by s (e.g., s * m). The picture estimation unit 226 may determine the location in the next picture indicated by the scaled motion vector and set the sample at the determined location as a sample in a particular block of the current picture. The picture estimation unit 226 may repeat this process for each interpretation block of the current picture.

[0067] In some examples, the picture estimation unit 226 may apply one or more filters to the prediction data. For example, the picture estimation unit 226 may apply one or more deblocking filters, smoothing filters, adaptive loop filters, or other types of filters to the prediction data.

[0068] The video encoder 228 may perform the same limited video encoding process as the video encoder 210 on the video data generated by the picture estimation unit 226. For example, the video encoder 228 may perform intra-prediction to generate prediction data. The video encoder 228 may use the prediction data and corresponding blocks of the video data generated by the picture estimation unit 226 to generate residual data. The video encoder 228 may apply a transformation (e.g., DCT transformation, DST transformation, etc.) to the residual data to generate transformation coefficients. The video encoder 228 may then apply quantization to the transformed coefficients. Furthermore, the video encoder 228 may apply entropy coding to the syntax elements representing the transformation coefficients.

[0069] As briefly mentioned above, the channel decoder 222 can apply the channel decoding process to channel encoded video data. Figures 3A and 3B provide further information regarding the channel encoding process performed by the channel encoder 212 and the channel decoding process performed by the channel decoder 222.

[0070] Specifically, Figure 3A is a block diagram illustrating an exemplary channel coding process using the technique of the present disclosure. For each picture of the coded video data, the channel encoder 212 of the transmitting device 102 may apply a systematic coding operation, such as a low-density parity check (LDPC) coding operation, to the systematic bits for the picture in order to generate error correction data for the picture. The systematic bits of the picture may include coded video data of the picture generated by the video encoder 210.

[0071] In the example in Figure 3A, the error correction data is labeled "error corr. bit". For picture n, the channel encoder 212 may generate error correction data 300A based on systematic bit 302A. Similarly, for picture n+1, the channel encoder 212 may generate error correction data 300B based on systematic bit 302B. The channel encoder 212 may classify the video data pictures as anchor video pictures and non-anchor pictures. The channel encoder 212 may classify pictures so that anchor pictures occur periodically between pictures. In some examples, the channel encoder 212 may classify a picture as an anchor picture if the channel decoder 222 receives an indication (e.g., from the receiving device 104) that there is an error in the picture. For each anchor picture, the transmitting device 102 may transmit the encoded anchor picture and error correction data for the anchor picture. However, in the case of a non-anchor picture, the transmitting device 102 may only transmit error correction data for the non-anchor picture.

[0072] For example, in the example in Figure 3A, picture n may be an anchor picture, and picture n+1 may be a non-anchor picture. Therefore, the transmitting device 102 can transmit systematic bit 302A for picture n, error correction data 300A for picture n, and error correction data 300B for picture n+1, but it cannot transmit systematic bit 302B for picture n+1.

[0073] Figure 3B is a block diagram illustrating an exemplary channel decoding process according to the technique of the present disclosure. As described above, the channel decoder 222 can perform a channel decoding process on the encoded video data to reconstruct the encoded video data. When processing an anchor picture (e.g., picture n), the channel decoder 222 processes the systematic bits (SI in Figure 3B) of the anchor picture. n(as shown) and error correction data for the anchor picture (in Figure 3B, y n The anchor picture may be obtained from the depuncturing unit 220 (as indicated by the anchor picture). The systematic bits of the anchor picture may represent the encoded video data of the anchor picture. The channel decoder 222 may use error correction data for the anchor picture to detect and / or correct errors in the systematic bits of the anchor picture. The video decoder 224 may use the obtained error-corrected encoded video data for the anchor picture to reconstruct the anchor picture. The receiving device 104 may store the reconstructed picture, including the reconstructed anchor picture and the reconstructed non-anchor picture, in the decoded picture buffer 350.

[0074] When processing non-anchor pictures, the channel decoder 222 represents the encoded video data of the non-anchor picture (SI n+1 Systematic bits (shown as shown) can be obtained. The video encoder 228 of the receiving device 104 can generate encoded video data for non-anchor pictures based on the video data generated by the picture estimation unit 226. The channel decoder 222 generates error correction data for non-anchor pictures (y in Figure 3B). n+1The channel decoder 222 may then obtain the (indicated as) from the depuncturing unit 220. The channel decoder 222 may then perform the same channel decoding process that the channel decoder 222 applied when processing the anchor picture. Thus, the channel decoder 222 may use the error correction data for the non-anchor picture to detect and / or correct any “errors” in the systematic bits of the non-anchor picture. However, “errors” in the systematic bits of the non-anchor picture are not due to noise in channel 230 (as in the case of errors in the systematic bits of the anchor picture). Rather, “errors” in the systematic bits for the non-anchor picture may be due to the difference between the predicted version of the non-anchor picture and the original version of the non-anchor picture. Thus, the channel decoder 222 may use the transmitted error correction data for the non-anchor picture as a mechanism for “correcting” prediction errors.

[0075] The video encoder 210 of the transmitting device 102, the video decoder 224 of the receiving device 104, and the video encoder 228 of the receiving device 104 may perform video encoding and video decoding processes based on the values ​​of one or more sets of parameters. In other words, the parameter values ​​may control various aspects of the video encoding process performed by the video encoders 210, 228, and 224. In some examples, the parameters may include one or more of the following: (A parameter that indicates the color space, such as red-green-blue, Y-Cb-Cr, etc.) Pixel decimation parameters A parameter indicating the DCT size Transmitted DCT coefficient A parameter indicating the number of bits per DCT coefficient. Quantization parameters (for example, parameters indicating the quantization scheme such as linear or Max-Lloyd)

[0076] Each of the video encoder 210, video decoder 224, and video encoder 228 may need to use the same parameter value. Accordingly, according to one or more techniques of this disclosure, the transmitting device 102 may transmit the parameter value to the receiving device 104. The receiving device 104 may receive the transmitted parameter value. The video decoder 224 and video encoder 228 may use the parameter value in the video decoding process and the video encoding process.

[0077] In some examples, the transmitted parameter values ​​are static or semi-static. For example, in an example where the transmitted parameter value is static, the transmitting device 102 may transmit the parameter value to the receiving device 104 once, and the receiving device 104 may operate using the parameter value for an indefinite period of time. In an example where the transmitted parameter value is semi-static, the transmitting device 102 may update the parameter value from time to time and retransmit the updated parameter value to the receiving device 104.

[0078] The transmitting device 102 may transmit the parameter value in one of several ways. For example, in some cases, the transmitting device 102 may transmit the parameter value to the receiving device 104 using an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, or another type of message.

[0079] In an example where the transmitting device 102 transmits parameter values ​​to the receiving device 104, the video encoder 210 may segment each of the color components of the video data picture (e.g., R, G, and B components, Y, Cb, and Cr components) into equally sized (MxM) blocks. Examples of such blocks may include macroblocks (MBs) and max coding units (LCUs). The video encoder 210 of the transmitting device 102 calculates a transformation (e.g., 2D-DCT) for each block, and M 2can generate a number of conversion factors. The video encoder 210 can assign an order to the conversion factors of the block. For example, the video encoder 210 can order the conversion factors of the block according to a zigzag scan order starting from the most important conversion factor (e.g., the lowest frequency) and ending with the least important conversion factor (e.g., the highest frequency). The video encoder 210 can select the first N c conversion factors, where N c is a parameter value indicating the amount of transmitted conversion factors. The video encoder 210 can discard the unselected conversion factors.

[0080] Furthermore, the video encoder 210 can quantize the selected conversion factors. For example, if the selected conversion factor has an index i in the range from 0 to N c -1, the parameter can include bitwidth parameters (e.g., B i , i = 0, 1,..., N c -1) corresponding to different index values. For each selected conversion factor d i , the video encoder 210 can quantize the selected conversion factor d i using the following formula.

[0081]

Equation

[0082] The video encoder 228 of the receiving device 104 may generate prediction data for a picture, generate residual data based on the prediction data, and apply one or more transformations to the residual data to generate a transformation block containing transformation coefficients. The video encoder 228 may need to use the same bit width parameters as the video encoder 210 so that the channel decoder 222 can correctly associate specific systematic bits with the corresponding error correction data received from the depunchaging unit 220.

[0083] In some examples of this disclosure, the receiving device 104 may determine the value of one or more of the parameters without the transmitting device 102 transmitting the values ​​of these parameters to the receiving device 104. Examples in which the receiving device 104 determines the value of one or more of the parameters without the transmitting device 102 transmitting the values ​​to the receiving device 104 may enable a better compressive strain trade-off with lower control signaling overhead. For example, the receiving device 104 may have N c Without sending the value of the DCT coefficient (N) to the receiving device 104, c ) can be determined. For example, in this example, the receiving device 104 can determine N to achieve the desired peak signal-to-noise ratio (P-SNR). c The minimum value of can be determined. In other words, the receiving device 104 can determine the N that achieves the desired P-SNR. c The minimum value can be determined as follows:

[0084]

number

[0085] In another example, the receiving device 104 may determine the number of quantized bits based on a predicted picture, instead of receiving the number of quantized bits from the transmitting device 102. For example, in this example, the receiving device 104 may calculate a probability distribution for the quantized and non-quantized coefficients.

[0086]

number

[0087]

number

[0088] The receiving device 104 determines the number of quantized bits B based on the entropy ratio of the quantization coefficient and the non-quantization coefficient, as follows: i It is possible to determine this.

[0089]

number

[0090] Figure 4 is a flowchart illustrating exemplary operation of a transmitting device 102 using the techniques of the present disclosure. In the example of Figure 4, the video encoder 210 of the transmitting device 102 may acquire video data (400). For example, the video encoder 210 may acquire video data from a video source 120. Furthermore, the video encoder 210 may perform video coding on the video data to produce coded video data (402). For example, the video encoder 210 may apply intra-prediction to generate prediction data, generate residual data based on the prediction data and the original video data, and apply a transformation (e.g., DCT) to the block of residual data to generate a transformation block. The video encoder 210 may quantize the transformation coefficients of the transformation block. Furthermore, in some examples, the video encoder 210 may apply entropy coding to syntax elements representing the quantized transformation coefficients. In some examples, the video encoder 210 may implement a reconstruction loop to reconstruct the residual data, which may apply entropy decoding, inverse quantization, and one or more inverse transformations. The video encoder 210 may apply and use the prediction data and the reconstructed residual data to reconstruct the video data. In some examples, the video encoder 210 applies one or more filters, such as a deblocking filter, an adaptive loop filter, or a sample-adaptive offset filter, to the reconstructed video data. The video encoder 210 may use the reconstructed video data as reference data for intra-prediction.

[0091] The video encoder 210 may perform a video encoding process based on the values ​​of one or more parameters. For example, the video encoder 210 may quantize the conversion coefficients according to specific quantization parameters, or use a specific color space.

[0092] The channel encoder 212 of the transmitting device 102 may perform channel coding on the encoded video data to generate error correction data (404). The transmitting device 102 may transmit the encoded video data and error correction data to the receiving device 104, for example, via channel 230 (406). In some examples, the transmitting device 102 may selectively transmit portions of the encoded video data and transmit other portions of the encoded video data. For example, the transmitting device 102 may transmit the encoded video data of some pictures and not the encoded video data of other pictures. In another example, the transmitting device 102 may transmit the encoded video data of some transformation blocks of a picture and not the encoded video data of other transformation blocks of the picture. In some examples, the transmitting device 102 may transmit the most significant bit of a certain amount of transformation coefficients and not the less significant bits of the transformation coefficients.

[0093] In some examples, the transmitting device 102 may also transmit values ​​for one or more parameters to the receiving device 104. The parameter values ​​may control how the receiving device 104 reconstructs the video data. For example, the parameters may include a transformation size parameter indicating the size of the transformation blocks in the encoded video data generated by the video encoder 210. In this example, the receiving device 104 may need to interpret the received encoded video data according to the same transformation block size in order to properly reconstruct the video data. In other examples, the parameters may include parameters indicating the amount of transformation coefficients, bit width parameters, and so on.

[0094] Therefore, in the example of Figure 4, the transmitting device 102 can acquire video data from the video source. Based on a set of parameters, the transmitting device 102 can generate encoded video data for a first picture of the video data and encoded video data for a second picture of the video data. The transmitting device 102 can perform channel coding on the encoded video data for the first picture and the encoded video data for the second picture in order to generate error correction data for the first picture and error correction data for the second picture. The transmitting device 102 can transmit the encoded video data for the first picture, the error correction data for the first picture, and the error correction data for the second picture. In some examples, the transmitting device 102 can transmit parameter values ​​to the receiving device 104.

[0095] Figure 5 is a flowchart illustrating exemplary operation of a receiving device 104 using the technique of the present disclosure. In the example of Figure 5, the receiving device 104 may receive first encoded video data and first error correction data from the transmitting device 102 (500). The first encoded data represents one or more blocks of first pictures of video data. The first error correction data may provide error correction information relating to the blocks of first pictures.

[0096] The receiving device 104 may use the first error correction data to generate first error-corrected encoded video data in order to perform an error correction operation on the first encoded video data (502). For example, the channel decoder 222 of the receiving device 104 may use the first error correction data to perform low-density parity check (LDPC) coding on the first encoded video data. In other examples, the channel decoder 222 may use the error correction data in other error correction algorithms such as forward error correction (FEC) or turbo coding. Performing an error correction operation on the first encoded video data may remove errors introduced by noise in channel 230.

[0097] The video decoder 224 of the receiving device 104 may perform a first reconstruction operation to reconstruct a block of a first picture based on the first error-corrected encoded video data (504). The first reconstruction operation is controlled by the values ​​of one or more parameters. For example, the video decoder 224 may perform an inverse transform on the transformed block of the first error-corrected encoded video data to obtain residual data. Furthermore, in this example, the video decoder 224 may generate prediction data using, for example, intra-prediction. In this example, the video decoder 224 may use the prediction data and residual data to reconstruct the block of the first picture.

[0098] In some examples, video encoders 210 and 228 may use quantization parameters to quantize conversion coefficients generated based on prediction data for a picture, thereby producing encoded video data. When performing a reconstruction operation, video decoder 224 may use quantization parameters to dequantize the conversion coefficients of the error-corrected encoded video data. In some examples, transmitting device 102 and / or receiving device 104 may calculate quantization parameters based on the entropy ratio of quantized conversion coefficients to unquantized conversion coefficients, for example, as described above.

[0099] In some examples, the parameters include a transformation size parameter. As part of generating encoded video data, video encoders 210 and 228 may apply a forward transformation to the sample region data of the picture (e.g., predicted sample data or residual data) having a transformation size indicated by the transformation size parameter. As part of performing a reconstruction operation, video decoder 224 may apply an inverse transformation to the transformation coefficients of the error-corrected encoded video data having a transformation size indicated by the transformation size parameter.

[0100] In some examples, the parameters include a parameter indicating the amount of conversion coefficients. As part of generating encoded video data, video encoders 210 and 228 may include a set of conversion coefficients in the encoded video data that includes the indicated amount of conversion coefficients. When performing a reconstruction operation, video decoder 224 may parse a set of conversion coefficients from the error-corrected encoded video data that includes the indicated amount of conversion coefficients. Furthermore, in some examples, receiving device 104 may receive encoded video data and error-corrected data from transmitting device via a communication channel, and receiving device 104 may apply an optimization process that determines the number of conversion coefficients based on the signal-to-noise ratio of the data transmitted over the communication channel, for example, as described above.

[0101] In some examples, the parameters include bit width parameters for multiple index values. Performing a reconstruction operation for each of the multiple index values ​​may involve parsing a first set of bits from the error-corrected encoded video data. The first set of bits may represent conversion coefficients having each index value, and the amount of bits in the first set of bits is equal to the bit width indicated by the bit width parameter for each index value. As part of generating the encoded video data, video encoders 210 and 228 may include a second set of bits in the encoded video data. The second set of bits may represent conversion coefficients having each index value, and the amount of bits in the second set of bits is equal to the bit width indicated by the bit width parameter for each index value. Video decoder 224 may parse a third set of bits from the error-corrected encoded video data. The third set of bits may represent conversion coefficients having each index value, and the amount of bits in the third set of bits is equal to the bit width indicated by the bit width parameter for each index value.

[0102] Other parameters may include one or more of the following: color space, conversion size, quantization parameters, number of conversion coefficients in the first encoded video data, or number of bits per conversion coefficient in the first encoded video data.

[0103] The receiving device 104 may obtain second error correction data from the transmitting device 102 (506). The second error correction data provides error correction information relating to one or more blocks of a second picture of the video data.

[0104] Furthermore, the picture estimation unit 226 of the receiving device 104 may estimate a second picture based on one or more previously reconstructed pictures, such as the first picture (508). The estimated second picture includes predictions of blocks of the second picture of video data based at least partially on blocks of the first picture. For example, the picture estimation unit 226 may generate prediction data using interpretation, a combination of intrapretation and interpretation, or other video coding tools, as described elsewhere in this disclosure.

[0105] The video encoder 228 of the receiving device 104 may generate second encoded video data based on the estimated second picture (510). For example, the video encoder 228 of the receiving device 104 may generate residual data based on prediction data. For example, the video encoder 228 may perform intra-prediction to generate second prediction data based on the estimated second picture. The video encoder 228 may then generate residual data by subtracting the second prediction data from the prediction data generated by the picture estimation unit 226. The video encoder 228 may then generate a transformation block by applying one or more forward transformations to the residual data. The video encoder 228 may perform the same process as the video encoder 210 of the transmitting device 102 and therefore may need to use the same parameters as the video encoder 210.

[0106] The channel decoder 222 of the receiving device 104 may perform a channel decoding process to generate second error-corrected encoded video data based on second error-corrected data and second encoded video data (512). The channel decoder 222 of the receiving device 104 may perform the same process as the channel decoder 222 performed when generating the first error-corrected encoded video data to generate the second error-corrected encoded video data.

[0107] The video decoder 224 of the receiving device 104 may perform a second video decoding process to reconstruct a second block of the picture based on the second error-corrected encoded video data (514). The second reconstruction operation is controlled by the value of a parameter. The video decoder 224 may perform the second reconstruction operation in the same manner as the first reconstruction operation. In this way, the receiving device 104 may reconstruct the video data of a picture (or block) without receiving all of the encoded video data of each picture (or block).

[0108] As described above, the puncturing unit 214 of the transmitting device 102 may perform a bit puncturing operation on the error correction data generated by the channel encoder 212. Bit puncturing involves selectively discarding some of the error correction data before the transmitting device 102 transmits the error correction data. The discarded bits are usually not the most important for performing error correction. The depuncturing unit 220 of the receiving device 104 may perform an inverse bit puncturing operation (i.e., a bit depuncturing operation) which reverses the bit puncturing operation performed by the puncturing unit 214. The puncturing unit 214 may perform a bit puncturing operation according to one or more sets of puncturing parameters. In different examples, the puncturing parameters may be predefined, static, or semi-static.

[0109] According to one or more techniques of this disclosure, the transmitting device 102 may perform a decimation procedure that can reduce memory bandwidth and enhance compression. For example, the video encoder 210 of the transmitting device 102 may segment the picture of video data into a grid of blocks (e.g., MB, LCU, etc.) and generate a transformed block for each block. The transmitting device 102 may need to store each transformed block intended for transmission to the receiving device 104 in memory (e.g., memory 116). The transmitting device 102 may then retrieve the stored transformed blocks from memory for channel coding and finally for transmission. These writes to and reads from memory may increase time and energy requirements. These time and energy requirements may be directly related to the amount of data written and read. Therefore, it may be advantageous to reduce the amount of data written to and read from memory.

[0110] Performing a decimation process can reduce the amount of data in a translated block that is written to and read from memory. Performing a decimation process can also reduce the amount of data that is sent by the transmitting device 102 to the receiving device 104. In some examples, when the video encoder 210 is encoding the current block of the current picture, the video encoder 210 may generate a translated block for the current block. Furthermore, the channel encoder 212 of the transmitting device 102 may determine, based on a decimation pattern, whether the current block is subject to decimation. If the current block is subject to decimation (i.e., the translated block is a "non-anchor translated block"), the channel encoder 212 may reduce the number of bits in the non-anchor translated block before storing the translated block in memory. If the current block is not subject to decimation (i.e., the translated block is an "anchor block"), the channel encoder 212 does not reduce the number of bits in the anchor block. The channel encoder 212 may perform a decimation process after generating error correction data. Therefore, the error correction data generated by the channel encoder 212 for non-anchor conversion blocks (and potentially transmitted to the receiving device 104) may be based on the full set of bits of the conversion block instead of a reduced number of bits.

[0111] Figure 6 is a conceptual diagram showing an exemplary decimation pattern 600 using the technique of the present disclosure. The example in Figure 6 shows a grid of DCT blocks. A DCT block is a block of transformation coefficients generated by applying a DCT transformation to video data, such as residual data or sample data. In other examples, a DCT block may be a transformation block generated using other types of transformations. In Figure 6, the "X" marks in the decimation pattern 600 indicate the DCT blocks (i.e., non-anchor transformation blocks) that are subject to decimation. Thus, in the example in Figure 6, the decimation pattern 600 decimates the DCT blocks by 2 horizontally and vertically. In some examples, which may be referred to herein as "full" decimation, the video encoder 210 may reduce the number of bits in the non-anchor transformation blocks to 0.

[0112] Therefore, in some examples, the decimation pattern defines the pattern of anchored and unanchored converted blocks in the picture. The receiving device 104 may receive systematic bits of the anchored converted blocks but not systematic bits of the unanchored converted blocks. The systematic bits of the anchored converted blocks may represent the conversion coefficients in the anchored converted blocks. The systematic bits of the unanchored converted blocks may represent reduced bit-depth versions of the original conversion coefficients in the unanchored converted blocks. Error correction data may include error correction data for the anchored converted blocks and error correction data for the unanchored converted blocks. Error correction data for the unanchored converted blocks is based on the original conversion coefficients in the unanchored converted blocks. As part of generating error-corrected encoded video data, the channel decoder 222 may use the error correction data for the anchored converted blocks to perform error correction on the systematic bits of the anchored converted blocks. The channel decoder 222 may use the error correction data for the unanchored converted blocks to perform error correction on the portions of the encoded video data corresponding to the unanchored converted blocks. In some examples, the receiving device 104 may determine a decimation pattern and transmit the decimation pattern to the transmitting device 102.

[0113] In some examples, the transmitting device 102 stores encoded bits (encoded video data and error correction data) in a cyclic buffer. The transmitting device 102 uses two parameters to select which bits in the cyclic buffer to transmit. The first parameter is the starting position, and the second parameter indicates the number of consecutive bits to transmit. The starting position can have various values ​​to support selective transmission and non-transmission of system bits. The starting position may be selected to skip the transmission of certain systematic bits (i.e., bits of encoded video data) without skipping the transmission of error correction data. Thus, decimation of non-anchor conversion blocks can be achieved simply by manipulating the first and second parameters so that the transmitting device 102 does not transmit bits of the non-anchor conversion block.

[0114] In some examples, the channel encoder 212 does not reduce the number of bits in any of the targeted "non-anchor" transformation blocks to zero, but applies a hybrid decimation technique that reduces the number of bits in the transformation coefficients within the non-anchor transformation blocks. For example, in the example in Figure 6, the channel encoder 212 may reduce the number of bits in each transformation coefficient in the DCT block marked "X" by a predetermined number (e.g., 2, 4, 5, etc.). The channel encoder 212 does not reduce the number of bits in transformation coefficients that are not targeted by the decimation pattern.

[0115] The channel decoder 222 of the receiving device 104 may receive the remaining reduced bits of the non-anchor conversion block and error correction data for the non-anchor conversion block. As part of the channel decoding process, the channel decoder 222 may perform an error correction process to restore the removed bits of the non-anchor conversion block using the error correction data for the non-anchor conversion block. This error correction process may be the same as the error correction process that the channel decoder 222 uses to correct errors introduced by noise in channel 230. In summary, the channel encoder 212 generates error correction data because error correction data will be needed to correct the unavoidable noise in channel 230, but these same error correction data are used to restore bits as if the noise in channel 230 had occurred in such a way that it corrupted the least significant bit of a particular conversion coefficient in a particular conversion block in a particular picture. Thus, the number of bits transmitted in channel 230 may be effectively reduced.

[0116] In some examples, the channel encoder 212 may generate a correlation matrix based on the picture's set of transformation blocks before performing any decimation process on any non-anchor transformation block in the set of transformation blocks. The correlation matrix contains values ​​indicating the level of correlation between transformation coefficients at corresponding positions within the transformation blocks. For example, the correlation matrix may contain correlation values ​​for the DC transformation coefficients (i.e., the top-left transformation coefficients) of the set of transformation blocks. If the differences between the DC transformation coefficients are relatively small, the correlation values ​​for the DC transformation coefficients may be relatively high. Conversely, if the differences between the DC transformation coefficients are relatively large, the correlation values ​​for the DC transformation coefficients may be relatively small. Each correlation value can be between 0 and 1.

[0117] In some examples, the channel encoder 212 may calculate the correlation value of the DC conversion coefficient using the following formula.

[0118]

number

[0119] The transmitting device 102 may transmit the correlation matrix along with the encoded video data and error correction data to the receiving device 104. The channel decoder 222 of the receiving device 104 may use the correlation matrix as part of the process to restore the non-anchor conversion blocks to their original bit widths. For example, continuing with the DC conversion coefficient example, after applying the error correction process, the channel decoder 222 may obtain the values ​​of the non-anchor DC conversion coefficients (i.e., the DC conversion coefficients in the non-anchor conversion block).

[0120] The error correction process may use correlation values ​​to estimate non-anchor transformation coefficients. For example, if all even transformation blocks are anchor transformation blocks and non-even transformation blocks are non-anchor transformation blocks, the channel decoder 222 may estimate the value of the non-anchor transformation coefficient in transformation block n by the following:

[0121]

number

[0122] In some examples, the video encoder 210 may reduce the bit width of each conversion coefficient in the target conversion block by the same amount. In other examples, the video encoder 210 may reduce the bit width of different conversion coefficients in the target conversion block by different amounts. In some examples, the amount by which the video encoder 210 reduces the bit width of a conversion coefficient depends on the distance of the conversion coefficient from the anchor conversion block. The anchor conversion block is a conversion block that is not targeted by the decimation pattern.

[0123] Figure 7 is a flowchart illustrating exemplary operation of a transmitting device 102 for hybrid decimation of transform blocks using the technique of the present disclosure. In the example of Figure 7, the transmitting device 102 may acquire video data from a video source 120 (Figure 7) (700). The video encoder 210 of the transmitting device 102 may generate transform blocks based on the video data (702). For example, the video encoder 210 may generate a prediction block by performing an intra-prediction on a block of pictures in the video data. The video encoder 210 may use the prediction block to generate residual data. The video encoder 210 may generate a transform block by applying a transform, such as DCT, DST, or other transform, to the residual data. In another example, the video encoder 210 may generate a transform block by directly applying a transform to a block of video data.

[0124] The channel encoder 212 may then determine, based on the decimation pattern, which of the transformation blocks are anchor transformation blocks (704). For example, in one example where the channel encoder 212 uses the decimation pattern 600 of Figure 6, the video encoder 210 may determine that every other transformation block in both the horizontal and vertical directions is an anchor transformation block. In other examples, the channel encoder 212 may use other decimation patterns to determine which of the transformation blocks are anchor transformation blocks. The channel encoder 212 may store the anchor transformation blocks in the memory of the transmitting device 102 (e.g., memory 116 (Figure 1)) (706).

[0125] The channel encoder 212 can calculate the correlation matrix of the transformation block sets (708). Each transformation block set includes one or more anchored transformation blocks and one or more non-anchored transformation blocks. For example, each transformation block set may correspond to a different row of transformation blocks in Figure 6. In another example, each transformation block set may correspond to a group of 2x2 transformation blocks. The number of values ​​in the correlation matrix of a transformation block set is individually the same as the number of transformation coefficients for each transformation block. Each value in the correlation matrix corresponds to a different position within the transformation coefficient block. For example, the value at position (0,0) in the correlation matrix corresponds to the transformation coefficient at position (0,0) of each transformation coefficient block in the transformation block set, and the value at position (0,1) in the correlation matrix corresponds to the transformation coefficient at position (0,1) of each transformation coefficient block in the transformation block set, and so on. The transformation coefficient at position (0,0) of the transformation coefficient block may be called the DC coefficient, and all other transformation coefficients may be called the AC coefficient.

[0126] Furthermore, the channel encoder 212 may generate a bit-reduced non-anchor transformation matrix (710). The non-anchor transformation matrix is ​​a transformation matrix other than the anchor transformation matrix. For example, referring to Figure 6, the transformation matrix marked with X may be a non-anchor transformation matrix. The transformation coefficients in the bit-reduced non-anchor transformation matrix may contain fewer bits than the original version of the non-anchor transformation matrix. The transmitting device 102 may then transmit the anchor transformation block, the non-anchor transformation block, the bit-reduced value, the correlation matrix, and the error correction data (712).

[0127] The channel encoder 212 may reduce the number of bits in the non-anchor conversion coefficients in one of several ways. For example, the channel encoder 212 may determine a bit reduction value for each conversion coefficient in the non-anchor conversion coefficient block. In this example, to calculate the bit reduction value, the channel encoder 212 may calculate the interpolated value of the conversion coefficient according to the correlation matrix. The channel encoder 212 may calculate the interpolated value using equation (6) above. The channel encoder 212 may then subtract the interpolated value of the conversion coefficient from the original value of the conversion coefficient to calculate a first strain value. The channel encoder 212 may then reduce the number of bits in the original value of the conversion coefficient by 1. The channel encoder 212 may then subtract the interpolated value from the reduced bit original value of the conversion coefficient to calculate a second strain value. Based on the first and second strain values, the channel encoder 212 may determine whether the second strain value is acceptable. For example, the channel encoder 212 may calculate the mean or greatest squares error from the interpolated value. The channel encoder 212 may determine that the second strain value is acceptable by comparing it with a predefined threshold value.

[0128] If the second distortion value is acceptable, the channel encoder 212 may reduce the number of bits in the original value of the conversion coefficient and repeat the process. If the second distortion value is unacceptable, the channel encoder 212 may increase the number of bits in the original value of the conversion coefficient. The resulting number of bits in the original value of the conversion coefficient is the bit reduction value.

[0129] In some examples, the channel encoder 212 may determine the bit reduction value for an entire block of non-anchor conversion coefficients. In this example, to calculate the bit reduction value for the block of conversion coefficients, the channel encoder 212 may calculate the interpolated value of each conversion coefficient according to the correlation matrix, for example, as described above. The video encoder 210 may then subtract the interpolated value of the conversion coefficient from the original value of the conversion coefficient and use the resulting difference to calculate a first distortion value. For example, the video encoder 210 may calculate the first distortion value as the mean squared error. The video encoder 210 may then reduce the number of bits of each original value of the conversion coefficient by 1. The video encoder 210 may subtract the interpolated value from the reduced bit original value of the conversion coefficient and use the resulting value to calculate a second distortion value (for example, using the mean squared error). Based on the first and second distortion values, the channel encoder 212 may determine whether the second distortion value is acceptable. If the second distortion value is acceptable, the video encoder 210 may again reduce the number of bits in the original value of the conversion coefficient and repeat the process. If the second distortion value is unacceptable, the video encoder 210 may increase the number of bits in the original value of the conversion coefficient. The resulting number of bits from the reduction of the original value of the conversion coefficient is the bit reduction value.

[0130] Therefore, in some examples, the transmitting device 102 may acquire video data from a video source. The transmitting device 102 may generate transformation blocks based on the video data. The transmitting device 102 may determine which of the transformation blocks are anchor transformation blocks. The transmitting device 102 may compute a correlation matrix for the set of transformation blocks. Furthermore, the transmitting device 102 may generate a bit-reduced non-anchor transformation matrix. The transmitting device 102 may send the anchor transformation blocks, the non-anchor transformation blocks, and the correlation matrix to the receiving device. In some examples, the transmitting device 102 may receive an indication of a decimation pattern from the receiving device 104.

[0131] Figure 8 is a flowchart illustrating exemplary operation of a receiving device 104 for hybrid decimation of a transform block using the technique of the present disclosure. In the example of Figure 8, the receiving device 104 may receive an anchored transform block, a non-anchored transform block, one or more bit reduction values, and a correlation matrix for the non-anchored block (800).

[0132] Furthermore, in the example in Figure 8, the channel decoder 222 of the receiving device 104 can calculate an interpolated value of the current non-anchor transformation coefficient (802). The current non-anchor transformation coefficient is one of the transformation coefficients in the non-anchor transformation block. The channel decoder 222 can calculate an interpolated value of the current non-anchor transformation coefficient based on a correlation matrix for the non-anchor block. The channel decoder 222 can calculate an interpolated value of the current non-anchor transformation coefficient in one of several ways. For example, in some examples, the channel decoder 222 may apply a machine learning model that takes one or more non-anchor transformation coefficients (including the current non-anchor transformation coefficient), one or more anchor transformation coefficients, and a correlation matrix for the non-anchor transformation coefficient block as input. In this example, the machine learning model may output an interpolated value of the current non-anchor transformation coefficient. In this example, the machine learning model may be implemented as a neural network model, a support vector machine, a regression model, or another type of machine learning model.

[0133] In another example, the correlation matrix may include values ​​showing the correlation between the current non-anchor transformation coefficients and each corresponding anchor transformation coefficient in one or more blocks of anchor transformation coefficients. The video decoder 224 may calculate the interpolated values ​​of the current non-anchor transformation coefficients as follows:

[0134]

number

[0135] Furthermore, the video decoder 224 may calculate the reconstructed value of the non-anchor conversion coefficient (804). The video decoder 224 may calculate the reconstructed value of the non-anchor conversion coefficient based on the interpolated value of the current non-anchor conversion coefficient and the transmitted value of the non-anchor conversion coefficient. The transmitted value of the non-anchor conversion coefficient is included in the received non-anchor conversion block. In some examples, the video decoder 224 calculates the reconstructed value of the non-anchor conversion coefficient as the average of the interpolated value of the current non-anchor conversion coefficient and the transmitted value of the non-anchor conversion coefficient.

[0136] The video decoder 224 may determine whether there are any remaining non-anchor conversion coefficients in the non-anchor conversion block (806). If there are one or more remaining non-anchor conversion coefficients in the non-anchor conversion block (the "yes" branch of 806), the video decoder 224 may repeat steps 802-806 with another of the non-anchor conversion coefficients. The video decoder 224 may continue this operation until there are no more remaining non-anchor conversion coefficients (the "no" branch of 806). In this way, the video decoder 224 may calculate the reconstructed value for each of the non-anchor conversion coefficients.

[0137] In this way, the receiving device 104 can receive systematic bits of the anchored transformation block, systematic bits of the non-anchored transformation block, and a correlation matrix. The systematic bits of the anchored transformation block may represent the transformation coefficients in the anchored transformation block. The systematic bits of the non-anchored transformation block may represent a reduced bit depth version of the original transformation coefficients in the non-anchored transformation block. As part of performing a reconstruction operation, the receiving device 104 may calculate an interpolated value of each non-anchored transformation coefficient in the non-anchored transformation block based on the correlation matrix and the corresponding anchored transformation coefficient. The receiving device 104 may calculate a reconstructed value of the non-anchored transformation coefficient based on the interpolated value of the non-anchored transformation coefficient and the value of the non-anchored transformation coefficient in the error-corrected encoded video data.

[0138] In some examples, the receiving device 104 may adaptively select a decimation pattern to be used to reduce or remove bits in a particular conversion block. In such examples, the receiving device 104 may communicate the selected decimation pattern back to the transmitting device 102. The transmitting device 102 may then use the selected decimation pattern in one or more pictures of the video data.

[0139] Figure 9 is a conceptual diagram showing an exemplary decimation pattern 900 adaptively selected by the receiving device 104 using one or more techniques of the present disclosure. In contrast to the decimation pattern 600 of Figure 6, the non-anchor blocks in the decimation pattern 900 do not necessarily occur at regular intervals or spacings.

[0140] The receiving device 104 may determine a decimation pattern based on information about the previous picture in the video data. The previous picture may be decimated or not. In some examples, the receiving device 104 may send a request to the transmitting device 102 for the undecimated version of the picture. The transmitting device 102 may respond to the request by sending the undecimated version of the picture to the receiving device 104. After receiving the undecimated version of the picture, the receiving device 104 may determine a decimation pattern based on the undecimated version of the picture. For example, the receiving device 104 may perform a rate-distortion optimization process that evaluates several potential decimation patterns to identify which of the decimation patterns results in the best combination of bitrate and distortion.

[0141] In some examples, the transmitting device 102 and the receiving device 104 may continue to use a selected decimation pattern for a predetermined number of pictures, after which the transmitting device 102 and / or the receiving device 104 may adaptively select a different decimation pattern. In some examples, the transmitting device 102 may send a message to the receiving device 104 requesting that the transmitting device 102 select a different decimation pattern. In some examples, the receiving device 104 may determine that an event or condition has occurred that would favor selecting a different decimation pattern. For example, the receiving device 104 may determine that it would be favorable to select a different decimation pattern when it determines that a scene change has occurred, when it determines that motion in the video data has exceeded one or more thresholds, or when it determines that other characteristics of the video data have changed.

[0142] The receiving device 104 may signal the selected decimation pattern to the transmitting device 102 in one of several ways. For example, the receiving device 104 may signal the selected decimation scheme to the transmitting device 102 by indicating the difference from an existing decimation scheme, such as the decimation scheme currently in use. For example, in this example, the receiving device 104 may indicate the selected decimation scheme to the transmitting device 102 by specifying a change in downsampling or upsampling along a particular axis, specifying a change in a particular area of ​​the picture, deleting a particular transformation block, or enabling a particular transformation block.

[0143] In some examples, there may be a predefined mapping of index values ​​to predefined decimation patterns. In such examples, the receiving device 104 may select a decimation pattern from the predefined decimation patterns and signal the index value of the selected decimation pattern to the transmitting device 102.

[0144] Figure 10 is a block diagram showing exemplary components of a transmitting device and a receiving device according to the technique of the present disclosure. In the example of Figure 10, the transmitting device 102 may include the same components as those shown in Figure 2. However, in the example of Figure 10, the receiving device 104 may further comprise a reliability unit 1002. Unless otherwise noted, the similarly named components of the transmitting device 102 and the receiving device 104 in Figures 2 and 10 perform the same function.

[0145] In general, accurately predicting the most significant bit (MSB) of a conversion coefficient can be easier than predicting the least significant bit (LSB). This is because the MSB is converted to a higher Euclidean distance in the video image region. Furthermore, blocks of a video picture experiencing high motion can be more difficult to predict accurately than blocks in the low motion region.

[0146] According to one or more techniques of this disclosure, the transmitting device 102 and the receiving device 104 may implement a system in which bit-level reliability values ​​are used. The use of bit-level reliability values ​​may enable the transmitting device 102 and the receiving device 104 to correctly weight a priori information. This may result in improved system performance (e.g., reduced amount of transmitted data and / or improved video quality). For example, if "soft" information is used, decoding performance may be improved. In other words, if the receiving device 104 is "notified" of the reliability of each bit (what the prior probability is that the bit value is "0" or "1"), the receiving device 104 can utilize this information and performance may be improved.

[0147] In the example shown in Figure 10, the reliability unit 1002 of the receiving device 104 receives encoded video data from the video encoder 228 of the receiving device 104. The encoded video data may include conversion coefficients for the video data conversion block. Furthermore, in some examples, the reliability unit 1002 receives predicted quality information from the picture estimation unit 226 of the receiving device 104.

[0148] In a constant scaling process, for each bit position of the conversion coefficients in a conversion block, the prediction quality information includes a confidence value for that bit position. For example, the most significant bit of the conversion coefficient has a first confidence value, the second most significant bit has a second confidence value, the third most significant bit has a third confidence value, and so on. The confidence value of a bit position is a measure of how likely it is that the bit at that position has an incorrect value. For example, a bit at a bit position may have an incorrect value if the bit value is predicted to be "0" but the actual value is "1", or if the bit is predicted to be "1" but the actual value is "0".

[0149] The picture estimation unit 226 may determine the reliability value of a bit position by collecting statistics on the rate of errors occurring in bits at that bit position. For example, the picture estimation unit 226 may determine the probability that the most significant bit of the conversion coefficient contains an error, the probability that the second most significant bit of the conversion coefficient contains an error, and the probability that the third most significant bit of the conversion coefficient contains an error. The picture estimation unit 226 may collect these statistics by counting the number of times the predicted bit (i.e., the bit in the estimated picture) was inaccurate out of the number of events tested. For example, the picture estimation unit 226 may estimate a picture, and the video encoder 228 may encode video data of the estimated picture. The video decoder 224 may then decode the error-corrected encoded video data of the picture. The picture estimation unit 226 may compare the bits in the conversion coefficient of the encoded video data of the estimated picture with the error-corrected video data of the picture to determine whether the bits in the encoded video data of the estimated picture are incorrect.

[0150] The reliability unit 1002 can convert probability values ​​to LLR values. In some examples, the reliability unit 1002 can convert probability values ​​to LLR values ​​using the following formula.

[0151]

number

[0152] Figure 11 shows a chart of exemplary error probabilities and corresponding log-likelihood ratio (LLR) absolute values ​​using one or more techniques of the present disclosure. In the example in Figure 11, each transform block is represented using 150 bits. Graph 1100 plots the error probabilities for individual bit positions within the transform block. As can be seen from Graph 1100, bits at certain positions have a higher probability of error. Graph 1102 shows the error probabilities transformed into LLR absolute values. LLR absolute values ​​can be scaled.

[0153] In some examples, the receiving device 104 uses a dynamic scaling process. In the dynamic scaling process, the picture estimation unit 226 dynamically determines prediction quality information based on the video data. For example, some areas of the picture are more difficult to predict than areas that are easier to predict (e.g., areas where prediction accuracy is reduced). Examples of areas of the picture that are more difficult to predict may include areas with higher motion. For example, if the sum of the motion vectors in an area exceeds a threshold, the picture estimation unit 226 may determine that the area is a difficult-to-predict area. Thus, the picture estimation unit 226 may identify such areas and generate confidence values ​​for bits of the transformation coefficients of the transformation block, at least partially based on whether the transformation block is inside or outside such areas. In some examples, the picture estimation unit 226 determines confidence values ​​for bits of the transformation coefficients based on general statistics of errors about the bit positions corrected based on whether the transformation block containing the transformation coefficients is in a difficult-to-predict area.

[0154] The reliability unit 1002 may use reliability values ​​to scale the bits of the conversion coefficients in the encoded video data generated by the video encoder 228. For example, bits of the encoded video data generated by the video encoder 228 may be considered "hard" bits and may have values ​​of exactly 0 or exactly 1. The reliability unit 1002 may determine "soft" bits between 0 and 1, which are the predicted quality information generated by the picture estimation unit 226 and the encoded video data generated by the video encoder 228. For example, if the value of a bit in the encoded video data is 1, the reliability unit 1002 may generate a "soft" value of the bit by multiplying the LLR absolute value of the bit by a negative 1 (i.e., -1). If the value of a bit in the encoded video data is 0, the reliability unit 1002 may generate a "soft" value of the bit by multiplying the LLR absolute value of the bit by a positive 1 (i.e., +1). Thus, the "soft" or scaled value of a bit of the conversion coefficient may be M or -M.

[0155] Therefore, in some examples, each bit can be converted to a scaled value where it has a more positive value if there is greater confidence that the bit has a value of 0, and a more negative value if there is greater confidence that the bit has a value of 1. The confidence unit 1002 provides the scaled value to the channel decoder 222 as a priori information.

[0156] The channel decoder 222 performs the channel decoding process using scaled values. For example, the reliability unit 1002 and the channel encoder 212 may encode the encoded video data into a codeword (e.g., a low-density parity check code). The bits of the codeword generated by the reliability unit 1002 may be scaled as described above. The bits of the codeword generated by the channel encoder 212 may be modified during passage through channel 230 so that the bits of the codeword may be received as values ​​between -1 and 1. The channel decoder 222 may apply the LPDC decoding process to the codeword to correct errors in the codeword. The bit values ​​in the corrected codeword are 0 or 1. The LPDC decoding process may then convert the codeword from the encoded video data to the original data. In other examples, other coding schemes may be used. In this example, the error-corrected data received by the channel decoder 222 may include cyclic redundancy check (CRC) data that is not used in the LPDC decoding process. In this way, the channel decoder 222 may determine the value of each bit of the conversion coefficient.

[0157] In some examples, the channel encoder 212 may use predictive quality feedback to sort bits of the conversion coefficients before applying unequal protection channel coding, such as polar coding or spinal coding. Unequal protection channel coding involves assigning coding redundancy according to the importance of the information bits. For example, the channel encoder 212 may use reliability data to protect different bits according to the predictability by the receiving device 104.

[0158] In some examples, the reliability unit 1002 sends predictive quality feedback to the transmitting device 102. The video encoder 210 may adjust one or more encoding parameters of the video encoding process that the video encoder 210 applies to the video data. For example, the transmitting device 102 may determine the compression ratio of a limited video encoding process based on the predictability (reliability) of the receiving device 104. For example, if the predictability is low, the transmitting device 102 may reduce the quality of the encoded video by reducing the number of conversion coefficients or the bit width per conversion coefficient so that less information is communicated.

[0159] In some examples, the transmitting device 102 may adjust one or more channel coding parameters used by the channel encoder 212 based on predictive quality feedback. For example, the channel encoder 212 may use unequal protection codes, where protection depends on predictive quality feedback. In some examples, the channel encoder 212 may select different LDPC graphs depending on reliability. In some examples, the channel encoder 212 may change the encoding scheme for generating error-corrected data to increase the error correction capability for bit positions or areas of pictures with lower reliability, or to decrease the error correction capability for bit positions or areas of pictures with higher reliability.

[0160] In some examples, the transmitting device 102 may update one or more bit puncturing parameters used by the puncturing unit 214 based on predictive quality feedback. For example, the puncturing unit 214 may modify the puncturing pattern to enable the transmission of more error correction data for bit locations and / or picture regions with lower reliability. Thus, the transmitting device 102 may avoid puncturing bits with lower reliability. In some examples, the puncturing unit 214 may modify the puncturing pattern to enable the transmission of less error correction data for bit locations and / or picture regions with higher reliability.

[0161] In some examples, the predictive quality feedback that the reliability unit 1002 sends to the transmitting device 102 is applicable to the entire picture. In some examples, the reliability unit 1002 may send predictive quality feedback to the transmitting device 102 on a region-by-region basis. Each region may be a defined area within the picture. The reliability unit 1002 may send predictive quality feedback for some regions of the picture but not for others.

[0162] In some examples, the reliability unit 1002 may periodically send predictive quality feedback to the transmitting device 102. For example, the reliability unit 1002 may send predictive quality feedback to the transmitting device 102 every N pictures, where N is an integer. In some examples, the reliability unit 1002 sends predictive quality feedback to the transmitting device 102 after the completion of a certain number of picture groups (GOPs). In other examples, the reliability unit 1002 may send predictive quality feedback aperiodically, such as in response to a specific condition or event.

[0163] In some examples, predictive quality feedback may include predictive quality data based on one or more noise models, such as a Gaussian noise model or a Laplacian noise model. Noise model parameters may control one or more noise models. Reliability unit 1002 may transmit noise model parameters to transmitting device 102. The use of noise models is an alternative for collecting error statistics. In this mode, the predictive error (a statistical value of the error between the estimated picture estimated by picture estimation unit 226 and the actual picture reconstructed by video decoder 224) may be modeled using several parameters that describe the error distribution function. Since noise model parameters may contain less data compared to bitwise statistics, noise model parameters may be easier for receiving device 104 to communicate to transmitting device 102. Transmitting device 102 may use noise models in the same way that transmitting device 102 may use other types of predictive quality feedback.

[0164] In some examples, instead of the reliability unit 1002 receiving predictive quality information from the picture estimation unit 226 of the receiving device 104, the video encoder 210 may generate the predictive quality information and send it to the reliability unit 1002. The video encoder 210 may determine the predictive quality information based on a priori information about the video coding process. For example, the video encoder 210 may evaluate bitwise reliability (e.g., MSB is more reliable than LSB, lower frequency conversion coefficients are more reliable than higher frequency conversion coefficients) based on compression parameters. The video encoder 210 may perform a prediction process to evaluate statistics itself and have predefined statistics for different light compression sets of parameters. Furthermore, the transmitting device 102 may use other sensors to evaluate instantaneous motion and adjust the reliability accordingly.

[0165] Figure 12 is a flowchart illustrating exemplary operation of a transmitting device 102 using scaled bits according to the technique of the present disclosure. In the example of Figure 12, the transmitting device 102 may, for example, acquire video data from a video source 120 (1200). Furthermore, the transmitting device 102 may acquire predictive quality feedback (1202). In some examples, the transmitting device 102 may acquire predictive quality feedback from a receiving device 104. The predictive quality feedback includes bit reliability information. In some examples, the predictive quality feedback is expressed in terms of noise model parameters, such as parameters of a Gaussian noise model or a Laplacian noise model.

[0166] The transmitting device 102 may adapt one or more of the video coding parameters, channel coding parameters, or bit puncturing parameters based on predictive quality feedback (1204). The video encoder 210 of the transmitting device 102 may perform a video coding process to generate coded video data (1206). The video coding process may be controlled by video coding parameters. For example, video coding parameters may control the number of conversion coefficients included in a conversion block, the number of bits included in the conversion coefficients, etc. In some examples, video coding parameters include quantization parameters, and the video encoder 210 may adapt the quantization parameters based on predictive quality feedback. For example, if the predictive quality feedback indicates low reliability, the quantization parameters may be reduced to lower the level of quantization. As part of performing the video coding process, the video encoder 210 may use quantization parameters to quantize the conversion coefficients of the conversion blocks of one or more pictures.

[0167] The channel encoder 212 of the transmitting device 102 may perform a channel coding process on the scaled bits to generate channel-coded data (1208). The channel coding process may be controlled by channel coding parameters. For example, the channel coding parameters may include control over which LDPC graph the channel coding process uses to generate the codeword, and the channel coding parameters may control the error correction capability, and so on. For example, the channel coding parameters may include an LDPC graph, and the channel encoder 212 may adapt an LDPC graph and use the LDPC graph to generate a codeword for transmission to the receiving device.

[0168] Furthermore, the puncturing unit 214 of the transmitting device 102 may perform a bit puncturing process on the error correction data generated by the channel encoder 212 (1210). The bit puncturing process may be controlled by bit puncturing parameters. For example, predictive quality feedback may indicate that certain parts of the encoded video data are less reliable. Therefore, the transmitting device 102 may adjust the bit puncturing parameters to reduce bit puncturing on the error correction data for the less reliable parts of the encoded video data. The transmitting device 102 may transmit the channel-encoded data and the bit-punctured error correction data to the receiving device 104 (1212).

[0169] Figure 13 is a flowchart illustrating exemplary operation of a receiving device 104 using scaled bits according to the technique of the present disclosure. In the example of Figure 13, the receiving device 104 may receive error correction data from the transmitting device (1300). The error correction data provides error correction information about the picture of the video data.

[0170] The picture estimation unit 226 may generate prediction data for the picture (1302). The prediction data for the picture may include predictions of blocks of the picture that are at least partially based on one or more previously reconstructed pictures of the video data. For example, the picture estimation unit 226 may use interpretation and / or intrapretation to generate block predictions.

[0171] Furthermore, the receiving device 104 may generate encoded video data based on predictive data for the picture (1304). For example, the video encoder 228 of the receiving device 104 may perform a video encoding process that generates encoded video data. The encoded video data includes a transformation block having transformation coefficients.

[0172] The receiving device 104 may scale the bits of the conversion coefficients of the conversion block based on a confidence value for bit positions (1306). In some examples, the receiving device 104 generates a confidence value. For example, the receiving device 104 may generate a confidence value based on statistics regarding the occurrence of errors at bit positions. In some examples, the receiving device 104 may generate a confidence value based on the confidence characteristics of individual regions of a picture of video data. In some examples, the receiving device 104 may generate a confidence value based on a noise model. Furthermore, in some examples, the receiving device 104 may send the confidence value to the transmitting device 102. In other examples, the receiving device 104 may receive a confidence value from the transmitting device 102.

[0173] Furthermore, the channel decoder 222 of the receiving device 104 may use error correction data to generate error-corrected encoded video data in order to perform error correction operations on the scaled bits of the conversion coefficients of the conversion block (1308).

[0174] The video decoder 224 of the channel decoder 222 can reconstruct the picture based on the error-corrected encoded video data (1310).

[0175] This disclosure describes techniques that can reduce the complexity of video encoding in transmitting devices such as extended reality (XR) headsets. A transmitting device may acquire multiview video data, which may include pictures from two or more viewpoints. For example, an XR headset may include two cameras for stereoscopic viewing of the scene the user is seeing. In this example, the multiview video data may include pictures from each of the cameras.

[0176] Processing the content of multi-view video data can consume significant processing resources. For example, in the context of augmented reality (AR) or mixed reality (MR), considerable processing resources may be required to determine where virtual elements should be placed and how they should appear. Multi-view video data can help in processing virtual elements. For instance, the same virtual element may need to be darker when it should be placed in a shaded area of ​​the scene and brighter when it should be placed in a sunny area of ​​the scene. As another example, the system may need to analyze the content of the scene to determine whether a virtual element is obscured by a physical element in the scene, such as a rock or a tree. Multi-view video data can be useful in determining the depth of objects in a scene. To keep the XR headset bright and conserve battery power, it may be desirable to minimize the processing of video data performed on the XR headset. Therefore, processing video data on another device, such as the user's smartphone or another nearby device, can help reduce the demand on processing resources for the XR headset.

[0177] While multiview video data can be very useful in certain situations, simply transmitting unencoded multiview video data can be impractical because the amount of data required to transmit multiple parallel streams of video data simultaneously can be very large. However, there is often considerable redundancy between the pictures of different views in multiview video data. For example, what a person sees with their left eye is often not so different from what their right eye sees. Therefore, video compression techniques have been developed to reduce this redundancy in order to reduce the amount of data required to transmit multiview video data.

[0178] However, some of the techniques for multiview video coding require considerable computational resources. For example, a video encoder may determine the difference between two or more sets of simultaneous pictures to determine a depth map of a scene. A depth map is an array of values ​​indicating the depth / distance from the camera to the objects shown in the pictures. In this example, one of the simultaneous pictures may be an anchor picture, and one or more of the simultaneous pictures may be non-anchor pictures. The video encoder may use the depth map to compute the disparity vector for blocks in the non-anchor pictures. The disparity vector for a block indicates the lateral displacement between the block and the corresponding block in another simultaneous picture, such as an anchor picture. Generally, blocks representing deeper objects have smaller disparity vectors than blocks representing closer objects. The video encoder may use the disparity vectors of the blocks to determine predicted blocks, generate residual data based on the predicted blocks, apply transformations to the residual data, quantize the transformation coefficients of the resulting transformed blocks, and signal the quantized transformation coefficients. In this example, considerable computational resources may be involved in generating the depth map.

[0179] In another example, illumination levels may differ between simultaneous pictures from different viewpoints. These illumination differences can impair the coding efficiency of multiview video coding. Therefore, illumination compensation may be applied to non-anchor pictures to temporarily adjust their illumination levels according to an illumination compensation factor in order to better match the illumination levels of non-anchor pictures to those of anchor pictures during video coding. The original illumination levels of non-anchor pictures can be restored during video decoding. Determining the illumination compensation factor can consume computational resources.

[0180] The techniques of this disclosure may shift some of the processing associated with multiview video coding from a transmitting device (e.g., an XR headset) to a receiving device (e.g., a mobile device). For example, the transmitting device may acquire a first set of multiview pictures of video data. The first set of multiview pictures includes a first picture and a second picture. The first picture is from a first viewpoint, and the second picture is from a second viewpoint. The transmitting device may transmit first encoded video data to the receiving device. The first encoded video data is based on the first set of multiview pictures. The transmitting device may receive a multiview coding queue from the receiving device. Furthermore, the transmitting device may acquire a second set of multiview pictures of video data. The second set of multiview pictures includes a third picture and a fourth picture. The third picture is from a first viewpoint, and the fourth picture is from a second viewpoint. The transmitting device may perform a multiview encoding process on a second set of multiview pictures to generate second encoded video data, based on a multiview encoding queue received from the receiving device. The multiview encoding process reduces interview redundancy between the third and fourth pictures. The transmitting device may then transmit the second encoded video data to the receiving device.

[0181] Similarly, a receiving device may receive first encoded video data from a transmitting device. The first encoded video data is based on a first set of multiview pictures of the video data. The first set of multiview pictures may include a first picture and a second picture. The first picture is from a first viewpoint, and the second picture is from a second viewpoint. The receiving device may determine a multiview coding queue based on the first encoded video data. The receiving device may send the multiview coding queue to the transmitting device. In addition, the receiving device may receive second encoded video data from the transmitting device. The second encoded video data is based on a second set of multiview pictures including a third picture and a fourth picture. The second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.

[0182] Since the receiving device determines the multiview coding queue and sends it to the transmitting device, the burden of determining the multiview coding queue can be shifted from the transmitting device to the receiving device. This can reduce the demand on resources at the transmitting device.

[0183] Referring to Figure 2, the video encoder 210 may perform a multiview coding process based on a multiview coding queue acquired from the receiving device 104. For example, the multiview coding queue may include a depth map. In this example, the video encoder 210 may use the depth map to estimate the disparity vector for blocks of pictures in non-anchor views of the multiview video data. The video encoder 210 may use the disparity vector of the current block of the current picture to determine a predicted block for the current block based on a sample of the concurrently referenced picture. The concurrently referenced picture has the same picture order count (POC) value as the current picture. The video encoder 210 may determine the residual data for the current block based on the original sample of the current block and the predicted block of the current block. The video encoder 210 may apply one or more transformations to the residual data to generate one or more transformed blocks. The video encoder 210 may quantize the transformation coefficients in the transformed blocks. The coded video data generated by the video encoder 210 may be based on the quantized transformation coefficients.

[0184] In some examples, a multiview coding queue may include one or more illumination compensation factors. When coding the current picture of multiview video data, the video encoder 210 may modify each sample of the current picture based on one or more illumination compensation factors. In some examples, different illumination compensation factors may be applied to different regions of the current picture. Modifying the samples of the current picture in this way may better match the illumination level of the current picture with the illumination level of the concurrently referenced picture. After modifying the samples of the current picture, the video encoder 210 may code blocks of the current picture by performing the multiview coding process as described in the previous paragraph.

[0185] Furthermore, according to some examples of this disclosure, the video decoder 224 may perform a multiview decoding process. For example, the video decoder 224 may use the disparity vector of the blocks of the current picture to generate predicted blocks. The video decoder 224 may use the predicted block and residual data received from the channel decoder 222 to reconstruct the samples of the blocks of the current picture. In some examples where the video encoder 210 has applied illumination compensation to the picture, the video decoder 224 may use illumination parameters to invert the illumination compensation applied to the picture. In other examples, the video decoder 224 may apply other multiview decoding operations.

[0186] Furthermore, according to one or more techniques of this disclosure, the video decoder 224 may determine a multiview coding queue based on the encoded video data received from the transmitting device 102. For example, the video decoder 224 may determine a depth map, illumination compensation parameters, and other information that may be used in the multiview coding operation. The receiving device 104 may send the multiview coding queue back to the transmitting device 102 so that the transmitting device 102 can use the multiview coding queue to perform the multiview coding process for subsequent pictures.

[0187] The picture estimation unit 226 may generate an estimate of the next picture in the video data. In some examples, the picture estimation unit 226 may estimate a picture based on one or more previously reconstructed reference pictures associated with different views. For example, the picture estimation unit 226 may extrapolate the content of a picture from a picture in the same time instant using information such as a disparity vector or depth map from a picture for a previous time instant. In another example, the picture estimation unit 226 may extrapolate a picture based on one or more pictures associated with the same view, regardless of pictures associated with other views, in much the same way as described elsewhere in this disclosure with respect to the picture estimation unit 226 for estimating a picture in single-view video data.

[0188] The video encoder 228 may perform the same operations as the video encoder 210 for the next estimated picture. For example, the video encoder 228 may perform intraprediction to predict blocks and use the predicted blocks of the picture generated by the picture estimation unit 226 and the corresponding predicted blocks to generate residual data. In some examples, the video encoder 228 may perform the same multiview coding process as the video encoder 210 using a multiview coding queue. The video encoder 228 may apply a transformation (e.g., DCT transformation) to the residual data to generate transformation coefficients. The video encoder 228 may then apply quantization to the transformed coefficients.

[0189] Figure 14 is a flowchart illustrating an exemplary data exchange between a transmitting device 102 and a receiving device 104 involved in multiview processing using one or more techniques of the present disclosure. In the example of Figure 14, the transmitting device 102 may acquire a first set of multiview pictures (1400). Based on the first set of multiview pictures, the transmitting device 102 may transmit first encoded video data to the receiving device 104. In some examples, the transmitting device 102 performs light compression on the first set of multiview pictures to generate the first encoded video data. In other examples, the first encoded video data may include unencoded versions of the first set of multiview pictures.

[0190] The receiving device 104 may perform multiview processing on the first set of multiview pictures (1402). For example, the receiving device 104 may decode the first set of multiview pictures if necessary. In addition, the receiving device 104 may determine a multiview coding queue, for example, as described elsewhere in this disclosure. The receiving device 104 may transmit the multiview coding queue to the transmitting device 102.

[0191] Furthermore, in the example of Figure 14, the transmitting device 102 may acquire a second set of multiview pictures (1404). The transmitting device 102 may perform a multiview coding process on the second set of multiview pictures to generate a second coded video data (1406). The second coded video data may include coded anchor pictures and secondary (non-anchor) pictures. The receiving device 104 may perform multiview decoding on the second coded video data to reconstruct the second set of multiview pictures (1408). The receiving device 104 may also perform multiview processing on the second set of multiview pictures to determine an updated multiview coding queue (1410). The receiving device 104 may send the updated multiview coding queue to the transmitting device 102. The transmitting device 102 may use the updated multiview coding queue for multiview coding of subsequent sets of multiview pictures.

[0192] Figure 15 is a flowchart illustrating exemplary operation of a transmitting device 102 for multiview processing using the technique of the present disclosure. In the example of Figure 5, the transmitting device 102 may acquire a first set of multiview pictures of video data (1500). The first set of multiview pictures includes a first picture and a second picture. The first picture is from a first viewpoint, and the second picture is from a second viewpoint. The communication interface 118 of the transmitting device 102 (Figure 1) may transmit first encoded video data to the receiving device 104 (1502). The first encoded video data is based on the first set of multiview pictures. The transmitting device 102 may receive a multiview encoding queue from the receiving device 104 (1504). In some examples, the multiview coding queue includes one or more motion data for relative shifts between blocks of the first picture (i.e., the first viewpoint picture) and blocks of the second picture (i.e., the second viewpoint picture), brightness correction between the first and second pictures, interblock shifts between anchor blocks and reconstruction blocks, or reference shifts. The transmitting device 102 may receive the multiview coding queue in one of several ways. For example, the transmitting device 102 may receive the multiview coding queue via uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, or other types of messages.

[0193] Furthermore, the transmitting device 102 may acquire a second set of multiview pictures of the video data (1506). The second set of multiview pictures includes a third picture and a fourth picture. The third picture is from the first viewpoint, and the fourth picture is from the second viewpoint. The video encoder 210 may perform a multiview encoding process on the second set of multiview pictures to generate second encoded video data based on the multiview encoding queue received from the receiving device (1508). The multiview encoding process reduces interview redundancy between the third picture and the fourth picture. The transmitting device 102 may then transmit the second encoded video data to the receiving device 104 (1510).

[0194] The operation shown in Figure 15 may be performed multiple times for subsequent sets of multiview pictures. For example, after sending a second encoded video data to a receiving device, the transmitting device 102 may receive an updated multiview encoding queue from the receiving device 104. The transmitting device 102 may obtain a third set of multiview pictures of the video data. The third set of multiview pictures may include a fifth and a sixth picture, where the fifth picture is from the first viewpoint and the sixth picture is from the second viewpoint. The video encoder 210 of the transmitting device 102 may encode the third set of multiview pictures based on the updated multiview encoding queue received from the receiving device in order to generate the third encoded video data. The transmitting device 102 may then transmit the third encoded video data to the receiving device 104.

[0195] Figure 16 is a flowchart illustrating exemplary operation of a receiving device 104 for multiview processing using the technique of the present disclosure. In the example of Figure 16, the receiving device 104 may receive first encoded video data from the transmitting device 102 (1600). For example, the receiving device 104 may receive first encoded video data via a communication interface 134 (Figure 1). The first encoded video data is based on a first set of multiview pictures of the video data. The first set of multiview pictures may include a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint. The receiving device 104 may determine a multiview encoding queue based on the first encoded video data (1602).

[0196] The receiving device 104 may transmit a multiview coded queue to the transmitting device 102 (1604). The receiving device 104 may transmit the multiview coded queue in one of several ways. For example, the receiving device 104 may transmit a decimation pattern indication via an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, or another type of message.

[0197] In addition, the receiving device 104 may obtain second encoded video data from the transmitting device 102 (1606). The second encoded video data is based on a second set of multiview pictures, including a third picture and a fourth picture. The second encoded video data is encoded using a multiview encoding process that reduces interview redundancy between the third picture and the fourth picture based on a multiview encoding queue. The video decoder 224 of the receiving device 104 may decode the second encoded video data.

[0198] In some examples, the multiview coding queue includes a depth map indicating the depth of objects represented in the first and second pictures. The receiving device 104 may determine the depth map based on the first and second pictures as part of determining the multiview coding queue. In some examples, the multiview coding queue includes one or more illumination compensation coefficients, and the receiving device 104 may determine the illumination compensation coefficients based on the first and second pictures as part of determining the multiview coding queue.

[0199] The process in Figure 16 may be repeated multiple times. For example, the receiving device 104 may determine a second multiview coding queue based on the second encoded video data. The receiving device 104 may send the second multiview coding queue to the transmitting device 102. The receiving device 104 may then receive third encoded video data from the transmitting device 102. The third encoded video data is based on a set of third multiview pictures including a fifth picture and a sixth picture, and the third encoded video data is encoded using a multiview coding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview coding queue.

[0200] According to one or more techniques of this disclosure, transmitting device 102 may receive a decimation pattern indication from receiving device 104. The decimation pattern indication may indicate a decimation pattern. As described in more detail elsewhere in this disclosure, receiving device 104 may determine the decimation pattern. The decimation pattern may be a non-transmitted pattern of encoded video data.

[0201] The transmitting device 102 may receive decimation pattern indications in one of several ways. For example, the transmitting device 102 may receive decimation pattern indications via uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, sidelink control information (SCI), or other types of messages.

[0202] The transmitting device 102 may apply a decimation pattern to the encoded video data generated by the video encoder 210, thereby generating decimated video data. For example, the decimation pattern may indicate a pattern that skips the transmission of the encoded video data for a full picture. Thus, in this example, the transmitting device 102 (e.g., the channel encoder 212 of the transmitting device 102) may transmit the encoded video data for some pictures and not for others, according to the indicated pattern. For example, the transmitting device 102 may skip the transmission of the encoded video data for every other picture. In another example, the transmitting device 102 may transmit the encoded video data for one picture and then not transmit the encoded video data for the next two or more pictures.

[0203] In another example, a decimation pattern may represent a pattern that skips the transmission of encoded video data for a specified region within a picture. Certain regions of a series of pictures may not change much, if any, from picture to picture. For example, the background of a scene from a static viewpoint may not change significantly while changes occur in a more limited region of interest. Since regions outside the region of interest do not change much, such regions may be easier to predict accurately. Thus, according to the technique of this disclosure, the receiving device 104 may identify regions outside the region of interest. Consequently, the transmitting device 102 may transmit encoded video data for the region of interest and not transmit encoded video data for other regions.

[0204] In another example, the video data may be multi-view video data, and the decimation pattern may indicate a pattern that skips the transmission of encoded video data for pictures from a given view. For example, two views may have very similar content, such as a view that primarily shows distant objects. In this example, the transmitting device 102 may transmit encoded video data for one of the views and not transmit encoded video data for one or more of the other views, as indicated by the decimation pattern.

[0205] In another example, a decimation pattern may indicate a pattern of bits to be omitted from the syntax elements representing the conversion coefficients. For example, a decimation pattern may indicate that the least significant bit of a particular number should be omitted from the conversion coefficients. In some examples, a decimation pattern may indicate that certain conversion coefficients (e.g., high-frequency conversion coefficients) should be omitted.

[0206] Figure 17 is a block diagram illustrating exemplary components of a transmitting and receiving device that perform decimation on encoded video data using the technique of the present disclosure. In the example of Figure 17, the transmitting device 102 comprises a video encoder 210, a channel encoder 212, a puncturing unit 214, and a transmitter decimation unit 1700. The receiving device 104 comprises a depuncturing unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, a video encoder 228, and similarly a receiver decimation unit 1702. The video encoder 210, channel encoder 212, puncturing unit 214, depuncturing unit 220, channel decoder 222, video decoder 224, picture estimation unit 226, and video encoder 228 may operate in the same manner as described elsewhere in the present disclosure.

[0207] However, in the example of Figure 17, the transmitter decimation unit 1700 may apply a decimation pattern to the encoded video data after the channel encoder 212 has generated error correction data for the encoded video data. The decimation pattern indicates a pattern of non-transmission of the encoded video data. For example, the transmitter decimation unit 1700 may prevent the transmitting device 102 from transmitting encoded video data for a particular picture, a region of a picture, a pattern of blocks within a picture, or a particular view. The receiver decimation unit 1702 may determine a decimation pattern indication based on the picture reconstructed by the video decoder 224. The receiver decimation unit 1702 may transmit a decimation pattern indication showing the decimation pattern to the transmitting device 102. The transmitter decimation unit 1700 may apply the decimation pattern indicated by the decimation pattern indication.

[0208] Figure 17 illustrates a DVC-based scheme, but the techniques of this disclosure relating to sending decimation pattern indications from the receiving device 104 to the transmitting device 102 are not necessarily limited thereto. For example, in some examples, the picture estimation unit 226 and the video encoder 228 may be omitted.

[0209] Figure 18 is a conceptual diagram illustrating an exemplary exchange of information, including decimation pattern indication, using the technique of the present disclosure. In the example of Figure 18, the transmitting device 102 transmits a first set of pictures (e.g., picture n-n1, picture nn) n1+1 The receiving device 104 may transmit encoded video data for the first set of pictures (n). The transmitting device 102 may also transmit error correction data for the first set of pictures.

[0210] The receiving device 104 may transmit and the transmitting device 102 may receive a decimation pattern indication that shows a decimation pattern determined based on the first set of encoded pictures. In the example in Figure 18, the decimation pattern decimates the pictures according to a 1:2 ratio. That is, video data encoded at a ratio of 2 pictures to 1 picture will be transmitted.

[0211] Therefore, the transmitting device 102 can transmit encoded video data for the second set of pictures to the receiving device 104. According to the decimation pattern indicated by the received decimation pattern indication, the transmitting device 102 skips transmitting the encoded video data for every other picture in the second set of pictures. As shown in the example in Figure 18, the index values ​​of the pictures in the second set of pictures (e.g., n+2, n+4, n+n2) are incremented by 2 instead of 1, as in the case of the first set of pictures.

[0212] Subsequently, the receiving device 104 may determine, based on the second set of pictures, that a more appropriate decimation pattern is the 1:1 decimation pattern (i.e., the decimation pattern in which the transmitting device 102 transmits encoded video data for each picture). Thus, in the example of Figure 18, the receiving device 104 may transmit a second decimation pattern indication indicating the second decimation pattern, which the transmitting device 102 may receive. The transmitting device 102 may then transmit encoded video data for a third set of pictures. According to the second decimation pattern, the transmitting device 102 does not skip transmitting encoded video data for any of the pictures in the third set of pictures. Thus, as shown in the example of Figure 18, the index values ​​(e.g., n+n2+1, n+n2+2, etc.) increment by 1 instead of 2.

[0213] Figure 19 is a flowchart illustrating exemplary operation of the transmitting device 102 in which the transmitting device 102 receives a decimation pattern indication using the technique of the present disclosure. In the example of Figure 19, the video encoder 210 of the transmitting device 102 may encode a first set of pictures of video data to generate first encoded video data (1900). The transmitting device 102 may transmit the first encoded video data to the receiving device 104 (1902).

[0214] Furthermore, the transmitting device 102 may receive a decimation pattern indication from the receiving device 104 indicating a decimation pattern determined based on the first set of pictures (1904). The decimation pattern may be a pattern of not transmitting encoded video data. For example, in some cases, the decimation pattern indicates a pattern of skipping the transmission of encoded video data for a full picture. In other words, the transmitting device 102 may not transmit encoded video data for a particular picture, but may transmit some or all of the encoded video data for other pictures. In some cases, the decimation pattern indicates a pattern of skipping the transmission of encoded video data for a particular region within a picture. For example, the decimation pattern may indicate that the transmitting device 102 skips the transmission of encoded video data associated with a particular block of a picture, as shown, for example, in Figures 6 and 9. In some cases where the video data is multiview video data, the decimation pattern may indicate a pattern of skipping the transmission of encoded video data for a picture from a particular view. In such examples, a view may be associated with a sensor on the same part of the user device (e.g., the same XR headset), or it may be associated with a sensor on a different part of the user device (e.g., a different XR headset worn by a different user) or a camera. In some examples, different decimation patterns may exist for different areas without pictures. For example, no decimation may be applied to the region of interest, while a decimation pattern may be applied to areas of the picture outside the region of interest that restrict the transmission of blocks, lower bits, or high-frequency conversion coefficients.

[0215] The video encoder 210 may encode a second set of pictures of video data to generate second encoded video data (1906). Furthermore, the transmitter decimation unit 1700 may apply a decimation pattern to the second encoded video data to generate decimated video data (1908). The transmitting device 102 may transmit the decimated video data to the receiving device (1910).

[0216] In some examples, the transmitter decimation unit 1700 may determine a decimation pattern. Thus, in the context of Figure 19, the transmitting device 102 may encode a third set of pictures of video data to generate a third encoded video data, determine a second decimation pattern indicating a second pattern of the encoded video data that is not transmitted, and apply the second decimation pattern to the third encoded video data to generate a second decimated video data. The transmitting device 102 may transmit the second decimated video data to the receiving device 104. The transmitting device 102 may transmit a second decimation pattern indication to the receiving device. The second decimation pattern indication indicates that the second decimation pattern has been applied to the third encoded video data.

[0217] The transmitter decimation unit 1700 can determine the decimation pattern in various ways. For example, the transmitter decimation unit 1700 can test various decimation patterns. When testing decimation patterns, the transmitter decimation unit 1700 can apply the decimation pattern to a picture and reconstruct the picture from error correction data for the picture and one or more previous original pictures of the video data. The transmitter decimation unit 1700 can compare the reconstructed picture to the original picture to determine the level of distortion. The transmitter decimation unit 1700 can compare the levels of distortion associated with different decimation patterns to determine the decimation pattern.

[0218] In some examples, the operation shown in Figure 19 is performed in the context of DVC. Thus, the channel encoder 212 of the transmitting device 102 may generate first error correction data based on first encoded video data. The transmitting device 102 may transmit the first error correction data to the receiving device 104. The channel encoder 212 may generate second error correction data based on second encoded video data. The transmitting device 102 may transmit the second error correction data to the receiving device.

[0219] Figure 20 is a flowchart illustrating exemplary operation of a receiving device 104 transmitting a decimation pattern indication using the technique of the present disclosure. In the example of Figure 20, the receiving device 104 may receive first encoded video data from the transmitting device 102 (2000). The video decoder 224 may perform a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data (2002).

[0220] Furthermore, the receiver decimation unit 1702 may determine a decimation pattern that indicates a pattern of non-transmission of encoded video data based on a first set of pictures (2004). In some examples, the decimation pattern indicates a pattern that skips the transmission of encoded video data for the full picture. In some examples, the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specified region or block within a picture, as in the examples in Figures 6 and 9. In some examples, the video data is multiview video data, and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a specified view.

[0221] The receiver decimation unit 1702 may determine a decimation pattern in one of several ways. For example, the receiver decimation unit 1702 may apply one or more trial decimation patterns to the error-corrected encoded video data generated by the channel decoder 222 for a first set of pictures to generate decimated encoded video data. The receiver decimation unit 1702 may then cause the channel decoder 222 to apply an error correction process to correct the decimated encoded video data based on the first error correction data to generate trial error-corrected video data. The receiver decimation unit 1702 may then cause the video decoder 224 to apply a decoding process to reconstruct the first set of pictures based on the trial error-corrected video data. The receiver decimation unit 1702 may determine whether a decimation pattern satisfies one or more criteria based on a comparison between a first set of pictures reconstructed based on trial error-corrected video data and a first set of pictures reconstructed based on the first error-corrected video data. For ease of explanation, the disclosure may refer to the pictures reconstructed based on trial error-corrected video data as “trial pictures” and the pictures reconstructed based on the first error-corrected video data as “baseline pictures.” The receiver decimation unit 1702 may repeat this procedure with multiple trial decimation patterns until the receiver decimation unit 1702 identifies a decimation pattern that satisfies the criteria.

[0222] For example, the receiver decimation unit 1702 may compare each trial picture with a corresponding baseline picture to determine whether the trial picture meets a criterion. For example, the receiver decimation unit 1702 may determine that a trial picture meets a criterion if the sum of the differences between the trial picture and the corresponding baseline picture is less than a certain amount. If at least a given number of trial pictures exceed the threshold, the receiver decimation unit 1702 may select a decimation pattern associated with the trial pictures.

[0223] In a more general example, the receiver decimation unit 1702 may apply a function to the trial picture and the corresponding baseline picture to generate a value. If the value is less than a threshold, the receiver decimation unit 1702 may select a decimation pattern associated with the trial picture.

[0224] Furthermore, in some examples, the receiver decimation unit 1702 may cancel the use of a decimation pattern (e.g., revert to a pattern where all encoded video data is transmitted) or change to a less aggressive decimation pattern if certain conditions occur. For example, the receiver decimation unit 1702 may cancel the use of a decimation pattern in response to determining that a given number of baseline pictures do not meet the criteria. For example, the receiver decimation unit 1702 may determine that a trial picture fails the criteria if the sum of the differences between the trial picture and the corresponding baseline picture is greater than a certain amount. If at least a given amount of trial pictures fail the criteria exceeds a threshold, the receiver decimation unit 1702 may cancel the use of the decimation pattern or revert to a less aggressive decimation pattern. In a more general example, the receiver decimation unit 1702 may apply a function to the trial picture and the corresponding baseline picture to generate a value. If the value is greater than the second threshold, the receiver decimation unit 1702 may cancel the use of the decimation pattern or revert to a less aggressive decimation pattern.

[0225] The receiver decimation unit 1702 may transmit a decimation pattern indication showing the determined decimation pattern to the transmitting device 102 (2006). The receiver decimation unit 1702 may transmit the decimation pattern indication using an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, a sidelink control information (SCI) message, or another type of message.

[0226] The receiving device 104 may receive decimated video data from the transmitting device 102 (2008). The decimated video data may include second encoded video data to which a decimation pattern has been applied. The second encoded video data is generated based on a second set of pictures of the video data.

[0227] The video decoder 224 may perform a decoding process to reconstruct a second set of pictures based on the second error-corrected encoded video data (2010). The video decoder 224 may perform the same decoding process as described elsewhere in this disclosure.

[0228] In some examples, the receiving device 104 may receive and use a decimation pattern indication from the transmitting device 102. Thus, in the example of Figure 20, the receiving device 104 may receive a second decimation pattern indication showing a second pattern of untransmitted encoded video data. The receiving device 104 may receive third error correction data and second decimated video data from the transmitting device 102. The second decimated video data may include third encoded video data to which the second decimation pattern has been applied. The third encoded video data may be generated based on a third set of pictures of video data. The channel decoder 222 may apply an error correction process to generate third error-corrected encoded video data based on the third encoded video data and the third error correction data. The video decoder 224 may apply a decoding process to reconstruct the third set of pictures based on the third error-corrected encoded video data.

[0229] In some examples, the process shown in Figure 20 may be performed in a DVC-based implementation. Thus, the receiving device 104 may receive first error correction data from the transmitting device 102. The receiving device 104 applies an error correction process to correct the first encoded video data based on the first error correction data to generate first error-corrected encoded video data. The receiving device 104 may perform a decoding process to reconstruct the first set of pictures based on the first error-corrected encoded video data. Furthermore, the receiving device 104 may receive second error correction data from the transmitting device 102. The receiving device 104 may apply an error correction process to generate second error-corrected encoded video data based on the second encoded video data, the predicted encoded video data generated by the video encoder 228, and the second error correction data. The receiving device 104 may perform a decoding process to reconstruct the second set of pictures based on the second error-corrected encoded video data.

[0230] During the video encoding process, a video encoder typically analyzes multiple encoding options and selects the best one. For example, a video encoder may analyze multiple ways of dividing a maximum coding unit (LCU) or macroblock into coding units (CUs) and / or prediction units (PUs). In another example, when a video encoder performs intra-prediction to generate prediction blocks for PUs, it may analyze multiple intra-prediction modes. In yet another example, when a video encoder performs inter-prediction to generate prediction blocks for PUs, it may analyze multiple reference pictures and motion vectors. Such analysis and selection can be resource-intensive. For example, to be efficient, a video encoder may need to process multiple options in parallel, which increases the hardware complexity of the video encoder and increases power requirements. Analysis and selection may also involve multiple requests to read and write data to memory, which further increases power requirements.

[0231] According to one or more techniques of this disclosure, the majority of the process of analyzing and selecting encoding operations is shifted from a transmitting device (e.g., transmitting device 102) to a receiving device (e.g., receiving device 104). For example, the transmitting device may encode a first picture of video data in order to generate first encoded video data. The transmitting device may transmit the first encoded video data to the receiving device. The receiving device may receive the first encoded video data from the transmitting device and reconstruct the first picture based on the first encoded video data. Furthermore, the receiving device may estimate a second picture of the video data based on the first picture. The second picture may be a picture that occurs after the first picture in the decoding order. The receiving device may generate encoding selection data for the estimated second picture. The encoding selection data indicates the encoding selection used to encode the estimated second picture. The receiving device may transmit the encoding selection data for the second picture. The transmitting device may receive the encoding selection data for the second picture of the video data. The transmitting device may encode a second picture based on encoding selection data to generate second encoded video data. The transmitting device may transmit the second encoded video data to the receiving device. The receiving device may receive the second encoded video data from the transmitting device. The receiving device may reconstruct the second picture based on the second encoded video data. In this way, since the transmitting device receives encoding selection data from the receiving device, the transmitting device does not need to perform resource-intensive analysis and selection processes while encoding the second picture, because the analysis and selection processes have already been performed on the second picture at the receiving device. This can reduce the resource requirements of the transmitting device.

[0232] Transmitting and receiving devices may communicate using low-range, low-power links over ultra-wideband (e.g., broadband) communication links. In some examples, transmitting and receiving devices may communicate using time-division duplex (TDD), subband non-overlapping full-duplex (SBFD), or SFFD schemes. The low latency associated with this type of communication may allow the transmitting device to receive encoded selection data quickly enough to continue transmitting encoded video data to meet a given picture rate.

[0233] Figure 21 is a block diagram showing exemplary components of a transmitting device 102 and a receiving device 104 that transmits encoded selection data to the transmitting device, according to the technique of the present disclosure. In the example of Figure 21, the transmitting device 102 may include a video encoder 210, a channel encoder 212, and a puncturing unit 214. The receiving device 104 may include a depuncturing unit 220, a channel decoder 222, a video decoder 224, a picture estimation unit 226, and a video encoder 228.

[0234] In the example shown in Figure 21, the video encoder 210 may encode pictures of video data. Unlike some of the examples provided above, the video encoder 210 may perform a complete video encoding process, which may include intra-prediction and inter-prediction. In some examples, the video encoder 210 may encode video data using video codecs such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, Essential Video Coding (EVC), and AV1. The channel encoder 212, puncturing unit 214, depuncturing unit 220, and channel decoder 222 may operate in the same manner as described elsewhere in this disclosure.

[0235] Furthermore, in the example shown in Figure 21, the video decoder 224 of the receiving device 104 may perform a video decoding process on the error-corrected encoded video data generated by the channel decoder 222. The video decoder 224 may perform a complete video decoding process, including intra-prediction and inter-prediction. The video decoder 224 may use the same video codec as the video encoder 210.

[0236] After the video decoder 224 has reconstructed at least a portion of the picture of the video data, the picture estimation unit 226 may estimate the corresponding portion of the subsequent picture following the reconstructed picture. The picture estimation unit 226 may estimate the subsequent picture in the same manner as described elsewhere in this disclosure. Furthermore, the video encoder 228 may apply a video encoding process to the subsequent picture. In examples where the video encoder 210 and video decoder 224 use a video codec, the video encoder 228 may use the same codec.

[0237] However, according to one or more techniques of the present disclosure, the receiving device 104 may transmit coding selection data 2100 to the transmitting device 102. The coding selection data 2100 indicates the coding selection used to encode the estimated subsequent picture. For example, the coding selection data may include motion parameters for a block in the estimated subsequent picture. The block motion parameters may include motion vectors, reference picture indicators, merge candidate indices, affine motion parameters, and other data used to determine the predicted block of a block in one or more reference pictures. Thus, in this example, the video encoder 228 may encode the block using interprediction, and the coding selection data 2100 may include motion parameters indicating how the video encoder 228 encoded the block using interprediction.

[0238] In some examples, the coding selection data may include intra-prediction parameters for the estimated blocks in the subsequent picture. The intra-prediction parameters may include data indicating the intra-prediction mode (e.g., planar mode, DC mode, directional prediction mode, etc.) used by the video encoder 228 for intra-prediction of the blocks. In some examples, the coding selection data may include other information, such as information describing how the video encoder 228 partitioned the estimated subsequent picture into blocks, whether residual prediction is used, whether intra-block copying (IBC) is used, whether a particular filter is used, etc.

[0239] The video encoder 210 may use encoding selection data 2100 when encoding actual (unpredicted) subsequent pictures. That is, instead of exploring different possibilities during the video encoding process, the video encoder 210 may use the video encoding selection indicated by the encoding selection data 2100. For example, the encoding selection data 2100 may indicate that a particular block of a subsequent picture is encoded in a particular intra-prediction mode. In this example, when encoding a subsequent picture, the video encoder 210 may encode a particular block using a particular intra-prediction mode without analyzing different potential intra-prediction modes to select a particular intra-prediction mode. In another example, the encoding selection data 2100 may indicate a motion vector and a reference picture for a particular block of a subsequent picture. In this example, when encoding a subsequent picture, the video encoder 210 may use a motion vector to determine the predicted block in the reference picture without analyzing a potential reference picture and motion vector. The video encoder 210 may use a predicted block to encode a particular block.

[0240] The transmitting device 102 can handle encoded video data for subsequent pictures in the same way as other pictures. Furthermore, the receiving device 104 can handle encoded video data for subsequent pictures in the same way as other encoded video data. Thus, after the video decoder 224 reconstructs at least a portion of the subsequent picture, the picture estimation unit 226 can predict the corresponding portion of the picture following the subsequent picture, the video encoder 228 can encode the video data of the estimated subsequent picture and transmit the encoded selection data for the estimated subsequent picture, and the cycle can be repeated. In this way, part of the burden of encoding the video data can be shifted from the video encoder 210 of the transmitting device 102 to the video encoder 228 of the receiving device 104. This can reduce the resource requirements of the transmitting device 102.

[0241] The process described with respect to Figure 21 may be adapted for use with DVC technology. For example, the transmitting device 102 may apply a decimation pattern to the encoded video data (e.g., according to one of the examples provided elsewhere in this disclosure) so that the transmitting device 102 transmits only some encoded video data, but still transmits error correction data for the decimated video data. The channel decoder 222 of the receiving device 104 may receive the error correction data for a particular picture from the transmitting device (e.g., via the depuncturing unit 220). In this example, the channel decoder 222 may apply an error correction process to generate error-corrected encoded video data based on the error correction data for the particular picture and the encoded video data for the particular picture generated by the video encoder 228 of the receiving device 104. The video decoder 224 may decode the error-corrected encoded video data to reconstruct a particular feature.

[0242] In some cases, the transmitting device 102 needs to transmit encoded video data according to a schedule. For example, the transmitting device 102 may need to transmit encoded video data to the receiving device 104 at a predetermined picture-per-minute rate to support a particular application. Therefore, a situation may arise where the transmitting device 102 encodes and transmits the encoded video data for a picture, but does not receive the encoding selection data for the picture in a timely manner. Thus, in some cases, the video encoder 210 may encode the picture without using the encoding selection data for the picture, based on the determination that the encoding selection data for the picture will not be received from the receiving device 104 before the expiration of the time limit. The video encoder 210 may use a limited video encoding process to encode the picture. Furthermore, in some cases, the encoded video data for the picture may include encoding selection data generated by the video encoder 210. In some cases, the encoded video data for the picture may include data indicating that the encoded video data for the picture was not generated based on the encoding selection data generated by the receiving device 104. The time limit may be defined according to or based on the capabilities of the transmitting device 102.

[0243] Encoded selection data and time limits can be defined for each image segment (e.g., slice, region, etc.). Therefore, in this disclosure, descriptions of encoding selection data for a picture, encoded video data, or other types of data may apply only to individual segments of a picture.

[0244] Figure 22 is a communication diagram illustrating exemplary data exchange between a transmitting device 102 and a receiving device 104, including the transmission and reception of encoding selection data using the technique of the present disclosure. In the example of Figure 22, the transmitting device 102 transmits encoded video data for picture n-1 to the receiving device 104. The receiving device 104 may reconstruct picture n-1 based on the encoded video data for picture n-1. Furthermore, the receiving device 104 may estimate and encode picture n based on picture n-1. The receiving device 104 may transmit encoding selection data for picture n to the transmitting device 102. The transmitting device 102 may encode picture n based on the encoding selection data for picture n and transmit the resulting encoded video data for picture n to the receiving device 104. This process can be repeated multiple times. Therefore, in the example of Figure 22, the receiving device 104 may reconstruct picture n based on the encoded video data for picture n, estimate picture n+1 based on picture n and / or one or more other previously reconstructed pictures, encode picture n+1, and transmit the encoding selection data for picture n+1 to the transmitting device 102.

[0245] Figure 23 is a flowchart illustrating exemplary operation of the transmitting device 102, in which case the transmitting device 102 receives encoding selection data using the technique of the present disclosure. In the example of Figure 23, the video encoder 210 encodes a first picture of video data (2300) to generate first encoded video data. In some examples, if the transmitting device 102 has not received encoding selection data for the first picture, the transmitting device 102 may perform a limited video encoding process for the first picture. The limited video encoding process may use coding tools that are relatively less computationally intensive than the full video encoding process. For example, the limited video encoding process may use intra-prediction but not inter-prediction.

[0246] The transmitting device 102 may transmit the first encoded video data to the receiving device 104 (2302). In some examples, the transmitting device 102 may apply a channeling encoding process to the first encoded video data to generate error correction data for the first encoded video data. The transmitting device 102 may transmit the first encoded video data and the error correction data to the receiving device 104.

[0247] Subsequently, the transmitting device 102 may receive coding selection data for a second picture of the video data from the receiving device 104 (2304). The coding selection data may indicate the coding selection used to code an estimate of the second picture. The second picture follows the first picture in the decoding order. In some examples, the second picture may occur before or after the first picture in the output codeder. In some examples, the coding selection data is entropy coded. Thus, in such examples, the transmitting device 102 may entropy decode the coding selection data. For example, the transmitting device 102 may apply CABAC decoding, Golomb-Rice decoding, or another type of entropy decoding to the coding selection data. In some examples, the coding selection data is channel coded. Thus, the transmitting device 102 may apply error correction operations to the coding selection data based on error correction data for the coding selection data.

[0248] The video encoder 210 of the transmitting device 102 may encode a second picture based on encoding selection data to generate second encoded video data (2306). For example, the encoding selection data may include data indicating how to partition a particular macroblock into CUs. In this example, the video encoder 210 may partition the macroblock into CUs in the manner indicated by the encoding selection data. In another example, the encoding selection data may indicate an intra-prediction mode for a block (e.g., CU or PU), and the video encoder 210 may use the indicated intra-prediction mode to encode the block. Thus, in this example, the encoding selection data received from the receiving device 104 may include intra-prediction parameters for blocks of the second picture, and the video encoder 210 may perform intra-prediction based on the intra-prediction parameters for blocks of the second picture to generate predicted blocks as part of encoding the second picture. The second encoded video data may include encoded video data based on predicted blocks.

[0249] In another example, the encoding selection data received from the receiving device 104 includes motion parameters for a block of the second picture, and the transmitting device 102 may perform motion compensation based on the motion parameters for the block of the second picture to generate a predicted block as part of encoding the second picture. The second encoded video data includes encoded video data based on the predicted block.

[0250] The transmitting device 102 may transmit the second encoded video data to the receiving device (2308). In some examples, the second encoded video data does not include encoding selection data indicating the encoding selection used when the transmitting device 102 encodes the second picture, or the encoding selection used when the receiving device 104 encodes an estimated value of the second picture. Since the receiving device 104 generates the encoding selection data and thus already has the encoding selection data, it may not be necessary for the second encoded video data to include the encoding selection data.

[0251] In some examples, the operations of FIG. 23 may be used in DVC techniques. For example, the transmitting device 102 may receive encoding selection data for a third picture of video data. The encoding selection data for the third picture may indicate the encoding selection used to encode an estimated value of the third picture. The video encoder 210 may encode the third picture based on the encoding selection data for the third picture to generate the third encoded video data. The channel encoder 212 may apply a channel encoding process to generate error correction data for the third encoded video data. The transmitting device 102 may transmit the error correction data for the third encoded video data to the receiving device without transmitting at least a portion of the third encoded video data.

[0252] FIG. 24 is a flowchart illustrating an exemplary operation of the receiving device 104 in which the receiving device 104 transmits encoding selection data according to the techniques of the present disclosure. In the example of FIG. 24, the receiving device 104 may receive the first encoded video data from the transmitting device (2400).

[0253] The video decoder 224 of the receiving device 104 may reconstruct a first picture of the video data based on the first encoded video data (2402). The picture estimation unit 226 may estimate a second picture of the video data by the receiving device 104 based on the first picture (2404). The second picture may be a picture that occurs after the first picture in the decoding order.

[0254] The video encoder 228 of the receiving device 104 may generate encoding selection data for the estimated second picture (2406). The encoding selection data indicates an encoding selection used to encode the estimated second picture. For example, as part of encoding the estimated second picture, the video encoder 228 may perform motion compensation based on motion parameters for blocks of the second picture to generate prediction blocks. In this example, the encoding selection data may include motion parameters for one or more processors of blocks of the second picture. In some examples, as part of encoding the second picture, the video encoder 228 may perform intra prediction based on intra prediction parameters for blocks of the second picture to generate prediction blocks. In this example, the encoding selection data may include intra prediction parameters for blocks of the second picture.

[0255] The receiving device 104 may transmit coded selection data for a second picture to the transmitting device 102 (2408). In some examples, the receiving device 104 may apply entropy coding (e.g., CABAC coding, Golomm-Rice coding, etc.) to the coded selection data for the second picture before transmitting it. In some examples, the receiving device 104 may perform a channel coding process on the coded selection data to generate error correction data for the coded selection data. The receiving device 104 may transmit the coded selection data and the error correction data for the coded selection data to the transmitting device 102. In some examples, the communication interface 134 of the receiving device 104 (Figure 1) may modulate the coded selection data at a lower modulation order compared to other data transmissions in the data link between the receiving device 104 and the transmitting device 102 (e.g., wireless sidelink channel 112, wireless uplink / downlink channel, etc.). This may increase the likelihood that the transmitting device 102 will correctly receive the coded selection data.

[0256] Subsequently, the receiving device 104 may receive second encoded video data from the transmitting device (2410). The video decoder 224 may reconstruct a second picture based on the second encoded video data (2412). In some examples, the second encoded video data does not include encoding selection data. The video decoder 224 may apply a decoding process that includes using encoding selection data to reconstruct a second picture based on the second encoded video data.

[0257] The process in Figure 24 can be used in conjunction with DVC technology. For example, a picture estimation unit 226 may estimate a third picture of video data based on one or more of the first or second pictures. A video encoder 228 may encode the estimated third picture to generate third encoded video data. A receiving device 104 may transmit third encoding selection data to a transmitting device 102. The third encoding selection data may indicate the encoding selection used to encode the estimated third picture. The receiving device 104 may then receive error correction data for the third picture. A channel decoder 222 may apply an error correction process to generate error-corrected encoded video data for the third picture based on the error correction data for the third picture and the third encoded video data. A video decoder 224 may apply a decoding process to reconstruct the third picture based on the error-corrected encoded video data for the third picture. In some examples where the transmitting device 102 and receiving device 104 use the DVC technique, the error-corrected video data for the third picture does not include the third encoding selection data. However, the video decoder 224 may apply a decoding process using the third encoding selection data generated by the video encoder 228 of the receiving device 104 to reconstruct the third picture based on the error-corrected encoded video data for the third picture.

[0258] Figure 25 is a conceptual diagram illustrating an exemplary hierarchy of encoded video data using the techniques of this disclosure. More specifically, Figure 25 shows the hierarchy of encoded video data generated using the H.264 / AVC video coding standard. As illustrated in the example in Figure 25, the Network Abstraction Layer (NAL) is the highest level of the hierarchy. At the Network Abstraction Layer, data is organized into NAL units. In some examples, NAL units are assigned to different packets or coding blocks for transmission. Network Abstraction Layer NAL units may include Sequence Parameter Sets (SPSs) and Picture Parameter Sets (PPSs) containing high-level syntax. Network Abstraction Layer NAL units may also include Video Coding Layer (VCL) NAL units. VCL NAL units may include slice NAL units containing slice-level data. A slice can be a series of macroblocks within a picture. Slices may include Instantaneous Decoder Refresh (IDR) slices and regular slices. Decoding of an IDR slice is independent of other slices. Regular slices may have dependencies on other slices.

[0259] Each slice NAL unit may contain a slice header and slice data. The slice header of a slice NAL unit contains information for decoding the slice data of the slice NAL unit. The slice data of a slice NAL unit contains a set of macroblocks (MBs). Skip instructions may be scattered between MBs. Each MB contains encoded video data for a particular block of the slice. Furthermore, as shown in Figure 25, an MB may contain a type indicator, prediction information, encoded block pattern, quantization parameters (QP), and encoded residual data. If an MB is encoded using intra-prediction, the prediction data may indicate one or more intra-modes used to encode the MB. If an MB is encoded using inter-prediction, the prediction data may indicate one or more reference pictures and one or more motion vectors. The encoded residual data for an MB may include encoded residual data for the rumor block in the MB, encoded residual data for the Cb block in the MB, and encoded residual data for the Cr block in the MB. In general, the encoded residual data is the largest portion of the encoded video data.

[0260] According to the techniques of this disclosure, everything in the hierarchy of the macroblock layer, excluding the encoded residual data, can be encoded selection data. Thus, in some examples, the receiving device 104 may transmit type data, prediction data, encoded block pattern, and QP to the transmitting device 102 for each MB of the estimated picture. Furthermore, in some examples, the transmitting device 102 may transmit only the encoded residual data of the MB to the receiving device 104, and not the type data, prediction data, encoded block pattern, or QP for the MB. In some examples, the encoded selection data transmitted by the receiving device 104 may include slice header data, SPS data, and PPS data. The transmitting device 102 and the receiving device 104 can exchange information or may be pre-configured with information indicating encoder and decoder capabilities.

[0261] In some examples, the transmitting device 102 may transmit data in addition to encoded residual data for several pictures, several MBs, or several slices. The transmitting device 102 may signal information (e.g., bits) in the network abstraction layer (e.g., as a picture-level control field) that may indicate whether a DVC-based technique is used or whether normal compression should be used for a particular picture.

[0262] Figure 26 is a block diagram showing an exemplary alternative component of the transmitting device 102 using one or more techniques of the present disclosure. In the example of Figure 26, the transmitting device 102 performs digital and analog coding on the video data. The transmitting device 102 transmits the digitally coded and analogously coded video data to the receiving device 104 via channel 230.

[0263] The example in Figure 26 includes a video encoder 2600, a residual generation unit 2602, an analog encoder 2604, a reliability sorting unit 2606, an interleaving unit 2608, a channel encoder 2610, and a puncturing unit 2612. The video encoder 2600 acquires video data and may operate in much the same manner as the video encoder 210 in Figure 2. As shown in the example in Figure 26A, the video encoder 2600 may receive values ​​of encoding parameters (e.g., encoding selection parameters) sent by the receiving device 104. In some examples, the video encoder 2600 may send values ​​of encoding parameters and / or encoding selection data to the receiving device 104. In this way, the video encoder 2600 of the transmitting device 102, the video encoder of the receiving device 104, and the video decoder of the receiving device 104 may operate based on the same values ​​of encoding parameters.

[0264] The video encoder 2600 can also output prediction data to the residual generation unit 2602. Furthermore, the video encoder 2600 can apply a higher level of quantization than the video encoder 210. The residual generation unit 2602 can generate residual data based on the prediction data and video data.

[0265] The analog encoder 2604 can perform analog coding operations on residual data. Illustrative details of analog coding operations can be found in U.S. Patent No. 11,553,184, “Hybrid Digital-Analog Modulation for Transmission of Video Data,” filed December 29, 2020; U.S. Patent No. 11,431,962, “Analog Modulated Video Transmission with Variable Symbol Rate,” filed December 29, 2020; and U.S. Patent No. 11,457,224, “Interlaced Coefficients in Hybrid Digital-Analog Modulation for Transmission of Video Data,” filed December 29, 2020.

[0266] Furthermore, in this example, the analog encoder 2604 may generate coefficients based on residual data. For example, the analog encoder 2604 may generate coefficients by binarizing the residual data. The analog encoder 2604 may quantize the coefficients. In other examples of generating coefficients based on video data, the analog encoder 2604 may perform more, fewer, or different steps. For example, in some examples, the analog encoder 2604 does not perform the quantization step. In yet another example, the analog encoder 2604 does not perform the step of binarizing the residual data.

[0267] Furthermore, the analog encoder 2604 can generate coefficient vectors. Each coefficient vector contains n of the coefficients. The analog encoder 2604 can generate coefficient vectors in one of several ways. For example, in one example, the analog encoder 2604 can generate coefficient vectors as groups of n consecutive coefficients according to a coefficient coding order. Various coefficient coding orders can be used, such as raster scan order, zigzag scan order, inverse raster scan order, and vertical scan order. In some examples, the coefficient vectors may contain one or more negative coefficients and one or more positive coefficients (i.e., signed coefficients). In some examples, the coefficient vectors may contain only non-negative coefficients (i.e., unsigned coefficients).

[0268] For each of the coefficient vectors, the analog encoder 2604 can determine the amplitude value of the coefficient vector based on the mapping pattern. For each of the multiple tolerance coefficient vectors, the mapping pattern maps each tolerance coefficient vector to each of the multiple amplitude values. Each amplitude value is adjacent in n-dimensional space to at least one other amplitude value among the multiple amplitude values ​​adjacent to it on the monotonic number line of amplitude values.

[0269] In some examples, the analog encoder 2604 may determine a position in n-dimensional space in order to determine the amplitude value of a coefficient vector. The coordinates of the position in n-dimensional space are based on the coefficients of the coefficient vector, and the mapping pattern maps different positions in n-dimensional space to different amplitude values ​​among multiple amplitude values. The analog encoder 2604 may determine the amplitude value of the coefficient vector as the amplitude value corresponding to the determined position in n-dimensional space.

[0270] The analog encoder 2604 can modulate an analog signal based on amplitude values ​​relating to a coefficient vector. For example, the analog encoder 2604 can determine an analog symbol based on a pair of amplitude values. The analog symbol may correspond to the phase shift and power of a point in the IQ plane having coordinates indicated by the amplitude value pair. Based on the determined phase shift and power, the analog encoder 2604 can modulate the analog signal during the symbol sampling time. The modem of the transmitting device 102 (e.g., the communication interface 118) may be configured to output an analog signal.

[0271] Furthermore, in the example in Figure 26A, the reliability sorting unit 2606 may acquire encoded video data generated by the video encoder 2600. The reliability sorting unit 2606 may acquire reliability-side information from the video encoder 2600. In some examples, the reliability sorting unit 2606 may receive values ​​for channel and compression state feedback (CCSF) parameters. The values ​​of the CCSF parameters may provide information about the state of channel 230 (e.g., signal-to-noise ratio, latency, network bandwidth congestion, etc.). In some examples, the values ​​of the CCSF parameters may provide information related to the reliability and quality of the prediction. For example, the values ​​of the CCSF parameters may include a decimation pattern indicator. In some examples, the values ​​of the CCSF parameters may enable the channel encoder 2610 to determine the decimation pattern.

[0272] The interleaving unit 2608 can perform interleaving operations that ensure trusted bits and untrusted bits are equally distributed across code blocks. For example, encoded video data may be divided into code blocks. The channel encoder 2610 may generate a separate set of error correction data for each code block. Before the channel encoder 2610 generates the error correction data, the interleaving unit 2608 may interleave the encoded video data across code blocks according to a predefined interleaving pattern. For example, encoded video data representing different adjacent pixels may be interleaved into different code blocks. The deinterleaving operation performed by the receiving device 104 reverses the interleaving operation after the channel decoding process is applied. Therefore, if one of the code blocks is corrupted during transmission, pixels decoded from the corrupted code block may be spatially distributed within the picture among pixels decoded from the uncorrupted code block.

[0273] The channel encoder 2610 of the transmitting device 102 may perform a channel coding process on the video data acquired from the interleaving unit 2608. The channel encoder 2610 may perform a channel coding process according to any of the examples provided with respect to the channel encoder 212 (Figure 2). The puncturing unit 2612 may perform a bit puncturing operation on the error correction data generated by the channel encoder 2610. The puncturing unit 2612 may perform a bit puncturing operation on the error correction data according to any of the examples provided with respect to the puncturing unit 214 (Figure 2). The transmitting device 102 may transmit the coded video data and error correction data (e.g., bit-punctured error correction data) to the receiving device 104 via channel 230.

[0274] FIG. 27 is a block diagram showing exemplary alternative components of the receiving device 104 according to one or more techniques of the present disclosure. The version of the receiving device 104 shown in FIG. 27 may be compatible with the version of the transmitting device 102 shown in FIG. 26. In the example of FIG. 27, the receiving device 104 includes an analog decoder 2700, a de-puncturing unit 2702, a channel decoder 2704, an interleaving unit 2706, a video decoder 2708, a reconstruction unit 2710, a picture estimation unit 2712, a video encoder 2714, a reliability unit 2716, and a feedback unit 2718.

[0275] The analog decoder 2700 may obtain the encoded video data. The analog decoder 2700 may perform an analog decoding operation to reconstruct the residual data. Detailed examples of the analog decoding operation can be found in U.S. Patent No. 11,553,184, U.S. Patent No. 11,431,962, and U.S. Patent No. 11,457,224.

[0276] For example, in some examples, the analog decoder 2700 may determine the amplitude values of a plurality of coefficient vectors based on an analog signal. For example, the analog decoder 2700 may determine the phase shift and power at the symbol sampling time of the analog signal. The analog decoder 2700 may determine a point in the I-Q plane indicated by the determined phase shift and power. The analog decoder 2700 may then determine an amplitude value pair as the coordinates of the point in the I-Q plane.

[0277] For each of the coefficient vectors, the analog decoder 2700 may determine the coefficients in the coefficient vector based on the amplitude values ​​and mapping pattern related to the coefficient vector. For each of the multiple allowable coefficient vectors, the mapping pattern may map each allowable coefficient vector to each of the amplitude values ​​of the multiple amplitude values. Each amplitude value is adjacent in n-dimensional space to at least one other amplitude value among the multiple amplitude values ​​adjacent to it on the monotonic number line of the amplitude values. Each of the coefficient vectors may contain n of the coefficients. The value n can be 2 or greater. In some examples, the analog decoder 2700 may determine the coefficients in the coefficient vector as coordinates of the positions in n-dimensional space corresponding to the amplitude values. The mapping pattern maps different positions in n-dimensional space to different amplitude values ​​among the multiple amplitude values. In some examples, the coefficient vector may contain one or more negative coefficients and one or more positive coefficients. In other examples, the coefficient vector may contain only non-negative coefficients.

[0278] In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain a sign value, which indicates the positive / negative sign of the coefficient in the coefficient vector. In such examples, the analog decoder 2700 may determine the absolute value of the coefficient in the coefficient vector based on the amplitude value and mapping pattern with respect to the coefficient vector. The analog decoder 2700 may reconstruct the coefficient in the coefficient vector by applying the sign value to the absolute value of the coefficient in the coefficient vector, at least partially. In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain data representing a shift value. In such examples, the shift value indicates the smallest negative coefficient among the coefficients in the coefficient vector. In addition, in such examples, the analog decoder 2700 may determine the midpoint of the coefficient in the coefficient vector based on the amplitude value and mapping pattern with respect to the coefficient vector. The analog decoder 2700 may reconstruct the coefficient in the coefficient vector by adding the shift value to each of the midpoints of the coefficient in the coefficient vector, at least partially.

[0279] Furthermore, the analog decoder 2700 can generate residual data based on the coefficients in the coefficient vector. For example, in one example, the analog decoder 2700 can inversely quantize the coefficients of the coefficient vector. In this example, the analog decoder 2700 can perform a de-binarization process to convert the coefficients into digital sample values. For example, the analog decoder 2700 can apply an inverse DCT to the coefficients to convert them into digital sample values. In this way, the analog decoder 2700 can generate digital residual sample values.

[0280] The depunching unit 2702 may obtain encoded video data and bit-punctured error-corrected data. The depunching unit 2702 may apply a depunching process to the bit-punctured error-corrected data to reconstruct the error-corrected data. The depunching unit 2702 may apply the depunching process according to any of the examples provided elsewhere in this disclosure with respect to the depunching unit 220 in Figure 2.

[0281] The channel decoder 2704 may perform a channel decoding process to correct the encoded video data (e.g., encoded video data received via channel 230, or encoded video data generated by the video encoder 2714 and, in some examples, corrected by the reliability unit 714) based on error correction data. The channel decoder 2704 may perform a channel decoding process according to any of the examples provided elsewhere in this disclosure with respect to the channel decoder 222 in Figure 2.

[0282] The deinterleaving unit 2706 may perform a deinterleaving operation on the error-corrected encoded video data generated by the channel decoder 2704. For example, the deinterleaving process may be the reverse of the interleaving process performed by the interleaving unit 2608 of the transmitting device 102. For example, the deinterleaving process may perform the deinterleaving process according to the interleaving pattern used by the interleaving unit 2608.

[0283] The video decoder 2708 may acquire encoded video data (e.g., deinterleaved encoded video data generated by the deinterleaving unit 2706). The video decoder 2708 may perform a video decoding process on the encoded video data to reconstruct the picture of the video data. The video decoding process performed by the video decoder 2708 may be the same as that described in any of the examples provided elsewhere in this disclosure with respect to the video decoder 224. The reconstruction unit 2710 of the receiving device 104 may add the residual data generated by the analog decoder 2700 to the corresponding samples of the reconstructed video data generated by the video decoder 2708, thereby completely reconstructing the picture of the video data.

[0284] Furthermore, in the example of Figure 27, the picture estimation unit 2712 may estimate one or more pictures based on a previously reconstructed picture. As previously mentioned in this disclosure, the picture description may apply to a segment of the picture, such as a slice. The picture estimation unit 2712 may estimate a picture according to any of the examples provided elsewhere in this disclosure with respect to the picture estimation unit 226. The video encoder 2714 may perform a video encoding process on the estimated picture. As part of performing the video encoding process, the video encoder 2714 may determine encoding selection data, such as encoding selection data 2100, as previously described. The receiving device 104 may transmit the encoding selection data to the transmitting device 102. In some examples, the video encoder 2714 may transmit encoding parameters, such as those described with respect to Figures 4 and 5, to the transmitting device 102, and the video encoder 2600 of the transmitting device 102 may perform a limited encoding process in the same manner as the video encoder 2714 of the receiving device 104. In some examples, the video encoder 2714 may determine a multiview coding queue and send it to the transmitting device 102.

[0285] Reliability unit 2716 may operate in much the same manner as reliability unit 1002 (Figure 10). Based on the output of reliability unit 1002, feedback unit 2718 may send predictive quality feedback (e.g., CCSF parameters) to transmission device 102.

[0286] In some examples of this disclosure, a video encoder 2600 of a transmitting device 102 may generate prediction data for a set of pictures, and a residual generation unit 2602 may generate residual data based on the first prediction data and the first set of pictures. The video encoder 2600 may apply a transformation of the prediction data to generate a transformation block, quantize the transformation coefficients of the transformation block, and apply entropy coding to the syntax elements representing the quantized transformation coefficients to generate entropy-coded syntax elements. The first encoded video data may include the entropy-coded syntax elements. A channel encoder 2610 may perform a channel coding process to generate error correction data for the encoded video data, which includes the entropy-coded syntax elements. An analog encoder 2604 may perform analog modulation on the residual data to generate analog-modulated residual data. A communication interface of the transmitting device 102 may transmit the analog-modulated residual data, the error correction data, and the encoded video data.

[0287] The communication interface of the receiving device 104 can receive analog modulation residual data and error correction data and decimated video data from the transmitting device. The decimated video data may include encoded video data to which a decimation pattern has been applied. The encoded video data is generated based on a set of pictures of video data. The channel decoder 2704 may apply an error correction process to generate error-corrected encoded video data based on the encoded video data and error correction data. The error-corrected encoded video data includes entropy-coded syntax elements representing quantized transformation coefficients. The video decoder 2708 may perform a decoding process to reconstruct a second set of pictures based on the error-corrected encoded video data. As part of applying the decoding process to reconstruct the set of pictures, the video decoder 2708 may apply entropy decoding to the syntax elements to obtain quantized transformation coefficients, dequantize the quantized transformation coefficients to generate inversely quantized transformation coefficients, and apply an inverse transform to the inversely quantized transformation coefficients to generate prediction data. The analog decoder 2700 can demodulate the analog modulation residual data to obtain residual data. The reconstruction unit 2710 can reconstruct the set of pictures based on the prediction data and the residual data.

[0288] The following is a non-exclusive list of the techniques used in this disclosure to define the provisions of one or more of these provisions.

[0289] Clause 1A. A method for decoding video data, comprising: in a receiving device, obtaining error correction data from a transmitting device, wherein the error correction data provides error correction information and is generated based on encoded video data of one or more blocks of picture in the video data; in a receiving device, generating prediction data for a picture, wherein the prediction data for a picture includes predictions of blocks of picture based at least partially on blocks of picture of one or more previously reconstructed video data, using one or more coding tools not used to generate encoded video data of one or more blocks; in a receiving device, generating encoded video data based on the prediction data for a picture; in a receiving device, generating error-corrected encoded video data using the error correction data in order to perform error correction processing on the encoded video data; and in a receiving device, performing a reconstruction operation to reconstruct blocks of image based on the error-corrected encoded video data, wherein the reconstruction operation is controlled by the values ​​of one or more parameters.

[0290] Clause 2A. The method of Clause 1A, further comprising receiving a parameter value from a transmitting device in a receiving device.

[0291] Clause 3A. The method according to Clause 1A, further comprising determining the value of a parameter in a receiving device without receiving the value of the parameter from a transmitting device.

[0292] Clause 4A. The method according to any one of Clauses 1A to 3A, wherein the parameter includes one or more quantization parameters, generating encoded video data includes using the quantization parameters to quantize transformation coefficients generated based on prediction data for a picture, and performing a reconstruction operation includes using the quantization parameters to dequantize the transformation coefficients of the error-corrected encoded video data.

[0293] Clause 5A. The method according to Clause 4A, further comprising calculating a quantization parameter based on the entropy ratio of the quantization conversion coefficients and the non-quantization conversion coefficients.

[0294] Clause 6A. The method according to any one of Clauses 1A to 5A, wherein the parameter includes a transformation size parameter, generating encoded video data includes applying a forward transformation to sample region data for a picture having a transformation size indicated by the transformation size parameter, and performing a reconstruction operation includes applying an inverse transformation to the transformation coefficients of the error-corrected encoded video data having a transformation size indicated by the transformation size parameter.

[0295] Clause 7A. The method according to any one of Clauses 1A to 6A, wherein the parameter includes a parameter indicating an amount of transformation coefficients, generating encoded video data includes including a set of transformation coefficients in the encoded video data that includes the indicated amount of transformation coefficients, and performing a reconstruction operation includes parsing a set of transformation coefficients that includes the indicated amount of transformation coefficients from the error-corrected encoded video data.

[0296] Clause 8A. The method according to Clause 7A, wherein the method for obtaining encoded video data and error correction data comprises a receiving device receiving encoded video data and error correction data from a transmitting device via a communication channel, and the method further comprises applying an optimization process that determines the number of conversion coefficients based on the signal-to-noise ratio of the data transmitted over the communication channel.

[0297] Clause 9A. The method according to any one of Clauses 1A to 8A, wherein the parameter includes a bit width parameter for a plurality of index values, and for each index value of the plurality of index values, the reconstruction operation includes parsing a first set of bits from error-corrected encoded video data, wherein the first set of bits represents a conversion coefficient having each index value, and the number of bits in the first set of bits is equal to the bit width indicated by the bit width parameter for each index value, and the encoded video data includes a second set of bits, wherein the second set of bits represents a conversion coefficient having each index value, and the number of bits in the second set of bits is equal to the bit width indicated by the bit width parameter for each index value, and the method according to any one of Clauses 1A to 8A.

[0298] Clause 10A. The method according to any one of Clauses 1A to 9A, further comprising performing a bit depuncturing operation on error-corrected data before generating error-corrected encoded video data.

[0299] Clause 11A. The method according to any one of Clauses 1A to 10A, wherein the parameter includes one or more of the following: color space, conversion size, quantization parameter, number of conversion coefficients in the first encoded video data, or number of bits per conversion coefficient in the first encoded video data.

[0300] Clause 12A. The method according to any one of Clauses 1A to 11A, wherein the decimation pattern defines a pattern of anchored and non-anchored converted blocks in a picture, and the method further comprises, at a receiving device, receiving systematic bits of anchored converted blocks rather than systematic bits of non-anchored converted blocks, the systematic bits of anchored converted blocks representing conversion coefficients in the anchored converted blocks, the systematic bits of non-anchored converted blocks representing bit-depth reduced versions of the original conversion coefficients in the non-anchored converted blocks, and the error correction data comprises error correction data for anchored converted blocks and error correction data for non-anchored converted blocks, and the method comprises using error correction data for anchored converted blocks to perform error correction on the systematic bits of anchored converted blocks and using error correction data for non-anchored converted blocks to perform error correction on the portion of encoded video data corresponding to the non-anchored converted blocks.

[0301] Clause 13A. The method according to Clause 12A, further comprising determining a decimation pattern in a receiving device and sending the decimation pattern to a transmitting device in a receiving device.

[0302] Clause 14A. The method according to any one of Clauses 1A to 13A, wherein the decimation pattern defines a pattern of anchored and non-anchored transformed blocks in a picture, the transformed coefficients in the non-anchored transformed blocks have a reduced bit depth with respect to the anchored transformed blocks, and the receiving device receives the systematic bits of the anchored transformed blocks, the systematic bits of the non-anchored transformed blocks, the systematic bits of the non-anchored transformed blocks, the systematic bits of the non-anchored transformed blocks, the systematic bits of the non-anchored transformed blocks, the systematic bits of the non-anchored transformed blocks, the systematic bits of the non-anchored transformed blocks, and a correlation matrix, and the reconstruction operation comprises, for each non-anchored transformed coefficient in the non-anchored transformed blocks, the receiving device calculating an interpolated value of the non-anchored transformed coefficient based on the correlation matrix and the corresponding anchored transformed coefficient, and the receiving device calculating a reconstructed value of the non-anchored transformed coefficient based on the interpolated value of the non-anchored transformed coefficient and the value of the non-anchored transformed coefficient in the error-corrected encoded video data.

[0303] Clause 15A. A method for encoding video data, comprising: a transmitting device acquiring video data from a video source; a transmitting device generating encoded video data of a first picture of the video data and encoded video data of a second picture of the video data based on a set of parameters; a transmitting device performing channel coding on the encoded video data of the first picture and the encoded video data of the second picture in order to generate error correction data for the first picture and error correction data for the second picture; and a transmitting device transmitting the encoded video data of the first picture, the error correction data for the first picture, and the error correction data for the second picture.

[0304] Clause 16A. The method of Clause 15A, further comprising transmitting a parameter value to a receiving device in a transmitting device.

[0305] Clause 17A. The method according to Clause 16A, wherein the parameter includes one or more quantization parameters, and generating encoded video data involves using the quantization parameters to quantize the transformation coefficients of a first picture and a transformation coefficient of a second picture.

[0306] Clause 18A. The method according to Clause 16A or 17A, wherein the parameters include a transform size parameter, and generating encoded video data involves applying a forward transform to a block of residual data of a first picture and a block of residual data of a second picture, the forward transform having a transform size indicated by the transform size parameter.

[0307] Clause 19A. The method according to any one of Clauses 16A to 18A, wherein the parameter includes a parameter indicating the amount of a conversion coefficient, and generating encoded video data of a first picture and encoded video data of a second picture includes including a set of conversion coefficients in the encoded video data of the first picture and encoded video data of the second picture, the amount of the conversion coefficient indicated.

[0308] Clause 20A. The method described in any one of Clauses 16A to 10A, wherein the parameter includes one or more of the following: color space, conversion size, quantization parameter, number of conversion coefficients in the encoded video data, or number of bits per conversion coefficient in the encoded video data.

[0309] Clause 21A. A method for encoding video data, comprising: a transmitting device acquiring video data from a video source; a transmitting device generating transformation blocks based on the video data; a transmitting device determining which of the transformation blocks are anchor transformation blocks; a transmitting device calculating a correlation matrix for the set of transformation blocks; a transmitting device generating a bit-reduced non-anchor transformation matrix; and a transmitting device transmitting the anchor transformation blocks, the non-anchor transformation blocks, and the correlation matrix to a receiving device.

[0310] Clause 22A. The method of Clause 21A, further comprising the transmitting device receiving an indication of a decimation pattern from a receiving device.

[0311] Clause 23A. A device comprising: a memory configured to store video data; a communication interface; and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors are configured to perform the method described in any one of Clauses 1A to 22A.

[0312] Clause 24A. A device comprising means for performing the method described in any one of Clauses 1A to 22A.

[0313] Clause 25A. A computer-readable data storage medium storing instructions, wherein, when the instructions are executed, causes a device to perform the method described in any one of Clauses 1A to 22A.

[0314] Clause 1B. A device for processing video data, comprising: a memory configured to store video data; a communication interface configured to obtain error correction data from a transmitting device, wherein the error correction data provides error correction information relating to pictures of video data; and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors are configured to generate prediction data for pictures, wherein the prediction data for pictures includes predictions of blocks of pictures based at least in part on one or more previously reconstructed pictures of video data; to generate encoded video data based on the prediction data for pictures, wherein the encoded video data includes a transformation block containing transformation coefficients; to scale the bits of the transformation coefficients of the transformation block based on confidence values ​​for bit positions; to generate error-corrected encoded video data using the error correction data to perform error correction operations on the scaled bits of the transformation coefficients of the transformation block; and to reconstruct a picture based on the error-corrected encoded video data.

[0315] Clause 2B. The device described in Clause 1B, further configured with one or more processors to generate a reliability value in the receiving device.

[0316] Clause 3B. The device described in Clause 2B, wherein one or more processors are configured to generate a confidence value based on statistics regarding the occurrence of errors at bit locations.

[0317] Clause 4B. The device described in Clause 2B or 3B, wherein one or more processors are configured to generate reliability values ​​based on reliability characteristics for individual regions of a picture of video data.

[0318] Clause 5B. The device described in any one of Clauses 2B to 4B, wherein one or more processors are further configured to generate reliability values ​​based on a noise model.

[0319] Clause 6B. A device described in any one of Clauses 1B to 5B, wherein the communication interface is further configured to send reliability values ​​to the transmitting device.

[0320] Clause 7B. A device described in any one of Clauses 1B to 5B, wherein the communication interface is further configured to receive a confidence value from a transmitting device.

[0321] Clause 8B. A device for processing video data, comprising: a memory configured to store video data; one or more processes implemented in the circuit and coupled to the memory; one or more processors configured to acquire video data, acquire predictive quality feedback, wherein the predictive quality feedback is based on the reliability of estimated pictures produced by a receiving device; apply one or more video coding parameters or channel coding parameters based on the predictive quality feedback; execute a video coding process, wherein the video coding process is controlled by video coding parameters, to produce channel-coded data based on one or more pictures of the acquired video data; and one or more processors configured to execute a channel coding process, wherein the channel coding process is controlled by channel coding parameters, on the encoded video data to produce channel-coded data; and a communication interface configured to transmit channel-coded data to a receiving device.

[0322] Clause 9B. The device described in Clause 8B, wherein the video coding parameters include quantization parameters, and one or more processors are configured to adapt the quantization parameters as part of adapting the video coding parameters, and one or more processors are configured to use the quantization parameters to quantize the conversion coefficients of one or more picture conversion blocks as part of performing the video coding process.

[0323] Clause 10B. The device described in Clause 8B or 9B, wherein the channel coding parameters include a low-density parity check (LDPC) graph, one or more processors are configured to adapt the LDPC graph as part of adapting the channel coding parameters, and one or more processors are configured to use the LDPC graph to generate codewords to be included in the channel-coded data as part of performing the channel coding process.

[0324] Clause 11B. A device according to any one of Clauses 8B to 10B, wherein channel-encoded data includes error-corrected data, one or more processors are further configured to adapt one or more bit-puncturing parameters based on predictive quality feedback, one or more processors are configured to perform a bit-puncturing process on the error-corrected data, and the bit-puncturing process is controlled by one or more bit-puncturing parameters.

[0325] Clause 11B. A method for processing video data, the receiving device comprising: obtaining error correction data from a transmitting device, wherein the error correction data provides error correction information relating to a picture of the video data; generating prediction data for a picture, wherein the prediction data for a picture includes predictions of blocks of the picture, at least in part on a previously reconstructed picture of one or more video data; generating encoded video data based on the prediction data for a picture, wherein the encoded video data includes a transformation block containing transformation coefficients; scaling the bits of the transformation coefficients of the transformation block based on a confidence value for bit positions; generating error-corrected encoded video data using the error correction data to perform error correction operations on the scaled bits of the transformation coefficients of the transformation block; and reconstructing a picture based on the error-corrected encoded video data.

[0326] Clause 12B. The method of Clause 11B, further comprising generating a reliability value in a receiving device.

[0327] Clause 13B. The method of Clause 12B, wherein generating a reliability value includes generating a reliability value in a receiving device based on statistics regarding the occurrence of errors in bit positions.

[0328] Clause 14B. The method of Clause 12B or 13B, wherein generating a reliability value in a receiving device includes generating a reliability value based on the reliability characteristics of individual areas of a picture of video data.

[0329] Clause 15B. The method described in any one of Clauses 12B to 14B, wherein generating a reliability value includes generating a reliability value based on a noise model in a receiving device.

[0330] Clause 16B. The method of any one of Clauses 11B to 15B, further comprising sending a confidence value to the transmitting device in the receiving device.

[0331] Clause 17B. The method of any one of Clauses 11B to 15B, further comprising receiving a confidence value from a transmitting device in a receiving device.

[0332] Clause 18B. A method for processing video data, comprising: acquiring video data; acquiring predictive quality feedback, wherein the predictive quality feedback is based on the reliability of estimated pictures produced by a receiving device; adapting one or more video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process, controlled by video coding parameters, to generate encoded video data based on one or more pictures of the acquired video data; performing a channel coding process, controlled by channel coding parameters, on the encoded video data to generate channel coded data; and transmitting the channel coded data to a receiving device.

[0333] Clause 19B. The method according to Clause 18B, wherein the video coding parameters include quantization parameters, adapting the video coding parameters includes adapting the quantization parameters, and performing the video coding process includes using the quantization parameters to quantize the conversion coefficients of one or more picture conversion blocks.

[0334] Clause 20B. The method according to Clause 18B or 19B, wherein the channel coding parameters include a low-density parity check (LDPC) graph, adapting the channel coding parameters includes adapting an LPDC graph, and performing the channel coding process includes using the LDPC graph to generate codewords to be contained in the channel-coded data.

[0335] Clause 21B. The method according to any one of Clauses 18B to 20B, wherein channel-encoded data includes error-corrected data, and the method further comprises adapting one or more bit-puncturing parameters based on predictive quality feedback, and performing a bit-puncturing process on the error-corrected data, wherein the bit-puncturing process is controlled by one or more bit-puncturing parameters.

[0336] Clause 22B. A device comprising means for performing the method described in any one of Clauses 11B to 21B.

[0337] Clause 23B. A computer-readable data storage medium storing instructions, wherein, when the instructions are executed, causes a device to perform the method described in any one of Clauses 11B to 21B.

[0338] Clause 1C. A device comprising: a memory configured to store video data; and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors acquire a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint; transmits first encoded video data to a receiving device, the first encoded video data being based on the first set of multiview pictures; and receives a multiview encoding queue from the receiving device. A device having a second set of multiview pictures of video data, the second set of multiview pictures including a third picture and a fourth picture, where the third picture is from a first viewpoint and the fourth picture is from a second viewpoint, and having a multiview coding process on the second set of multiview pictures to generate second encoded video data based on a multiview coding queue received from a receiving device, the multiview coding process having a multiview coding process that reduces interview redundancy between the third picture and the fourth picture, and having transmitted the second encoded video data to the receiving device.

[0339] Clause 2C. The device according to Clause 1, wherein one or more processors, after transmitting second encoded video data to a receiving device, receive an updated multiview encoding queue from the receiving device, obtain a third set of multiview pictures of the video data, wherein the third set of multiview pictures includes a fifth picture and a sixth picture, the fifth picture being from a first viewpoint and the sixth picture being from a second viewpoint, encode the third set of multiview pictures based on the updated multiview encoding queue received from the receiving device to generate third encoded video data, and transmit the third encoded video data to the receiving device.

[0340] Clause 3C. The device described in Clause 1C or 2C, wherein the multiview coding queue includes one or more motion data for relative shifts between blocks of the first picture and blocks of the second picture, brightness corrections between the first picture and the second picture, interblock shifts between anchor blocks and reconstruction blocks, or reference shifts.

[0341] Clause 4C. The device described in any one of Clauses 1C to 3C, wherein the device is an Extended Reality (XR) headset, and one or more processors are configured to receive virtual element data generated from a receiving device based on a first set of multiview pictures and a second set of multiview pictures, and to output the virtual element data for display in an XR scene.

[0342] Clause 5C. A device comprising a memory configured to store video data, and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors receive first encoded video data from a transmitting device, the first encoded video data being based on a first set of multiview pictures of the video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint, and the first encoded A device configured to determine a multiview coding queue based on video data, send the multiview coding queue to a transmitting device, and retrieve second coded video data from the transmitting device, wherein the second coded video data is based on a second set of multiview pictures including a third picture and a fourth picture, and the second coded video data is coded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.

[0343] Clause 6C. The device described in Clause 5C, wherein one or more processors are further configured to decode second encoded video data.

[0344] The device according to Clause 7C, wherein the multiview coding queue is a first multiview coding queue, and one or more processors determine a second multiview coding queue based on second coded video data, transmit the second multiview coding queue to a transmitting device, and is further configured to obtain third coded video data from the transmitting device, wherein the third coded video data is based on a set of third multiview pictures including a fifth picture and a sixth picture, and the third coded video data is coded using a multiview coding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview coding queue.

[0345] Clause 8C. A device according to any one of Clauses 5C to 7C, wherein the multiview coding queue includes a depth map indicating the depth of objects represented in a first picture and a second picture, and one or more processors are configured to determine the depth map based on the first picture and the second picture as part of determining the multiview coding queue.

[0346] Clause 9C. A device according to any one of Clauses 5C to 8C, wherein the multiview coding queue includes one or more illumination compensation factors, and one or more processors are configured to determine the illumination compensation factors based on a first picture and a second picture as part of determining the multiview coding queue.

[0347] Clause 10C. The device described in any one of Clauses 5C to 9C, wherein the transmitting device is an Extended Reality (XR) headset, and one or more processors are further configured to process a second set of pictures to generate virtual element data and transmit the virtual element data to the XR headset.

[0348] Clause 11C. A method for processing video data, comprising: obtaining a first set of multiview pictures of the video data, wherein the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multiview pictures; receiving a multiview encoding queue from the receiving device; and obtaining a second set of multiview pictures of the video data, where the second A method comprising: obtaining a second set of multiview pictures, the set of multiview pictures including a third picture and a fourth picture, the third picture being from a first viewpoint and the fourth picture being from a second viewpoint; performing a multiview coding process on the second set of multiview pictures to generate second coded video data based on a multiview coding queue received from a receiving device, the multiview coding process reducing interview redundancy between the third picture and the fourth picture; and transmitting the second coded video data to a receiving device.

[0349] The method according to Clause 12C, further comprising: transmitting second encoded video data to a receiving device, receiving an updated multiview encoding queue from the receiving device; obtaining a third set of multiview pictures of the video data, wherein the third set of multiview pictures includes a fifth picture and a sixth picture, where the fifth picture is from a first viewpoint and the sixth picture is from a second viewpoint; encoding the third set of multiview pictures based on the updated multiview encoding queue received from the receiving device to generate third encoded video data; and transmitting the third encoded video data to the receiving device.

[0350] Clause 13C. The method according to Clause 11C or 12C, wherein the multiview coding queue includes one or more motion data for relative shifts between blocks of a first picture and blocks of a second picture, brightness corrections between the first picture and the second picture, interblock shifts between anchor blocks and reconstruction blocks, or reference shifts.

[0351] Clause 14C. The method according to any one of Clauses 11C to 13C, further comprising receiving virtual element data generated based on a first set of multiview pictures and a second set of multiview pictures from a receiving device, and outputting the virtual element data for display in an Extended Reality (XR) scene.

[0352] Clause 15C. A method for processing video data, comprising: obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint; determining a multiview coding queue based on the first encoded video data; transmitting the multiview coding queue to the transmitting device; and obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multiview pictures comprising a third picture and a fourth picture, the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.

[0353] Clause 16C. The method of Clause 15C, further comprising decoding a second encoded video data.

[0354] Clause 17C. The method according to Clause 15C or 16C, further comprising: multiview coding queue being a first multiview coding queue, the method determining a second multiview coding queue based on second coded video data; transmitting the second multiview coding queue to a transmitting device; and obtaining third coded video data from the transmitting device, wherein the third coded video data is based on a set of third multiview pictures including a fifth picture and a sixth picture, and the third coded video data is coded using a multiview coding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview coding queue.

[0355] Clause 18C. The method according to any one of Clauses 15C to 17C, wherein the multiview coding queue includes a depth map indicating the depth of objects represented in a first picture and a second picture, and determining the multiview coding queue includes determining the depth map based on the first picture and the second picture.

[0356] Clause 19C. The method according to any one of Clauses 15C to 18C, wherein the multiview coding queue includes one or more illumination compensation coefficients, and determining the multiview coding queue includes determining the illumination compensation coefficients based on a first picture and a second picture.

[0357] Clause 20C. The method according to any one of Clauses 15C to 19C, wherein the transmitting device is an Extended Reality (XR) headset, and the method further comprises processing a second set of pictures to generate virtual element data, and transmitting the virtual element data to the XR headset.

[0358] Clause 21C. A device comprising means for acquiring a first set of multiview pictures of video data, wherein the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint; means for transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multiview pictures; means for receiving a multiview encoding queue from the receiving device; and a second set of multiview pictures of video data, wherein the second multiview picture A device comprising: means for acquiring a second set of multiview pictures, the set of view pictures including a third picture and a fourth picture, the third picture being from a first viewpoint and the fourth picture being from a second viewpoint; means for performing a multiview coding process on the second set of multiview pictures to generate second coded video data based on a multiview coding queue received from a receiving device, the multiview coding process reducing interview redundancy between the third picture and the fourth picture; and means for transmitting the second coded video data to a receiving device.

[0359] Clause 22C. A device comprising: means for acquiring first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of video data, the first set of multiview pictures comprising a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint; means for determining a multiview coding queue based on the first encoded video data; means for transmitting the multiview coding queue to the transmitting device; and means for acquiring second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multiview pictures comprising a third picture and a fourth picture, the second encoded video data is encoded using a multiview coding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview coding queue.

[0360] Clause 1D. A device comprising a memory configured to store video data, and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors are configured to encode a first set of pictures of video data to generate first encoded video data, transmit the first encoded video data to a receiving device, receive a decimation pattern indication from the receiving device indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data, encode a second set of pictures of video data to generate second encoded video data, apply the decimation pattern to the second encoded video data to generate decimated video data, and transmit the decimated video data to a receiving device.

[0361] Clause 2D. The device according to Clause 1D, wherein one or more processors are configured to generate first error correction data based on first encoded video data and transmit the first error correction data to a receiving device, and to generate second error correction data based on second encoded video data and transmit the second error correction data to a receiving device.

[0362] Clause 3D: A device as described in Clause 1D or 2D, in which the decimation pattern indicates a pattern that skips the transmission of encoded video data of the full picture.

[0363] Clause 4D. A device as described in any one of Clauses 1D to 3D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specific area within a picture.

[0364] Clause 5D. A device as described in any one of Clauses 1D to 4D, wherein the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a particular view.

[0365] Clause 6D. A device according to any one of Clauses 1D to 5D, wherein the decimation pattern indication is a first decimation pattern indication, the untransmitted pattern of the encoded video data is a first untransmitted pattern of the encoded video data, the decimated video data is the first decimated video data, and one or more processors encode a third set of pictures of video data to generate a third encoded video data, determine a second decimation pattern indicating a second untransmitted pattern of the encoded video data, apply the second decimation pattern to the third encoded video data to generate the second decimated video data, transmit the second decimated video data to a receiving device, and further configured to transmit a second decimation pattern indication to the receiving device, the second decimation pattern indication indicating that the second decimation pattern has been applied to the third encoded video data.

[0366] Clause 7D. A device according to any one of Clauses 1D to 6D, wherein one or more processors are configured to generate first prediction data for a first set of first pictures as part of encoding a first set of pictures, generate residual data based on the first prediction data and the first set of pictures, apply a transformation of the first prediction data to generate a transformation block, quantize the transformation coefficients of the transformation block, and apply entropy coding to syntax elements representing the quantized transformation coefficients to generate a first entropy-coded syntax element, the first encoded video data comprising the first entropy-coded syntax element, and one or more processors are further configured to perform analog modulation on the residual data to generate first analog-modulated residual data, the device further comprising a communication interface configured to transmit the first analog-modulated residual data and the first encoded video data.

[0367] Clause 8D. The device described in any one of Clauses 1D to 7D, wherein the device is an Extended Reality (XR) headset, comprising a display system, one or more processors further configured to receive virtual element data from a receiving device, and the display system configured to display one or more virtual elements in an XR scene based on the virtual element data.

[0368] Clause 9D. A device comprising a memory configured to store video data, and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors receive first encoded video data from a transmitting device, perform a decoding process to reconstruct a first set of pictures based on the first encoded video data, determine a decimation pattern indicating a non-transmitted pattern of the encoded video data based on the first set of pictures, transmit a decimation pattern indication indicating the determined decimation pattern to the transmitting device, receive decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data, and perform a decoding process to reconstruct a second set of pictures based on the second encoded video data.

[0369] Clause 10D. The device according to Clause 9D, wherein one or more processors are configured to receive first error correction data from a transmitting device and apply an error correction process to correct the first encoded video data based on the first error correction data in order to generate first error-corrected encoded video data; one or more processors are configured to perform a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data; one or more processors are further configured to receive second error correction data from a transmitting device and apply an error correction process to generate second error-corrected encoded video data based on the second encoded video data and the second error correction data; and one or more processors are configured to perform a decoding process to reconstruct a second set of pictures based on the second error-corrected encoded video data.

[0370] Clause 11D. A device as described in Clause 9D or 10D, in which the decimation pattern indicates a pattern that skips the transmission of encoded video data of the full picture.

[0371] Clause 12D. A device according to any one of Clauses 9D to 11D, wherein one or more processors are configured to apply a decimation pattern to first encoded video data in order to generate decimated encoded video data as part of determining a decimation pattern; to apply an error correction process to correct the decimated encoded video data based on first error correction data in order to generate trial error corrected encoded video data; to apply a decoding process to reconstruct a first set of pictures based on the trial error corrected video data; and to determine whether the decimation pattern meets a criterion based on a comparison between the first set of pictures reconstructed based on the trial error corrected video data and the first set of pictures reconstructed based on the first video data.

[0372] Clause 13D. A device as described in any one of Clauses 9D to 12D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specified area within a picture.

[0373] Clause 14D. A device as described in any one of Clauses 9D to 13D, wherein the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data of pictures from a specified view.

[0374] Clause 15D. A device according to any one of Clauses 9D to 14D, wherein the decimation pattern indication is a first decimation pattern indication, the untransmitted pattern of the encoded video data is a first untransmitted pattern of the encoded video data, the decimated video data is the first decimated video data, and one or more processors receive a second decimation pattern indication showing a second untransmitted pattern of the encoded video data, and are further configured to receive a second decimated video data from a transmitting device, the second decimated video data comprising a third encoded video data to which the second decimation pattern has been applied, and the third encoded video data is generated based on a third set of pictures of the video data, and to apply a decoding process to reconstruct the third set of pictures based on the third encoded video data.

[0375] Clause 16D. The device according to any one of Clauses 9D to 15D, further comprising a communication interface configured to receive analog modulation residual data, wherein second encoded video data includes entropy coding syntax elements representing quantization conversion coefficients, and one or more processors are configured to apply entropy decoding to the syntax elements to obtain quantized conversion coefficients as part of applying a decoding process to reconstruct a second set of pictures, inverse quantization of the quantization conversion coefficients to produce inverse quantized conversion coefficients, inverse transformation to the inverse quantized conversion coefficients to produce prediction data, demodulate the analog modulation residual data to obtain residual data, and reconstruct a second set of pictures based on the prediction data and residual data.

[0376] Clause 17D. A device as described in any one of Clauses 9D to 16D, wherein one or more processors are further configured to process a second set of pictures to generate virtual element data, and the transmitting device is an XR headset configured to display one or more virtual elements in an Extended Reality (XR) scene based on the virtual element data.

[0377] Clause 18D. A method comprising: encoding a first set of pictures of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving a decimation pattern indication from the receiving device, which indicates a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; encoding a second set of pictures of video data to generate second encoded video data; applying the decimation pattern to the second encoded video data to generate decimated video data; and transmitting the decimated video data to a receiving device.

[0378] The method according to Clause 18D, further comprising generating first error correction data based on first encoded video data, transmitting the first error correction data to a receiving device, generating second error correction data based on second encoded video data, and transmitting the second error correction data to a receiving device.

[0379] Clause 20D. The method of Clause 18D or 19D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data of the full picture.

[0380] Clause 21D. The method described in any one of Clauses 18D to 20D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specific area within a picture.

[0381] Clause 22D. The method described in any one of Clauses 18D to 21D, wherein the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a particular view.

[0382] Clause 23D. The method according to any one of Clauses 18D to 22D, wherein the decimation pattern indication is a first decimation pattern indication, the untransmitted pattern of the encoded video data is a first untransmitted pattern of the encoded video data, and the decimated video data is the first decimated video data, and the method further comprises encoding a third set of pictures of video data to generate a third encoded video data, determining a second decimation pattern indicating a second untransmitted pattern of the encoded video data, applying the second decimation pattern to the third encoded video data to generate the second decimated video data, transmitting the second decimated video data to a receiving device, and transmitting a second decimation pattern indication to the receiving device, the second decimation pattern indication indicating that the second decimation pattern has been applied to the third encoded video data.

[0383] Clause 24D. The method according to any one of Clauses 18D to 23D, wherein encoding a first set of pictures includes generating first prediction data for the first set of pictures, generating residual data based on the first prediction data and the first set of pictures, applying a transformation to the first prediction data to generate a transformation block, quantizing the transformation coefficients of the transformation block, and applying entropy coding to syntax elements representing the quantized transformation coefficients to generate a first entropy-coded syntax element, wherein the first encoded video data includes the first entropy-coded syntax element, and the method further includes performing analog modulation on the residual data to generate first analog-modulated residual data, and transmitting the first analog-modulated residual data and the first encoded video data.

[0384] The method described in any one of the clauses 18D to 24D, further comprising receiving virtual element data from a receiving device and displaying one or more virtual elements in an Extended Reality (XR) scene based on the virtual element data.

[0385] Clause 26D. A method comprising: receiving first encoded video data from a transmitting device; applying a decoding process to reconstruct a first set of pictures based on the first encoded video data; determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; transmitting a decimation pattern indication indicating the determined decimation pattern to the transmitting device; receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.

[0386] The method according to Clause 27D, further comprising: receiving first error correction data from a transmitting device; applying an error correction process to correct the first encoded video data based on the first error correction data in order to generate first error-corrected encoded video data; and performing a decoding process to reconstruct a first set of pictures based on the first error-corrected encoded video data; and further comprising: receiving second error correction data from a transmitting device; applying an error correction process to generate second error-corrected encoded video data based on the second encoded video data and the second error correction data; and performing a decoding process to reconstruct a second set of pictures based on the second error-corrected encoded video data.

[0387] Clause 28D. The method of Clause 26D or 27D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data of a full picture.

[0388] Clause 29D. The method of any one of Clauses 26D to 28D, wherein determining a decimation pattern includes applying a decimation pattern to first encoded video data to generate decimated encoded video data; applying an error correction process to correct the decimated encoded video data based on first error correction data to generate trial error corrected encoded video data; applying a decoding process to reconstruct a first set of pictures based on the trial error corrected video data; and determining whether the decimation pattern meets a criterion based on a comparison of the first set of pictures reconstructed based on the trial error corrected video data with the first set of pictures reconstructed based on the first video data.

[0389] Clause 30D. The method described in any one of Clauses 26D to 29D, wherein the decimation pattern indicates a pattern that skips the transmission of encoded video data for a specified area within a picture.

[0390] Clause 31D. The method described in any one of Clauses 26D to 30D, wherein the video data is multiview video data and the decimation pattern indicates a pattern that skips the transmission of encoded video data for pictures from a specified view.

[0391] Clause 32D. The method according to any one of Clauses 26D to 31D, wherein the decimation pattern indication is a first decimation pattern indication, the untransmitted pattern of the encoded video data is a first untransmitted pattern of the encoded video data, and the decimated video data is the first decimated video data, the method further comprising: receiving a second decimation pattern indication indicating a second untransmitted pattern of the encoded video data; receiving a second decimated video data from a transmitting device, wherein the second decimated video data includes a third encoded video data to which the second decimation pattern has been applied, and the third encoded video data is generated based on a third set of pictures of the video data; and applying a decoding process to reconstruct the third set of pictures based on the third encoded video data.

[0392] Clause 33D. The method according to any one of Clauses 26D to 32D, further comprising receiving analog modulation residual data, the second encoded video data comprising entropy coding syntax elements representing quantization transformation coefficients, and applying a decoding process to reconstruct a second set of pictures, comprising applying entropy decoding to the syntax elements to obtain quantized transformation coefficients, inverse quantization of the quantization transformation coefficients to produce inverse quantized transformation coefficients, applying inverse transformation to the inverse quantized transformation coefficients to produce prediction data, demodulating the analog modulation residual data to obtain residual data, and reconstructing a second set of pictures based on the prediction data and the residual data.

[0393] Clause 34D. The method according to any one of Clauses 26D to 33D, further comprising processing a second set of pictures to generate virtual element data, wherein the transmitting device is an XR headset configured to display one or more virtual elements in an Extended Reality (XR) scene based on the virtual element data.

[0394] Clause 35D. A device comprising: means for encoding a first set of pictures of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving a decimation pattern indication from the receiving device, the decimation pattern indicating a decimation pattern determined based on the first set of pictures, wherein the decimation pattern is a non-transmitted pattern of encoded video data; means for encoding a second set of pictures of video data to generate second encoded video data; means for applying the decimation pattern to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to a receiving device.

[0395] Clause 36D. A device comprising: means for receiving first encoded video data from a transmitting device; means for applying a decoding process to reconstruct a first set of pictures based on the first encoded video data; means for determining a decimation pattern indicating a non-transmission pattern of the encoded video data based on the first set of pictures; means for transmitting a decimation pattern indication indicating the determined decimation pattern to the transmitting device; means for receiving decimated video data from the transmitting device, wherein the decimated video data comprises second encoded video data to which the decimation pattern has been applied, and the second encoded video data is generated based on a second set of pictures of the video data; and means for performing a decoding process to reconstruct a second set of pictures based on the second encoded video data.

[0396] Clause 1E. A device comprising a memory configured to store video data, and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors encode a first picture of video data to generate first encoded video data, transmit the first encoded video data to a receiving device, receive from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, the second picture follows the first picture in decoding order, encodes the second picture based on the encoded selection data for the second picture to generate second encoded video data, and transmits the second encoded video data to a receiving device.

[0397] Clause 2E. The device described in Clause 1E, wherein the encoding selection data received from the receiving device includes motion parameters for a block of a second picture, and one or more processors are configured to perform motion compensation based on the motion parameters for a block of a second picture in order to generate a prediction block as part of encoding the second picture, and the second encoded video data includes video data encoded based on the prediction block.

[0398] Clause 3E. The device described in Clause 1E or 2E, wherein the encoding selection data received from the receiving device includes intra-prediction parameters for a block of a second picture, and one or more processors are configured to perform intra-prediction based on the intra-prediction parameters for a block of a second picture in order to generate a predictive block as part of encoding the second picture, and the second encoded video data includes video data encoded based on the predictive block.

[0399] Clause 4E. A device described in any one of Clauses 1E to 3E, in which the second encoded video data does not include the encoded selection data.

[0400] Clause 5E. A device according to any one of Clauses 1E to 4E, wherein one or more processors are configured to entropy decode the encoding selection data prior to the second picture.

[0401] Clause 6E. A device according to any one of Clauses 1E to 5E, further configured to have one or more processors that generate first error correction data based on first encoded video data and transmit the first encoded video data and the first error correction data to a receiving device.

[0402] Clause 7E. A device according to any one of Clauses 1E to 6E, wherein one or more processors are further configured to encode a third picture without using encoding selection data for the third picture, based on the determination that encoding selection data for the third picture will not be received from the receiving device before the expiration of the time limit.

[0403] Clause 8E. A device according to any one of Clauses 1E to 7E, wherein one or more processors receive encoding selection data for a third picture of video data, wherein the encoding selection data for the third picture indicates an encoding selection used to encode an estimate of the third picture; encodes the third picture based on the encoding selection data for the third picture to generate third encoded video data; applies a channel coding process to generate error correction data for the third encoded video data; and is further configured to transmit the error correction data for the third encoded video data to a receiving device without transmitting at least a portion of the third encoded video data.

[0404] Clause 9E. A device as described in any one of Clauses 1E to 8E, wherein the device is an Extended Reality (XR) headset, comprising a display system, one or more processors further configured to receive virtual element data from a receiving device, and the display system configured to display one or more virtual elements in an XR scene based on the virtual element data.

[0405] Clause 10E. A device comprising a memory configured to store video data, and one or more processors implemented in the circuit and coupled to the memory, wherein one or more processors are configured to receive first encoded video data from a transmitting device, reconstruct a first picture of the video data based on the first encoded video data, estimate a second picture of the video data based on the first picture, wherein the second picture is a picture that occurs after the first picture in the decoding order, generate encoding selection data for the estimated second picture, wherein the encoding selection data indicates the encoding selection used to encode the estimated second picture, transmit the encoding selection data for the second picture to the transmitting device, receive second encoded video data from the transmitting device, and reconstruct a second picture based on the second encoded video data.

[0406] Clause 11E. The device described in Clause 10E, wherein one or more processors are configured to perform motion compensation based on motion parameters for blocks of a second picture in order to generate predicted blocks as part of encoding a predicted second picture, and the encoded selection data includes motion parameters for blocks of a second picture, and the second encoded video data includes video data encoded based on the predicted blocks.

[0407] Clause 12E. The device according to Clause 10E or 11E, wherein one or more processors are configured to perform intra-prediction based on intra-prediction parameters for blocks of a second picture in order to generate predictive blocks as part of encoding a second picture, and the encoded selection data includes intra-prediction parameters for blocks of a second picture, and the second encoded video data includes video data encoded based on the predictive blocks.

[0408] Clause 13E. A device as described in any one of Clauses 10E to 12E, wherein the second encoded video data does not include encoding selection data, and one or more processors are configured to use the encoding selection data to reconstruct a second picture based on the second encoded video data as part of applying a decoding process.

[0409] Clause 14E. A device according to any one of Clauses 10E to 13E, wherein one or more processors are configured to entropically encode the encoding selection data for the second picture before transmitting the encoding selection data for the second picture.

[0410] Clause 15E. A device according to any one of Clauses 10E to 14E, wherein one or more processors are further configured to estimate a third picture of video data based on one or more of the first or second pictures, and to generate a third encoded video data, by performing an encoding process to encode the estimated third picture, wherein the third encoding selection data indicates the encoding selection used to encode the estimated third picture, transmit the third encoding selection data to a transmitting device, receive error correction data for the third picture from the transmitting device, apply an error correction process to generate error-corrected encoded video data of the third picture based on the error correction data of the third picture and the third encoded video data, and apply a decoding process to reconstruct the third picture based on the error-corrected encoded video data of the third picture.

[0411] Clause 16E. The device described in Clause 15E, wherein the error-corrected encoded video data for the third picture does not include third encoding selection data, and one or more processors are configured to use the third encoding selection data to reconstruct the third picture based on the error-corrected encoded video data for the third picture as part of applying a decoding process.

[0412] Clause 17E. A device according to any one of Clauses 10E to 16E, wherein one or more processors are configured to apply a channel coding process to the coded selection data for the second picture in order to generate error correction data for the coded selection data for the second picture, and to transmit the error correction data for the coded selection data for the second picture to a transmitting device.

[0413] Clause 18E. A device as described in any one of Clauses 10E to 17E, wherein the device has a communication interface configured to modulate coded selection data at a lower modulation order compared to other data transmissions in the data link between the device and the transmitting device.

[0414] Clause 19E. A device as described in any one of Clauses 10E to 18E, wherein one or more processors are further configured to process a second set of pictures to generate virtual element data, and the transmitting device is an XR headset configured to display one or more virtual elements in an Extended Reality (XR) scene based on the virtual element data.

[0415] Clause 20E. A method for processing video data, comprising: encoding a first picture of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in decoding order; encoding a second picture based on the encoding selection data for the second picture to generate second encoded video data; and transmitting the second encoded video data to a receiving device.

[0416] Clause 21E. The method according to Clause 20E, wherein the encoding selection data received from the receiving device includes motion parameters for a block of a second picture, and encoding the second picture includes performing motion compensation based on the motion parameters for the block of the second picture to generate a predicted block, and the second encoded video data includes video data encoded based on the predicted block.

[0417] Clause 22E. The method according to Clause 20E or 21E, wherein the encoding selection data received from the receiving device includes intra-prediction parameters for a block of a second picture, and encoding the second picture includes performing an intra-prediction based on the intra-prediction parameters for a block of a second picture in order to generate a prediction block, and the second encoded video data includes video data encoded based on the prediction block.

[0418] Clause 23E. The method described in any one of Clauses 20E to 22E, wherein the second encoded video data does not include the encoded selection data.

[0419] Clause 24E. The method of any one of Clauses 20E to 23E, further comprising entropy decoding the encoding selection data prior to the second picture.

[0420] The method described in any one of the clauses 20E to 24E, further comprising generating first error correction data based on first encoded video data, and transmitting the first encoded video data and the first error correction data to a receiving device.

[0421] Clause 26E. The method of any one of Clauses 20E to 25E, further comprising encoding the third picture without using the encoding selection data for the third picture, based on the determination that the encoding selection data for the third picture is not received from the receiving device before the expiration of the time limit.

[0422] Clause 27E. The method of any one of Clauses 20E to 26E, further comprising: receiving coding selection data for a third picture of video data, wherein the coding selection data for the third picture indicates a coding selection used to encode an estimate of the third picture; encoding a third picture based on the coding selection data for the third picture to generate a third encoded video data; applying a channel coding process to generate error correction data for the third encoded video data; and transmitting the error correction data for the third encoded video data to a receiving device without transmitting at least a portion of the third encoded video data.

[0423] Clause 28E. The method according to any one of Clauses 20E to 27E, wherein the device is an Extended Reality (XR) headset and comprises a display system, and the method further comprises receiving virtual element data from a receiving device and displaying one or more virtual elements in an XR scene on the display system based on the virtual element data.

[0424] Clause 29E. A method for processing video data, comprising: receiving first encoded video data from a transmitting device; reconstructing a first picture of the video data based on the first encoded video data; estimating a second picture of the video data based on the first picture, wherein the second picture is a picture that occurs after the first picture in the decoding order; generating encoding selection data for the estimated second picture, wherein the encoding selection data indicates the encoding selection used to encode the estimated second picture; transmitting the encoding selection data for the second picture to a transmitting device; receiving second encoded video data from a transmitting device; and reconstructing the second picture based on the second encoded video data.

[0425] Clause 30E. The method according to Clause 29E, wherein encoding an estimated second picture includes performing motion compensation based on motion parameters for a block of the second picture to generate a predicted block, the encoded selection data includes motion parameters for a block of the second picture, and the second encoded video data includes video data encoded based on the predicted block.

[0426] Clause 31E. The method according to Clause 29E or 30E, wherein encoding a second picture comprises performing an intra-prediction based on intra-prediction parameters for a block of the second picture in order to generate a predictive block, the encoded selection data comprises the intra-prediction parameters for a block of the second picture, and the second encoded video data comprises the video data encoded based on the predictive block.

[0427] Clause 32E. The method described in any one of Clauses 29E to 31E, wherein the second encoded video data does not include encoded selection data, and the decoding process involves using the encoded selection data to reconstruct a second picture based on the second encoded video data.

[0428] Clause 33E. The method described in any one of Clauses 29E to 32E, wherein the encoding selection data for the second picture is entropy encoded before the encoding selection data for the second picture is transmitted.

[0429] Clause 34E. The method of any one of Clauses 29E to 33E, further comprising: estimating a third picture of video data based on one or more of the first or second pictures; performing an encoding process to encode the estimated third picture, wherein the estimated third picture is such that the third encoding selection data indicates the encoding selection used to encode the estimated third picture, in order to generate third encoded video data; transmitting the third encoding selection data to a transmitting device; receiving error correction data for the third picture from the transmitting device; applying an error correction process to generate error-corrected encoded video data of the third picture based on the error correction data of the third picture and the third encoded video data; and applying a decoding process to reconstruct the third picture based on the error-corrected encoded video data of the third picture.

[0430] Clause 35E. The method according to Clause 34E, wherein the error-corrected encoded video data for the third picture does not include the third encoded selection data, and the decoding process involves using the third encoded selection data to reconstruct the third picture based on the error-corrected encoded video data for the third picture.

[0431] The method of any one of the clauses 29E to 35E, further comprising applying a channel coding process to the coded selection data for the second picture in order to generate error correction data for the coded selection data for the second picture, and transmitting the error correction data for the coded selection data for the second picture to a transmitting device.

[0432] Clause 37E. The method of any one of Clauses 29E to 36E, further comprising modulating the encoded selected data at a lower modulation order compared to other data transmissions in a data link between the device and the transmitting device.

[0433] The method described in any one of the clauses 29E to 37E, further comprising processing a second set of pictures to generate virtual element data, wherein the transmitting device is an XR headset configured to display one or more virtual elements in an Extended Reality (XR) scene based on the virtual element data.

[0434] Clause 39E. A device comprising: means for encoding a first picture of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device encoded selection data for a second picture of video data, wherein the encoded selection data for the second picture indicates an encoding selection used to encode an estimate of the second picture, and the second picture follows the first picture in decoding order; means for encoding a second picture based on the encoding selection data for the second picture to generate second encoded video data; and means for transmitting the second encoded video data to a receiving device.

[0435] Clause 40E. A device comprising: means for receiving a first encoded video data from a transmitting device; means for reconstructing a first picture of the video data based on the first encoded video data; means for estimating a second picture of the video data based on the first picture, wherein the second picture is a picture that occurs after the first picture in the decoding order; means for generating encoding selection data for the second picture, wherein the enco...

Claims

1. It is a device, A memory configured to store video data, The circuit comprises one or more processors implemented in the circuit and coupled to the memory, wherein the one or more processors Obtain a first set of multiview pictures of the video data, wherein the first set of multiview pictures includes a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint. A first encoded video data, wherein the first encoded video data is based on the first set of multiview pictures and is transmitted to a receiving device. The receiving device receives a multiview coding queue, A second set of multiview pictures of the video data is obtained, wherein the second set of multiview pictures includes a third picture and a fourth picture, the third picture being from the first viewpoint and the fourth picture being from the second viewpoint. Based on the multiview coding queue received from the receiving device, a multiview coding process is performed on the second set of multiview pictures to generate second coded video data, wherein the multiview coding process reduces interview redundancy between the third picture and the fourth picture. A device configured to transmit the second encoded video data to the receiving device.

2. The aforementioned one or more processors After transmitting the second encoded video data to the receiving device, the system receives an updated multiview encoding queue from the receiving device. A third set of multiview pictures of the video data is obtained, wherein the third set of multiview pictures includes a fifth picture and a sixth picture, the fifth picture is from the first viewpoint and the sixth picture is from the second viewpoint. To generate third encoded video data, the set of third multiview pictures is encoded based on the updated multiview encoding queue received from the receiving device, The device according to claim 1, further configured to transmit the third encoded video data to the receiving device.

3. The multiview coding queue, The relative shift between the first picture block and the second picture block, Brightness correction between the first picture and the second picture, Inter-block shift between anchor blocks and reconstruction blocks, or The device according to claim 1, comprising one or more motion data for reference shift.

4. The aforementioned device is an Extended Reality (XR) headset, The aforementioned one or more processors The receiving device receives virtual element data generated based on the first set of multiview pictures and the second set of multiview pictures. The device according to claim 1, configured to output the virtual element data for display in an XR scene.

5. It is a device, A memory configured to store video data, The circuit comprises one or more processors implemented in the circuit and coupled to the memory, wherein the one or more processors Obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of the video data, the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint. Based on the first encoded video data, a multiview encoding queue is determined. The multiview coding queue is transmitted to the transmitting device. A device configured to receive second encoded video data from the transmitting device, wherein the second encoded video data is encoded based on a second set of multiview pictures including a third picture and a fourth picture, and the second encoded video data is encoded using a multiview encoding process that reduces interview redundancy between the third picture and the fourth picture based on the multiview encoding queue.

6. The device according to claim 5, wherein one or more processors are further configured to decode the second encoded video data.

7. The multi-view coding queue is a first multi-view coding queue, and the one or more processors are A second multiview coding queue is determined based on the second encoded video data. The second multiview coding queue is transmitted to the transmitting device. The device according to claim 5, further configured to obtain third encoded video data from the transmitting device, wherein the third encoded video data is encoded based on a set of third multiview pictures including a fifth picture and a sixth picture, and the third encoded video data is encoded using a multiview encoding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview encoding queue.

8. The multiview coding queue includes a depth map indicating the depth of objects represented in the first picture and the second picture. The device according to claim 5, wherein one or more processors are configured to determine the depth map based on the first picture and the second picture as part of determining the multiview coding queue.

9. The multiview coding queue includes one or more illumination compensation coefficients, The device according to claim 5, wherein one or more processors are configured to determine the illumination compensation coefficient based on the first picture and the second picture as part of determining the multiview coding queue.

10. The aforementioned transmitting device is an Extended Reality (XR) headset, One or more processors, The second set of pictures is processed in order to generate virtual element data. The device according to claim 5, further configured to transmit the virtual element data to the XR headset.

11. A method for processing video data, Obtaining a first set of multiview pictures of the video data, wherein the first set of multiview pictures includes a first picture and a second picture, the first picture being from a first viewpoint and the second picture being from a second viewpoint. A first encoded video data, wherein the first encoded video data is based on the first set of multiview pictures and is transmitted to a receiving device. Receiving a multiview coded queue from the aforementioned receiving device, Obtaining a second set of multiview pictures of the video data, wherein the second set of multiview pictures includes a third picture and a fourth picture, the third picture being from the first viewpoint and the fourth picture being from the second viewpoint. Based on the multiview coding queue received from the receiving device, a multiview coding process is performed on the second set of multiview pictures to generate second coded video data, wherein the multiview coding process reduces interview redundancy between the third picture and the fourth picture. A method comprising transmitting the second encoded video data to the receiving device.

12. After transmitting the second encoded video data to the receiving device, the system receives an updated multiview encoding queue from the receiving device. To obtain a third set of multiview pictures of the video data, wherein the third set of multiview pictures includes a fifth picture and a sixth picture, the fifth picture being from the first viewpoint and the sixth picture being from the second viewpoint. To generate third encoded video data, the third set of multiview pictures is encoded based on the updated multiview encoding queue received from the receiving device, The method according to claim 11, further comprising transmitting the third encoded video data to the receiving device.

13. The multiview coding queue, The relative shift between the first picture block and the second picture block, Brightness correction between the first picture and the second picture, Inter-block shift between anchor blocks and reconstruction blocks, or The method according to claim 11, comprising one or more motion data for reference shift.

14. The method described above is The receiving device receives virtual element data generated based on the first set of multiview pictures and the second set of multiview pictures, The method according to claim 11, further comprising outputting the virtual element data for display in an Extended Reality (XR) scene.

15. A method for processing video data, To obtain first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multiview pictures of the video data, and the first set of multiview pictures includes a first picture and a second picture, where the first picture is from a first viewpoint and the second picture is from a second viewpoint. Determining a multiview coding queue based on the first encoded video data, The multiview coding queue is transmitted to the transmitting device, A method comprising: obtaining second encoded video data from the transmitting device, wherein the second encoded video data is encoded using a multiview encoding process that reduces interview redundancy between the third picture and the fourth picture based on a set of second multiview pictures including a third picture and a fourth picture.

16. The method according to claim 15, further comprising decoding a second encoded video data.

17. The multiview coding queue is a first multiview coding queue, and the method is Determining a second multiview coding queue based on the second encoded video data, Transmitting the second multiview coding queue to the transmitting device, The method according to claim 15, further comprising obtaining third encoded video data from the transmitting device, wherein the third encoded video data is encoded based on a set of third multiview pictures including a fifth picture and a sixth picture, and the third encoded video data is encoded using a multiview encoding process that reduces interview redundancy between the fifth picture and the sixth picture based on the second multiview encoding queue.

18. The multiview coding queue includes a depth map indicating the depth of objects represented in the first picture and the second picture. The method according to claim 15, wherein determining the multiview coding queue includes determining the depth map based on the first picture and the second picture.

19. The multiview coding queue includes one or more illumination compensation coefficients, The method according to claim 15, wherein determining the multiview coding queue includes determining the illumination compensation coefficient based on the first picture and the second picture.

20. The aforementioned transmitting device is an Extended Reality (XR) headset, The method described above is Processing the second set of pictures in order to generate virtual element data, The method according to claim 15, further comprising transmitting the virtual element data to the XR headset.