Receiver selection decimation scheme for video coding
By using distributed video decoding technology, the transmitting device performs limited encoding and sends error correction data, while the receiving device performs complex decoding and reconstruction. This solves the problems of resource-intensive and energy-intensive video encoding processes and achieves low-latency and low-power video data transmission.
Patent Information
- Application Number
- CN202480025289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2024-04-05
- Publication Date
- 2025-11-11
AI Technical Summary
Existing video encoding processes are resource-intensive and complex, resulting in excessive latency and energy consumption when transmitting high-quality video in wireless communication, especially in short-range wireless communication systems where it is difficult to achieve low-latency and low-power video data transmission.
Distributed video decoding (DVC) technology is employed, where the transmitting device performs limited video coding and transmits error correction data, and the receiving device estimates video data based on the error correction data and previously reconstructed images, uses more sophisticated decoding tools for reconstruction and error correction coding, reduces the resource consumption of the transmitting device, and optimizes data transmission through decimation mode and channel coding.
It effectively reduces the resource consumption and data transmission volume of the transmitting device, improves decoding efficiency, and realizes low-latency and low-power video data transmission, making it suitable for extended reality devices and wireless communication systems.
Smart Images

Figure CN120937348A_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 18 / 306,142, filed April 24, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to video encoding and decoding. Background Technology
[0003] The adoption of virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies is growing rapidly and is expected to be widely used in applications beyond gaming, such as healthcare, education, social networking, and retail. VR, AR, and MR can be collectively referred to as extended reality (XR). Due to this increasing popularity, the demand for XR devices, such as XR glasses, with high-quality 3D graphics, higher video resolution, and low latency response is also increasing. Summary of the Invention
[0004] This disclosure describes techniques for processing video data in a transmitting and receiving device. The transmitting device may be an XR device or other type of device. The receiving device may be a user equipment (UE) device, such as a smartphone or tablet. The transmitting device may perform a limited video coding process on the video data to generate coded video data. The transmitting device may apply channel coding to the coded video data to generate error correction data. The transmitting device may transmit the error correction data and at least some coded video data to the receiving device. The receiving device may estimate the video data based on one or more previously reconstructed images. The receiving device may then encode the estimated video data. The receiving device may use one or more decoding tools to encode the estimated video data that the transmitting device did not use when performing the limited video coding process on the video data. The receiving device may use the error correction data and the predicted video data to regenerate the portions of the coded video data that the transmitting device did not transmit. This process avoids the need to transmit portions of the coded video data.
[0005] In one example, this disclosure describes a method for decoding video data, comprising: obtaining error correction data at a receiving device and from a transmitting device, wherein the error correction data provides error correction information and is generated from encoded video data of one or more blocks of an image of the video data; generating prediction data of the image at the receiving device using one or more decoding tools not used to generate the encoded video data of one or more blocks, wherein the prediction data of the image includes predictions of blocks of the image based at least partially on one or more previously reconstructed blocks of the image from the video data; generating encoded video data at the receiving device based on the prediction data of the image; generating error-corrected encoded video data at the receiving device using the error correction data to perform an error correction operation on the encoded video data; and performing a reconstruction operation at the receiving device, the reconstruction operation reconstructing blocks of the image based on the error-corrected encoded video data, wherein the reconstruction operation is controlled by the values of one or more parameters.
[0006] In another example, this disclosure describes a method for encoding video data, comprising: obtaining video data from a video source at a transmitting device; generating encoded video data of a first image of the video data and encoded video data of a second image of the video data based on a set of parameters at the transmitting device; performing channel coding on the encoded video data of the first image and the encoded video data of the second image at the transmitting device to generate error-corrected data of the first image and error-corrected data of the second image; and transmitting the encoded video data of the first image, the error-corrected data of the first image, and the error-corrected data of the second image at the transmitting device.
[0007] In another example, this disclosure describes a method for encoding video data, comprising: obtaining video data from a video source at a transmitting device; generating transform blocks based on the video data at the transmitting device; determining which of the transform blocks are anchor transform blocks at the transmitting device; calculating a correlation matrix of the set of transform blocks at the transmitting device; generating a bit-reduced non-anchor transform matrix at the transmitting device; and transmitting the anchor transform blocks, non-anchor transform blocks, and correlation matrix to a receiving device at the transmitting device.
[0008] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; a communication interface; and one or more processors implemented in a circuit and coupled to the memory, the one or more processors being configured to perform the method according to any one of claims 1-22.
[0009] In another example, this disclosure describes an apparatus for processing video data, comprising: a memory configured to store video data; and a communication interface configured to obtain error correction data from a transmitting device, wherein the error correction data provides error correction information about images of the video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: generate prediction data for images, wherein the prediction data for images includes predictions of blocks of images based at least in part on one or more previously reconstructed images of the video data; generate coded video data based on the prediction data for images, wherein the coded video data includes transform blocks, the transform blocks including transform coefficients; scale the bits of the transform coefficients of the transform blocks based on reliability values of bit positions; generate error-corrected coded video data using the error correction data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstruct images based on the error-corrected coded video data.
[0010] In another example, this disclosure describes an apparatus for processing video data, comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: acquire the video data; acquire predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated picture generated by a receiving device; adjust one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; perform a video coding process to generate coded video data based on one or more pictures of the acquired video data, wherein the video coding process is controlled by the video coding parameters; perform a channel coding process on the coded video data to generate channel-coded data, wherein the channel coding process is controlled by the channel coding parameters; and a communication interface configured to transmit the channel-coded data to the receiving device.
[0011] In another example, this disclosure describes a method for processing video data, comprising: obtaining error correction data at a receiving device and from a transmitting device, wherein the error correction data provides error correction information about a picture of the video data; generating prediction data of the picture at the receiving device, wherein the prediction data of the picture includes predictions of blocks of the picture based at least in part on one or more previously reconstructed pictures of the video data; generating coded video data at the receiving device based on the prediction data of the picture, wherein the coded video data includes transform blocks, the transform blocks including transform coefficients; scaling the bits of the transform coefficients of the transform blocks at the receiving device based on reliability values of bit positions; generating error-corrected coded video data at the receiving device using the error correction data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstructing a picture at the receiving device based on the error-corrected coded video data.
[0012] In another example, this disclosure describes a method for processing video data, comprising: acquiring video data; acquiring predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated picture generated by a receiving device; adjusting one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process to generate coded video data based on one or more pictures of the acquired video data, wherein the video coding process is controlled by the video coding parameters; performing a channel coding process on the coded video data to generate channel-coded data, wherein the channel coding process is controlled by the channel coding parameters; and transmitting the channel-coded data to a receiving device.
[0013] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: acquire a first multiview image set of video data, wherein the first multiview image set includes a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; transmit first encoded video data to a receiving device, wherein the first encoded video data is based on the first multiview image set; receive a multiview encoding cue from the receiving device; acquire a second multiview image set of video data, wherein the second multiview image set includes a third image and a fourth image, the third image being from the first viewpoint and the fourth image being from the second viewpoint; perform a multiview encoding process on the second multiview image set based on the multiview encoding cue received from the receiving device to generate second encoded video data, wherein the multiview encoding process reduces interview redundancy between the third and fourth images; and transmit the second encoded video data to the receiving device.
[0014] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: obtain first encoded video data from a transmitting device, wherein the first encoded video data is based on a first multiview image set of the video data, the first multiview image set including a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; determine a multiview encoding cue based on the first encoded video data; send the multiview encoding cue to the transmitting device; and obtain second encoded video data from the transmitting device, wherein the second encoded video data is based on a second multiview image set including a third image and a fourth image, the second encoded video data being encoded using a multiview encoding process that reduces interview redundancy between the third and fourth images based on the multiview encoding cue.
[0015] In another example, this disclosure describes a method for processing video data, comprising: obtaining a first multiview image set of video data, wherein the first multiview image set includes a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first multiview image set; receiving a multiview encoding cue from the receiving device; obtaining a second multiview image set of video data, wherein the second multiview image set includes a third image and a fourth image, the third image being from the first viewpoint and the fourth image being from the second viewpoint; performing a multiview encoding process on the second multiview image set based on the multiview encoding cue received from the receiving device to generate second encoded video data, wherein the multiview encoding process reduces interview redundancy between the third image and the fourth image; and transmitting the second encoded video data to the receiving device.
[0016] In another example, this disclosure describes a method for processing video data, comprising: obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first multiview image set of the video data, the first multiview image set including a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; determining a multiview encoding cue based on the first encoded video data; sending the multiview encoding cue to the transmitting device; and obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second multiview image set including a third image and a fourth image, the second encoded video data being encoded using a multiview encoding process that reduces interview redundancy between the third image and the fourth image based on the multiview encoding cue.
[0017] In another example, this disclosure describes an apparatus comprising: means for acquiring a first multiview image set of video data, wherein the first multiview image set includes a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; means for transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first multiview image set; means for receiving a multiview encoding cue from the receiving device; means for acquiring a second multiview image set of video data, wherein the second multiview image set includes a third image and a fourth image, the third image being from the first viewpoint and the fourth image being from the second viewpoint; means for performing a multiview encoding process on the second multiview image set based on the multiview encoding cue received from the receiving device to generate second encoded video data, wherein the multiview encoding process reduces interview redundancy between the third and fourth images; and means for transmitting the second encoded video data to the receiving device.
[0018] In another example, this disclosure describes an apparatus comprising: components for obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first multiview image set of video data, the first multiview image set including a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; components for determining multiview encoding cues based on the first encoded video data; components for sending multiview encoding cues to the transmitting device; and components for obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second multiview image set including a third image and a fourth image, the second encoded video data being encoded using a multiview encoding process that reduces interview redundancy between the third and fourth images based on the multiview encoding cues.
[0019] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: encode a first set of images of the video data to generate first encoded video data; transmit the first encoded video data to a receiving device; receive a decimation mode indication from the receiving device, the decimation mode indication indicating a decimation mode determined based on the first set of images, the decimation mode being a mode in which the encoded video data is not transmitted; encode a second set of images of the video data to generate second encoded video data; apply the decimation mode to the second encoded video data to generate decimated video data; and transmit the decimated video data to the receiving device.
[0020] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: receive first encoded video data from a transmitting device; perform a decoding process to reconstruct a first set of images based on the first encoded video data; determine a decimation mode based on the first set of images, the decimation mode indicating a mode in which the encoded video data is not transmitted; send a decimation mode indication to the transmitting device indicating the determined decimation mode; receive decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation mode has been applied, wherein the second encoded video data is generated based on a second set of images of the video data; and perform a decoding process to reconstruct a second set of images based on the second encoded video data.
[0021] In another example, this disclosure describes a method comprising: encoding a first set of images of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving a decimation mode indication from the receiving device, the decimation mode indication indicating a decimation mode determined based on the first set of images, the decimation mode being a mode in which the encoded video data is not transmitted; encoding a second set of images of video data to generate second encoded video data; applying the decimation mode to the second encoded video data to generate decimated video data; and transmitting the decimated video data to the receiving device.
[0022] In another example, this disclosure describes a method comprising: receiving first encoded video data from a transmitting device; applying a decoding process to reconstruct a first set of images based on the first encoded video data; determining a decimation mode based on the first set of images, the decimation mode indicating a mode in which the encoded video data is not transmitted; sending a decimation mode indication to the transmitting device indicating the determined decimation mode; receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation mode has been applied, wherein the second encoded video data is generated based on a second set of images of the video data; and performing a decoding process to reconstruct a second set of images based on the second encoded video data.
[0023] In another example, this disclosure describes an apparatus comprising: means for encoding a first set of images of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device a decimation mode indication indicating a decimation mode determined based on the first set of images, the decimation mode being a mode in which the encoded video data is not transmitted; means for encoding a second set of images of video data to generate second encoded video data; means for applying the decimation mode to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to the receiving device.
[0024] In another example, this disclosure describes an apparatus comprising: components for receiving first encoded video data from a transmitting device; components for performing a decoding process to reconstruct a first set of images based on the first encoded video data; components for determining a decimation mode based on the first set of images, the decimation mode indicating a mode in which the encoded video data is not transmitted; components for sending a decimation mode indication to the transmitting device indicating the determined decimation mode; components for receiving decimated video data from the transmitting device, wherein the decimated video data includes second encoded video data to which the decimation mode has been applied, wherein the second encoded video data is generated based on a second set of images of the video data; and components for performing a decoding process to reconstruct a second set of images based on the second encoded video data.
[0025] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: encode a first image of the video data to generate first encoded video data; transmit the first encoded video data to a receiving device; receive encoding selection data of a second image of the video data from the receiving device, wherein: the encoding selection data of the second image indicates encoding selections for encoding an estimate of the second image, and the second image follows the first image in decoding order; encode the second image based on the encoding selection data of the second image to generate second encoded video data; and transmit the second encoded video data to the receiving device.
[0026] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: receive first coded video data from a transmitting device; reconstruct a first image of the video data based on the first coded video data; estimate a second image of the video data based on the first image, the second image being an image that appears after the first image in decoding order; generate encoding selection data for the second image, wherein the encoding selection data for the second image indicates encoding selections for encoding the second image; transmit the encoding selection data for the second image to the transmitting device; receive second coded video data from the transmitting device; and reconstruct the second image based on the second coded video data.
[0027] In another example, this disclosure describes a method for processing video data, comprising: encoding a first image of the video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving encoding selection data of a second image of the video data from the receiving device, wherein: the encoding selection data of the second image indicates encoding selection for encoding an estimate of the second image, and the second image follows the first image in decoding order; encoding the second image based on the encoding selection data of the second image to generate second encoded video data; and transmitting the second encoded video data to the receiving device.
[0028] In another example, this disclosure describes a method for processing video data, comprising: receiving first coded video data from a transmitting device; reconstructing a first image of the video data based on the first coded video data; estimating a second image of the video data based on the first image, the second image being an image that appears after the first image in decoding order; generating encoding selection data for the second image, wherein the encoding selection data for the second image indicates encoding selections for encoding the second image; transmitting the encoding selection data for the second image to the transmitting device; receiving second coded video data from the transmitting device; and reconstructing the second image based on the second coded video data.
[0029] In another example, this disclosure describes an apparatus comprising: means for encoding a first image of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving encoding selection data of a second image of video data from the receiving device, wherein: the encoding selection data of the second image indicates encoding selections for encoding an estimate of the second image, and the second image follows the first image in decoding order; means for encoding the second image based on the encoding selection data of the second image to generate second encoded video data; and means for transmitting the second encoded video data to the receiving device.
[0030] In another example, this disclosure describes an apparatus comprising: means for receiving first coded video data from a transmitting device; means for reconstructing a first picture of the video data based on the first coded video data; means for estimating a second picture of the video data based on the first picture, the second picture being a picture that appears after the first picture in decoding order; means for generating encoding selection data for the second picture, wherein the encoding selection data for the second picture indicates encoding selections for encoding the second picture; means for transmitting the encoding selection data for the second picture to the transmitting device; means for receiving second coded video data from the transmitting device; and means for reconstructing the second picture based on the second coded video data.
[0031] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description
[0032] Figure 1 This is a block diagram illustrating an example system according to the technology of this disclosure.
[0033] Figure 2 This is a block diagram illustrating example components of a transmitting device and a receiving device according to the technology of this disclosure.
[0034] Figure 3AThis is a conceptual diagram illustrating an example channel coding process according to the technology disclosed herein.
[0035] Figure 3B This is a block diagram illustrating an example channel decoding process according to the technology of this disclosure.
[0036] Figure 4 This is a flowchart illustrating an example operation of a transmitting device according to the technology of this disclosure.
[0037] Figure 5 This is a flowchart illustrating an example operation of a receiving device according to the technology disclosed herein.
[0038] Figure 6 This is a conceptual diagram illustrating an example extraction pattern according to the technology disclosed herein.
[0039] Figure 7 This is a flowchart illustrating an example operation of a transmitting device for hybrid extraction of transform blocks according to the technology of this disclosure.
[0040] Figure 8 This is a flowchart illustrating an example operation of a receiving device for hybrid decimation of a transform block according to the technology of this disclosure.
[0041] Figure 9 This is a conceptual diagram illustrating an example extraction mode adaptively selected by a receiving device according to one or more techniques of this disclosure.
[0042] Figure 10 This is a block diagram illustrating example components of a transmitting device and a receiving device according to the technology of this disclosure.
[0043] Figure 11 A graph showing example error probabilities and corresponding absolute values of log-likelihood ratios (LLRs) for one or more techniques according to this disclosure is provided.
[0044] Figure 12 This is a flowchart illustrating an example operation of a transmitting device using scaled bits according to the technology of this disclosure.
[0045] Figure 13 This is a flowchart illustrating an example operation of a receiving device using scaled bits according to the technology of this disclosure.
[0046] Figure 14 This is a flowchart illustrating an example of data exchange between a transmitting device and a receiving device in relation to multi-view processing according to one or more techniques of this disclosure.
[0047] Figure 15 This is a flowchart illustrating an example operation of a transmitting device for multi-view processing according to the technology of this disclosure.
[0048] Figure 16 This is a flowchart illustrating an example operation of a receiving device for multi-view processing according to the technology of this disclosure.
[0049] Figure 17 This is a block diagram illustrating example components of a transmitting device and a receiving device that perform extraction of coded video data according to the technology of this disclosure.
[0050] Figure 18 This is a conceptual diagram illustrating an example exchange of information including an extraction pattern indication according to the technology of this disclosure.
[0051] Figure 19 This is a flowchart illustrating an example operation of a transmitting device according to the technology of this disclosure, wherein the transmitting device receives a decimation mode indication.
[0052] Figure 20 This is a flowchart illustrating an example operation of a receiving device according to the technology of this disclosure, wherein the receiving device sends a decimation mode indication.
[0053] Figure 21 This is a block diagram illustrating a transmitting device according to the technology of this disclosure and an example component of a receiving device that transmits encoded selection data to the transmitting device.
[0054] Figure 22 This is a communication diagram illustrating an example of data exchange between a transmitting device and a receiving device, including the transmission and reception of encoded selection data according to the technology of this disclosure.
[0055] Figure 23 This is a flowchart illustrating an example operation of a transmitting device according to the technology of this disclosure, wherein the transmitting device receives encoded selection data.
[0056] Figure 24 This is a flowchart illustrating an example operation of a receiving device according to the technology of this disclosure, wherein the receiving device transmits encoded selection data.
[0057] Figure 25 This is a conceptual diagram illustrating an example hierarchical structure of encoded video data according to the technology of this disclosure.
[0058] Figure 26 This is a block diagram illustrating alternative example components of a transmitting device according to one or more technologies of this disclosure.
[0059] Figure 27 This is a block diagram illustrating alternative example components of a receiving device according to one or more technologies of this disclosure. Detailed Implementation
[0060] While modern video coding processes can significantly reduce the amount of data required to represent video data, these processes are typically resource-intensive and can involve numerous memory operations. Therefore, modern video coding processes may require sophisticated processors, fast memory, and consume considerable power. However, for some contemporary and planned wireless communication systems, such as 5G and 6G wireless communication systems, wireless transmission bandwidth may be less constrained, especially when communicating over short distances (such as the distance between devices on a person).
[0061] This disclosure describes techniques for reducing the complexity of video coding at a transmitting device using error correction, which is performed as part of channel decoding using error-corrected data. The transmitting device can perform a finite video coding process that generates coded video data. Finite video coding processes typically use relatively low-resource-intensive decoding tools, such as intra-frame prediction. Because the video coding process uses less complex decoding tools, the resulting coded video data can be larger than video data encoded using more complex and resource-intensive decoding tools. Error-corrected data is based on the coded video data. The transmitting device can send the error-corrected data to a receiving device. The transmitting device may not need to send all the coded video data for one or more frames to the receiving device.
[0062] The receiving device can estimate the images of the video data based on one or more previously reconstructed images. In some examples, to estimate the images, the receiving device can extrapolate the contents of blocks from previously reconstructed images. The receiving device can then perform a full video coding process on the estimated images to generate estimated coded video data for those images. When performing a full video coding process, the receiving device can use more complex decoding tools, such as inter-frame prediction, than the limited video coding process performed by the transmitting device. The receiving device can perform a channel decoding process that generates error-corrected coded video data based on the estimated coded video data of the images and the error-correcting data of the images. In some cases, the channel decoding process can generate error-corrected coded video data based on the error-correcting data of the images and a combination of the estimated coded video data and the coded video data of the images transmitted by the transmitting device. The receiving device can reconstruct the images based on the error-corrected coded video data. In this way, the receiving device is able to reconstruct each image of the video data even if the transmitting device has not transmitted all the coded video data of the images.
[0063] As further described in this disclosure, various techniques, such as the application of decimation modes, can be applied to specify which transform blocks of lightly coded video data are not signaled or have a reduced bit depth. Furthermore, in some examples of this disclosure, reliability values can be determined for bit positions, and these reliability values can be used to scale the bits of the transform coefficients of the transform block, and the scaled values can be used in channel coding and channel decoding.
[0064] As further described in this disclosure, the receiving device can determine a decimation mode based on a first set of images. The decimation mode is a mode in which the encoded video data is not transmitted. The receiving device can send a decimation mode indication, indicating the determined decimation mode, to the transmitting device. The transmitting device can receive the decimation mode indication from the receiving device and apply the indicated decimation mode to the encoded video data to generate decimated video data. The transmitting device can then send the decimated video data to the receiving device. In this way, the technology of this disclosure can further reduce resource consumption at the transmitting device while still avoiding the transmission of excessive data. This can further improve decoding efficiency.
[0065] Figure 1 This is a block diagram illustrating an example system 100 according to the technology of this disclosure. Figure 1 In this example, system 100 includes a transmitting device 102, a receiving device 104, and a base station 106. The transmitting device 102 can be a device configured to perform actions, including extended reality (XR) devices (e.g., XR headsets), mobile devices, wearable devices, sensor devices, Internet of Things (IoT) devices, intermediate networking devices, or other types of devices. In some examples, the transmitting device 102 may be included in a robot or vehicle. The receiving device 104 can be a computing device, such as a mobile device (e.g., a mobile phone or tablet), a personal computer, a vehicle-based computing device, a wireless base station, a wearable computing device, an intermediate networking device, a dedicated device, an Internet of Things (IoT) device, or other types of devices. In some examples, the receiving device 104 is a device that a user of the transmitting device 102 may have in addition to the transmitting device 102.
[0066] Transmitting device 102 and receiving device 104 can communicate with base station 106. In some examples, transmitting device 102 and receiving device 104 can communicate with base station 106 using fifth-generation (5G) wireless communication protocols, sixth-generation (6G) wireless communication protocols, WiFi protocols, Bluetooth protocols, or another type of wireless communication protocol. Base station 106 can transmit data from network 115 to transmitting device 102 and receiving device 104 via wireless downlink channels 108A and 108B (collectively referred to as "wireless downlink channel 108"). Base station 106 can receive data from transmitting device 102 and receiving device 104 for transmission to other devices connected to network 115 via wireless uplink channels 110A and 110B (collectively referred to as "wireless uplink channel 110"). Transmitting device 102 and receiving device 104 can communicate directly with each other via wireless sidelink channel 112. In other examples, transmitting device 102 and receiving device 104 can communicate via other types of channels. In other examples, the transmitting device 102 and the receiving device 104 may communicate via other types of channels.
[0067] exist Figure 1 In the example, transmitting device 102 includes one or more processors 114, memory 116, communication interface 118, video source 120, and display system 122. Receiving device 104 includes one or more processors 130, memory 132, and communication interface 134. Processors 114 and 130 may include circuitry configured to perform various information processing tasks, including the execution of computer-readable instructions. Processors 114 and 130 may include microprocessors, digital signal processors, and other types of circuitry. Memory 116 and 132 may be configured to store data, such as computer-readable instructions, video data, and other types of data. Communication interfaces 118 and 134 may be configured to transmit and receive data, for example, via wireless downlink channel 108, wireless uplink channel 110, and wireless sidelink channel 112.
[0068] Typically, video source 120 refers to a source of video data (e.g., raw, unencoded video data). Video source 120 may include one or more video capture devices, such as cameras, video archives containing previously captured raw video, and / or video feed interfaces that receive video from video content providers. As a further alternative, video source 120 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video.
[0069] In the example where the transmitting device 102 is an XR device that presents MR and AR images to a user, it may be necessary to analyze video data from video source 120 so that the display system 122 of the transmitting device 102 can display virtual elements in the correct positions. Processing video data in this way may require significant computing resources. In other words, powerful processors and a large amount of energy may be used when processing video data. Because the transmitting device 102 may be designed to be worn on the user's head, minimizing the weight and power consumption of the transmitting device 102 while supporting high-quality, low-latency video may be important.
[0070] Furthermore, in some examples, the transmitting device 102 is an XR headset, and the transmitting device 102 can be configured to process images of video data to generate virtual element data. The receiving device 104 can be configured to transmit (and the transmitting device 102 is configured to receive) the virtual element data. The transmitting device 102 may include a display system 122 configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0071] Therefore, it may be desirable to offload video data processing to a device other than transmitting device 102, such as receiving device 104. Receiving device 104 may have significantly more resources than transmitting device 102, either permanently or temporarily. For example, receiving device 104 may be equipped with a larger battery and a relatively powerful processor. However, for receiving device 104 to process video data, transmitting device 102 may need to transmit video data to receiving device 104 via wireless sidelink channel 112. Because a very large number of bits may be required to represent unencoded high-quality video data, transmitting unencoded high-quality video data from transmitting device 102 to receiving device 104 will consume a significant amount of time and energy. The transmission time may undermine the goal of providing low-latency video to the user. The energy required for transmission may undermine the goal of minimizing power consumption. Encoding video data using video decoding specifications such as H.264 / Advanced Video Decoding (AVC), H.265 / High-Efficiency Video Decoding (HEVC), or H.266 / Various Video Decoding (VVC) can significantly reduce the amount of data required to represent video data. However, the encoding process itself may introduce its own latency and power consumption requirements.
[0072] This disclosure describes techniques that can solve these problems. According to the techniques of this disclosure, transmitting device 102 and receiving device 104 can use a distributed video decoding (DVC) process. The DVC process reduces the amount of encoding work performed by transmitting device 102 and offloads some encoding work to receiving device 104. Receiving device 104 may have more resources (e.g., computing power, access power, etc.) than transmitting device 102 and is therefore better equipped to perform encoding work. In some examples, the DVC process can be used for load balancing computational tasks between devices. For example, the system may determine that receiving device 104 is generally more efficient than transmitting device 102 in performing a specific video-related computational task.
[0073] In addition to the video encoding process, transmitting device 102 may also perform a channel coding process to prepare encoded video data for transmission to receiving device 104. The channel coding process can generate error correction data for the data sequence within the encoded video data. Typically, receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, according to the techniques of this disclosure, transmitting device 102 may transmit error correction data for some encoded video data, but not the encoded video data corresponding to the error correction data. Receiving device 104 may estimate one or more subsequent frames. Receiving device 104 may perform a video encoding process on the subsequent frames to generate estimated encoded video data. The receiving device may use the estimated encoded video data and the received error correction data to generate error-corrected encoded video data. Receiving device 104 may then decode the error-corrected encoded video data to reconstruct the video data that transmitting device 102 did not transmit.
[0074] Therefore, in some examples, receiving device 104 can obtain first coded video data and first error correction data from transmitting device 102. The first coded video data may represent one or more blocks of a first image of video data. The first error correction data can provide error correction information about the blocks of the first image. Receiving device 104 can use the first error correction data to perform an error correction operation on the first coded video data to generate first error-corrected coded video data. Additionally, receiving device 104 can perform a first reconstruction operation that reconstructs the blocks of the first image based on the first coded video data. The first reconstruction operation can be controlled by the values of one or more parameters.
[0075] Furthermore, receiving device 104 can obtain second error correction data from transmitting device 102. The second error correction data can provide error correction information about one or more blocks of a second image of video data. Receiving device 104 can generate prediction data for the second image. The prediction data for the second image can include predictions of blocks of the second image of video data based at least in part on blocks of one or more previously reconstructed images, such as a first image. Receiving device 104 can use one or more decoding tools to generate prediction data for encoded video data not used to generate the second image. Receiving device 104 can generate second encoded video data based on the predictions of blocks of the second image. Receiving device 104 can use the second error correction data to perform error correction operations on the second encoded video data to generate second error-corrected encoded video data. Receiving device 104 can perform a second reconstruction operation, which reconstructs blocks of the second image based on the second error-corrected encoded video data. The second reconstruction operation is controlled by the values of parameters.
[0076] Furthermore, according to one or more techniques of this disclosure, receiving device 104 can receive a decimation mode indication. Receiving device 104 can determine the decimation mode indication based on a previously reconstructed image. The decimation mode indication can indicate a mode in which encoded video data is not transmitted. For example, the decimation mode can indicate a mode in which encoded video data for the entire image is skipped. In some examples, the decimation mode indicates a mode in which encoded video data for a specific region within the image is skipped. In some examples where the video data is multi-view video data, the decimation mode can indicate a mode in which encoded video data for images from a specific viewpoint is skipped.
[0077] Transmitting device 102 can perform a video encoding process on images of video data. This video encoding process can compress images less than "heavy" or more complex compression operations described in H.264, H.265, and H.266 video decoding standards. In addition to the video encoding process, transmitting device 102 can also perform a channel coding process to prepare encoded video data for transmission to receiving device 104. The channel coding process can generate error correction data for the data sequence within the encoded video data. Typically, receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, receiving device 104 can also use the error correction data to recover information intentionally not transmitted to the receiving device. Therefore, transmitting device 102 can apply a decimation mode to second encoded video data to generate decimated video data. Transmitting device 102 can send error correction data (generated based on undecimated encoded video data) and decimated video data to receiving device 104.
[0078] Receiving device 104 can obtain first coded video data and first error correction data from transmitting device 102. The first coded data can represent one or more blocks of a first image of the video data. The first error correction data can provide error correction information about the blocks of the first image. Receiving device 104 can use the first error correction data to perform an error correction operation on the first coded video data to generate first error-corrected coded video data. Additionally, receiving device 104 can perform a first reconstruction operation, which reconstructs blocks of the first image based on the first coded video data. The first reconstruction operation can be controlled by the values of one or more parameters.
[0079] Furthermore, receiving device 104 can obtain first error-corrected data and first encoded video data from transmitting device 102. Receiving device 104 can apply an error correction process to modify the first encoded video data based on the first error-corrected data to generate first error-corrected encoded video data. Receiving device 104 can also apply a decoding process to reconstruct a first set of images based on the first error-corrected encoded video data. Receiving device 104 can determine a decimation mode based on the first set of images, indicating a mode in which encoded video data has not been transmitted. Receiving device 104 can send a decimation mode indication to transmitting device 102, indicating the determined decimation mode. Receiving device 104 can receive second error-corrected data and decimated video data from transmitting device 102. Decimated video data may include second encoded video data for which a decimation mode has been applied. The second encoded video data is generated based on a second set of images of video data. Receiving device 104 can apply an error correction process to modify the second encoded video data based on the second error-corrected data to generate second error-corrected encoded video data. Receiving device 104 can apply a decoding process to reconstruct a second set of images based on the second error-corrected encoded video data.
[0080] Figure 2 This is a block diagram illustrating example components of a transmitting and receiving device according to the technology of this disclosure. System 200 includes a transmitting device 102 and a receiving device 104. The transmitting device 102 is configured to transmit encoded video data to the receiving device 104. Figure 2 In one example, transmitting device 102 includes a video encoder 210, a channel encoder 212, and a punching unit 214. Receiving device 104 includes a de-punching unit 220, a channel decoder 222, a video decoder 224, an image estimation unit 226, and a video encoder 228. In other examples, transmitting device 102 and receiving device 104 may include more, fewer, or different units. The processor 114 of transmitting device 102 ( Figure 1The receiving device 104's processor 130 can implement a video encoder 210, a channel encoder 212, and a punching unit 214. The receiving device 104's processor 130 can implement a de-punching unit 220, a channel decoder 222, a video decoder 224, an image estimation unit 226, and a video encoder 228. Communication interface 118 ( Figure 1 () can represent sending and receiving data on behalf of sending device 102. Communication interface 134 ( Figure 1 () can represent receiving device 104 sending and receiving data.
[0081] The video encoder 210 of the transmitting device 102 can transmit from a video source (e.g., video source 120). Figure 1 The transmitting device 102 receives video data. The video data may include, for example, raw, unencoded video images from video source 120. In some examples, the transmitting device 102's memory (e.g., memory 116) receives video data. Figure 1 The video encoder 210 can store video data. It can perform a video encoding process on the video data to generate encoded video data. The video encoding process can be "limited" because it may be relatively fast and consume fewer resources compared to more robust video compression processes such as H.264 / AVC, H.265 / HEVC, or H.266 / VVC. The video encoding process may not reduce the number of bits representing the video data to the same extent as a more robust or complete video encoding process.
[0082] The video encoder 210 can perform a limited video encoding process in one of a variety of ways. For example, in some examples, the video encoder 210 can perform a prediction process (such as intra-frame prediction) on each frame of video data to produce prediction data. The video encoder 210 can generate residual data based on the prediction data. For example, the video encoder 210 can subtract the samples of the prediction data from the corresponding samples of the original frame to determine the samples of the residual data. The samples can be values indicating color values (such as Y, Cb, or Cr values in the YCbCr color gamut, or red, green, or blue values in the RGB color gamut).
[0083] Video encoder 210 may apply a transform, such as Discrete Cosine Transform (DCT), to the residual data to produce a transform block including transform coefficients. Additionally, video encoder 210 may quantize the transform coefficients. Video encoder 210 may apply entropy coding (such as Context Adaptive Binary Arithmetic Decoding (CABAC) coding or Exponential Golomb-Rice decoding) to the syntax elements representing the quantized transform coefficients. Encoded video data may include entropy-coded syntax elements. In some examples, video encoder 210 applies the transform and / or quantization directly to the video data without first using intra-frame prediction. In some examples where video encoder 210 does not apply entropy coding, the encoded video data includes syntax elements representing quantized transform coefficients, unquantized transform coefficients, or residual data.
[0084] In the example where the video encoder 210 does not use inter-picture prediction, fewer memory read requests may be required compared to a more robust video compression process that might need to read data about previously decoded pictures from memory. Such memory read requests can be relatively time- and energy-intensive.
[0085] In some examples where the video data is multi-view video data, the video encoder 210 can perform multi-view video coding to generate prediction data. For example, the video encoder 210 can use inter-view prediction to generate prediction data for blocks (e.g., macroblocks, decoding units, etc.) of non-anchor images. In some cases, inter-view prediction may include determining a disparity vector for a block, which indicates the lateral displacement between the block and a corresponding block in the images of one or more reference viewpoints.
[0086] The channel encoder 212 of the transmitting device 102 can apply a channel coding process to encode video data. The channel coding process prepares the encoded video data for transmission over a wireless communication channel (such as channel 230). Channel 230 can be a wireless sidelink channel 112 (…). Figure 1 The channel can be another communication channel. Channel-coded video data may include error correction data. Channel encoder 212 can generate error correction data in various ways. For example, channel encoder 212 can generate error correction data as convolutional codes or turbo codes. Error correction data can help receiving device 104 determine whether the received coded video data has changed during transmission via channel 230, and can help receiving device 104 correct such changes. A more detailed discussion of channel coding and channel decoding is provided below with reference to FIG3.
[0087] In addition, Figure 2In the example, the puncturing unit 214 of the transmitting device 202 can apply a bit puncturing process to the error-correcting data to generate bit-punctured error-correcting data. The bit puncturing process can reduce the number of bits in the error-correcting data. For example, the puncturing unit 214 can perform an operation to remove bits from the error-correcting data according to a puncturing pattern.
[0088] Transmitting device 102 can transmit data, such as encoded video data and error correction data (e.g., bit-puncturing error correction data), to receiving device 104 via channel 230. Channel 230 may introduce noise into the transmitted data. In some examples, channel 230 is a multipath channel, and the data transmitted in channel 230 may be time-varying. Receiving device 104 can receive the noise-modified data. Receiving device 104 can store the noise-modified data, at least temporarily, in a memory, such as memory 132. Figure 1 ).
[0089] The de-puncturing unit 220 can perform a de-puncturing operation on the received bit-punctured error-corrected data to reconstruct the error-corrected data. The de-puncturing operation can replace the punctured symbols with neutral values according to the puncturing pattern indication. The de-puncturing operation can generate erase bits that indicate the presence of a neutral symbol in the error-corrected data.
[0090] Channel decoder 222 can apply a channel decoding process to generate error-corrected coded video data based on error-corrected data and coded video data (such as coded video data received from the transmitting device and / or coded video data generated by the receiving device 104). For example, channel decoder 222 can modify the value of bit-coded video data according to any of a variety of error correction schemes (such as low-density parity-check (LDPC) decoding or forward error correction (FEC)).
[0091] Video decoder 224 can perform a video decoding process to reconstruct images based on error-corrected coded video data. For example, video decoder 224 can apply an entropy decoding process to the bits of error-corrected coded video data to obtain quantization transform coefficients. Video decoder 224 can apply an inverse quantization operation to the quantization transform coefficients, apply an inverse transform to the inverse quantization transform coefficients to generate residual data, generate prediction data, and use the prediction data and residual data to reconstruct images from the video data. Video decoder 224 can generate prediction data in the same manner as video encoder 210.
[0092] Image estimation unit 226 can generate an estimate of the next image from the video data. For example, image estimation unit 226 can extrapolate the next image from two or more previously reconstructed images. For example, in this example, image estimation unit 226 can divide a first previously reconstructed image into blocks. For each block of the first previously reconstructed image, image estimation unit 226 can determine one or more corresponding blocks for blocks in one or more additional previously reconstructed images. The corresponding block of the block can be the best available match for that block. Image estimation unit 226 can generate a prediction of the block based on one or more corresponding blocks of the block. Image estimation unit 226 can use one-way prediction or two-way prediction to generate the prediction. Therefore, by generating a prediction for each block of the next image, image estimation unit 226 can generate an estimate of the next image. In some examples, image estimation unit 226 generates the next image by applying global motion to the previously reconstructed images.
[0093] In some examples, the image estimation unit 226 can re-encode the current image that has already been decoded by the video decoder 224. The next image of the video data can be the image that follows the image just decoded by the video decoder 224 in the decoding order. In this example, the image estimation unit 226 can perform intra-frame prediction or inter-frame prediction on blocks of the current image. When performing inter-frame prediction on blocks, the image estimation unit 226 can determine one or more motion vectors of the block. For example, the image estimation unit 226 can determine that a particular block of the current image has a motion vector with an amplitude m relative to a reference block in a reference image at a picture order count (POC) distance from the current image p1. In this example, the current image and the next image can have a POC distance of p2. The image estimation unit 226 can determine a scaling factor s as p2 / p1. The image estimation unit 226 can then scale the motion vector of the particular block by s (e.g., s*m). The image estimation unit 226 can determine the position in the next image indicated by the scaled motion vector and set the sample at the determined position as the sample of the particular block of the current image. Image estimation unit 226 can repeat this process for each inter-frame prediction block of the current image.
[0094] In some examples, the image estimation unit 226 may apply one or more filters to the prediction data. For example, the image estimation unit 226 may apply one or more deblocking filters, smoothing filters, adaptive loop filters, or other types of filters to the prediction data.
[0095] Video encoder 228 can perform the same finite video coding process as video encoder 210 on the video data generated by image estimation unit 226. For example, video encoder 228 can perform intra-frame prediction to generate prediction data. Video encoder 228 can use the prediction data and corresponding blocks of video data generated by image estimation unit 226 to generate residual data. Video encoder 228 can apply transforms (e.g., DCT transform, DST transform, etc.) to the residual data to produce transform coefficients. Video encoder 228 can apply quantization to the transform coefficients. In addition, video encoder 228 can apply entropy coding to the syntax elements representing the transform coefficients.
[0096] As briefly described above, the channel decoder 222 can apply the channel decoding process to channel-coded video data. Figure 3A and Figure 3B More information is provided about the channel encoding process performed by the channel encoder 212 and the channel decoding process performed by the channel decoder 222.
[0097] Specifically, Figure 3A This is a block diagram illustrating an example channel coding process according to the technology of this disclosure. For each frame of encoded video data, the channel encoder 212 of the transmitting device 102 may apply systematic decoding operations (such as low-density parity-check (LDPC) decoding operations) to the systematic bits of the frame to generate error-correcting data for the frame. The systematic bits of the frame may include the encoded video data of the frame generated by the video encoder 210.
[0098] exist Figure 3A In the example, error correction data is labeled "error correction bits." For image n, channel encoder 212 can generate error correction data 300A based on systematic bit 302A. Similarly, for image n+1, channel encoder 212 can generate error correction data 300B based on systematic bit 302B. Channel encoder 212 can classify images of video data into anchor video images and non-anchor images. Channel encoder 212 can classify images such that anchor images appear periodically in the images. In some examples, if channel decoder 222 (e.g., from receiving device 104) receives an indication that an error exists in an image, channel encoder 212 can classify the image as an anchor image. For each anchor image, transmitting device 102 can transmit the encoded anchor image and the error correction data for the anchor image. However, for non-anchor images, transmitting device 102 can only transmit the error correction data for the non-anchor images.
[0099] For example, in Figure 3AIn the example, image n can be an anchor image, while image n+1 is a non-anchor image. Therefore, transmitting device 102 can transmit the systematic bit 302A of image n, the error correction data 300A of image n, and the error correction data 300B of image n+1, but does not transmit the systematic bit 302B of image n+1.
[0100] Figure 3B This is a block diagram illustrating an example channel decoding process according to the technology of this disclosure. As described above, channel decoder 222 can perform a channel decoding process on encoded video data to reconstruct the encoded video data. When processing anchor images (e.g., image n), channel decoder 222 can obtain the systematic bits of the anchor images (in...) from de-puncturing unit 220. Figure 3B The Chinese character is represented as SI. n ) and anchor image error correction data (in Figure 3B The middle is represented as y n The systematic bits of the anchor image can represent the encoded video data of the anchor image. The channel decoder 222 can use the error correction data of the anchor image to detect and / or correct errors in the systematic bits of the anchor image. The video decoder 224 can reconstruct the anchor image using the obtained error-corrected encoded video data for the anchor image. The receiving device 104 can store the reconstructed images, including the reconstructed anchor image and the reconstructed non-anchor image, in the decoded image buffer 350.
[0101] When processing non-anchor images, the channel decoder 222 can obtain the systematic bits (denoted as SI) representing the encoded video data of the non-anchor images. n+1 The video encoder 228 of the receiving device 104 can generate coded video data of the non-anchor image based on the video data generated by the image estimation unit 226. The channel decoder 222 can obtain error correction data of the non-anchor image from the de-puncturing unit 220 (in...). Figure 3B The middle is represented as y n+1 Channel decoder 222 can then perform the same channel decoding process as when processing the anchor image. Therefore, channel decoder 222 can use error correction data from the non-anchor image to detect and / or correct "errors" in the systematic bits of the non-anchor image. However, "errors" in the systematic bits of the non-anchor image cannot be attributed to noise in channel 230 (the same applies to errors in the systematic bits of the anchor image). Instead, "errors" in the systematic bits of the non-anchor image may be due to differences between the predicted version of the non-anchor image and the original version of the non-anchor image. Therefore, channel decoder 222 can use the transmitted error correction data for the non-anchor image as a mechanism for "correcting" prediction errors.
[0102] The video encoder 210 of transmitting device 102, the video decoder 224 of receiving device 104, and the video encoder 228 of receiving device 104 can perform video encoding and video decoding processes based on the values of one or more sets of parameters. In other words, the values of the parameters can control various aspects of the video encoding process performed by the video encoder 210, video encoder 228, and video decoder 224. In some examples, the parameters may include one or more of the following:
[0103] • Parameters indicating the color space (e.g., red-green-blue, Y-Cb-Cr, etc.)
[0104] Pixel extraction parameters
[0105] Parameters indicating DCT size
[0106] • Transmitted DCT coefficients
[0107] • A parameter indicating the number of bits for each DCT coefficient
[0108] • Quantization parameters (e.g., parameters indicating quantization schemes such as linear, Max Lloyd, etc.)
[0109] Each of the video encoder 210, video decoder 224, and video encoder 228 may need to use the same parameter values. Therefore, according to one or more techniques of this disclosure, the transmitting device 102 can send parameter values to the receiving device 104. The receiving device 104 can receive the sent parameter values. The video decoder 224 and video encoder 228 can use the parameter values during video decoding and video encoding processes.
[0110] In some examples, the transmitted parameter values are static or semi-static. For instance, in an example where the transmitted parameter values are static, the transmitting device 102 may transmit the parameter value once to the receiving device 104, and the receiving device 104 may use the parameter value to operate within an indeterminate time period. In an example where the transmitted parameter values are semi-static, the transmitting device 102 may occasionally update the parameter values and retransmit the updated parameter values to the receiving device 104.
[0111] Transmitting device 102 can transmit parameter values in one of a variety of ways. For example, in some examples, transmitting device 102 can use uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, or another type of message to transmit parameter values to receiving device 104.
[0112] In the example where transmitting device 102 sends parameter values to receiving device 104, video encoder 210 can divide each color component (e.g., R, G, and B components; Y, Cb, and Cr components) of the video data image into uniformly sized (M×M) blocks. Examples of such blocks can include macroblocks (MBs) and maximum decoding units (LCUs). The video encoder 210 of transmitting device 102 can compute a transform (e.g., 2D-DCT) for each block, thereby producing M... 2 There are N transform coefficients. The video encoder 210 can assign a sorting order to the transform coefficients of a block. For example, the video encoder 210 can sort the transform coefficients of a block according to a zigzag scan order, starting with the most important transform coefficient (e.g., the lowest frequency) and ending with the least important transform coefficient (e.g., the highest frequency). The video encoder 210 can select the first N c There are N transformation coefficients, where N c This is a parameter value that indicates the number of transform coefficients transmitted. The video encoder 210 can discard unselected transform coefficients.
[0113] Additionally, the video encoder 210 can quantize the selected transform coefficients. For example, the selected transform coefficients can have a range from 0 to N. c In the case of index i = -1, the parameter can include a bit width parameter corresponding to different index values (e.g., Bi, i = 0, 1, ..., N). c -1). For the selected transformation coefficients d i For each of the selected transform coefficients d, the video encoder 210 can quantize the selected transform coefficients d using the following equation. i .
[0114] c i = round ( α d i B i (1)
[0115] In the equation above, c i The transformation coefficient d i The quantized version, α is the scaling constant, B i `i` is the bit width parameter for index `i`, and `round` is the function for rounding to the nearest integer. Therefore, in the example where B0 is 8, B1 is 4, and B2 is 4, the quantization transform coefficients could be, for example, c0 = 00100011, c1 = 0110, c2 = 1001, and so on.
[0116] The video encoder 228 of the receiving device 104 can generate predicted data for an image, generate residual data based on the predicted data, and apply one or more transforms to the residual data to generate a transform block including transform coefficients. The video encoder 228 may need to use the same bit width parameters as the video encoder 210 so that the channel decoder 222 can correctly associate specific systematic bits with the corresponding error correction data received from the de-puncturing unit 220.
[0117] In some examples of this disclosure, receiving device 104 can determine the values of one or more parameters without requiring transmitting device 102 to send those values to receiving device 104. Examples where receiving device 104 determines the values of one or more parameters without transmitting device 102 to receiving device 104 can achieve a better trade-off with lower control signaling overhead. For example, receiving device 104 can determine the values of one or more parameters without transmitting device 102 to receiving device 104. c Determine the number of DCT coefficients (N) given the value of N. c For example, in this example, receiving device 104 can determine N to achieve the desired peak signal-to-noise ratio (P-SNR). c The minimum value. In other words, the receiving device 104 can achieve the desired P-SNR of N. c The minimum value is determined as follows:
[0118] (2)
[0119] In the equation above, c i The transform coefficients (e.g., DCT coefficients) with index i, SNR d It is the expected P-SNR, M 2 -1 represents the maximum number of transform coefficients. In some examples, the receiving device 104 can make robust predictions based on the number of transform coefficients (N) for each image (or other fragment). c An evaluation should be performed once. Periodic reset of parameter values can be applied to prevent error propagation.
[0120] In another example, receiving device 104 can determine the number of quantized bits based on the predicted image, rather than the number of quantized bits received from transmitting device 102. For example, in this example, receiving device 104 can calculate the probability distribution of quantized and unquantized coefficients:
[0121] (3)
[0122] In the equation above, It is the probability distribution of the quantization coefficients, p i It is the probability distribution of the unquantized coefficients, c iThe transformation coefficient d i The quantified version.
[0123] The receiving device 104 can determine the number of quantized bits B based on the entropy of the quantized and unquantized coefficients as follows: i :
[0124] (4)
[0125] Therefore, the receiving device 104 can evaluate B for each picture (or fragment) based on the prediction of the picture generated by the picture estimation unit 226. i once.
[0126] Figure 4 This is a flowchart illustrating an example operation of the transmitting device 102 according to the technology of this disclosure. Figure 4 In the example, the video encoder 210 of the transmitting device 102 can obtain video data (400). For example, the video encoder 210 can obtain video data from the video source 120. Furthermore, the video encoder 210 can perform video encoding on the video data to generate encoded video data (402). For example, the video encoder 210 can apply intra-frame prediction to generate prediction data, generate residual data based on the prediction data and the original video data, and apply a transform (e.g., DCT) to blocks of the residual data to generate transform blocks. The video encoder 210 can quantize the transform coefficients of the transform blocks. Additionally, in some examples, the video encoder 210 can apply entropy encoding to syntax elements representing the quantized transform coefficients. In some examples, the video encoder 210 can implement a reconstruction loop that can apply entropy decoding, inverse quantization, and one or more inverse transforms to reconstruct the residual data. The video encoder 210 can apply the prediction data and reconstruct the residual data to reconstruct the video data. In some examples, the video encoder 210 applies one or more filters to reconstruct the video data, such as deblocking filters, adaptive loop filters, sampling adaptive offset filters, etc. The video encoder 210 can use reconstructed video data as reference data for intra-frame prediction.
[0127] The video encoder 210 can perform the video encoding process based on the values of one or more parameters. For example, the video encoder 210 quantizes the transform coefficients according to specific quantization parameters, using a specific color space, etc.
[0128] The channel encoder 212 of transmitting device 102 can perform channel coding on the encoded video data to generate error correction data (404). Transmitting device 102 can, for example, transmit the encoded video data and error correction data to receiving device 104 via channel 230 (406). In some examples, transmitting device 102 can selectively transmit a portion of the encoded video data and transmit other portions of the encoded video data. For example, transmitting device 102 may transmit encoded video data of some images but not encoded video data of other images. In another example, transmitting device 102 may transmit encoded video data of some transform blocks of images but not uncoded video data of other transform blocks of images. In some examples, transmitting device 102 may transmit a certain number of the most significant bits of the transform coefficients but not the less significant bits of the transform coefficients.
[0129] In some examples, transmitting device 102 may also send the values of one or more parameters to receiving device 104. The values of these parameters can control how receiving device 104 reconstructs the video data. For example, the parameters may include a transform size parameter, which indicates the size of the transform blocks in the encoded video data generated by video encoder 210. In this example, receiving device 104 may need to interpret the received encoded video data based on the same transform block size in order to correctly reconstruct the video data. In other examples, the parameters may include parameters indicating the number of transform coefficients, bit width parameters, etc.
[0130] Therefore, in Figure 4 In the example, transmitting device 102 can obtain video data from a video source. Transmitting device 102 can generate encoded video data of a first image and encoded video data of a second image based on a set of parameters. Transmitting device 102 can perform channel coding on the encoded video data of the first and second images to generate error-corrected data for the first and second images. Transmitting device 102 can transmit the encoded video data of the first image, the error-corrected data of the first image, and the error-corrected data of the second image. In some examples, transmitting device 102 can transmit the values of the parameters to receiving device 104.
[0131] Figure 5 This is a flowchart illustrating an example operation of the receiving device 104 according to the technology of this disclosure. Figure 5 In the example, receiving device 104 can obtain first coded video data and first error correction data (500) from transmitting device 102. The first coded data represents one or more blocks of a first image of the video data. The first error correction data can provide error correction information about the blocks of the first image.
[0132] Receiver 104 can use the first error-correcting data to generate first error-corrected coded video data to perform error correction operations on the first coded video data (502). For example, channel decoder 222 of receiver 104 can use the first error-correcting data to perform low-density parity-check (LDPC) decoding on the first coded video data. In other examples, channel decoder 222 can use the error-correcting data in other error correction algorithms, such as forward error correction (FEC) or turbo decoding. Performing error correction operations on the first coded video data can remove errors introduced by noise in channel 230.
[0133] The video decoder 224 of the receiving device 104 can perform a first reconstruction operation to reconstruct blocks (504) of a first image based on first error-corrected coded video data. The first reconstruction operation is controlled by the values of one or more parameters. For example, the video decoder 224 can perform an inverse transform on the transformed blocks of the first error-corrected coded video data to obtain residual data. Furthermore, in this example, the video decoder 224 can, for example, use intra-frame prediction to generate prediction data. In this example, the video decoder 224 can use the prediction data and the residual data to reconstruct blocks of the first image.
[0134] In some examples, video encoders 210 and 228 can use quantization parameters to generate coded video data to quantize transform coefficients generated from image-based prediction data. When performing a reconstruction operation, video decoder 224 can use the quantization parameters to inversely quantize the transform coefficients of the error-correcting coded video data. In some examples, transmitting device 102 and / or receiving device 104 can calculate the quantization parameters based on the entropy ratio of quantized transform coefficients to unquantized transform coefficients, for example, as described above.
[0135] In some examples, the parameters include a transform size parameter. As part of generating encoded video data, video encoders 210 and 228 may apply a forward transform with a transform size indicated by the transform size parameter to the sampled domain data of the image (e.g., predicted sampled data or residual data). As part of performing a reconstruction operation, video decoder 224 may apply an inverse transform with a transform size indicated by the transform size parameter to the transform coefficients of the error-correcting encoded video data.
[0136] In some examples, the parameters include a parameter indicating the number of transform coefficients. As part of generating coded video data, video encoders 210 and 228 may include a set of transform coefficients in the coded video data, which includes a set of transform coefficients indicating the number. When performing a reconstruction operation, video decoder 224 can parse the set of transform coefficients including the indicated number of transform coefficients from the error-corrected coded video data. Furthermore, in some examples, receiving device 104 may receive coded video data and error-corrected data from transmitting device via a communication channel. For example, as described above, receiving device 104 may apply an optimization process that determines the number of transform coefficients based on the signal-to-noise ratio of the data transmitted over the communication channel.
[0137] In some examples, the parameters include a bit width parameter for multiple index values. For each corresponding index value among the multiple index values, performing the reconstruction operation may include parsing a first bit set from the error-corrected coded video data. The first bit set may indicate transform coefficients with corresponding index values, and the number of bits in the first bit set is equal to the bit width indicated by the bit width parameter of the corresponding index value. As part of generating coded video data, video encoders 210 and 228 may include a second bit set in the coded video data. The second bit set may indicate transform coefficients with corresponding index values, and the number of bits in the second bit set is equal to the bit width indicated by the bit width parameter of the corresponding index value. Video decoder 224 may parse a third bit set from the error-corrected coded video data. The third bit set may indicate transform coefficients with corresponding index values, and the number of bits in the third bit set is equal to the bit width indicated by the bit width parameter of the corresponding index value.
[0138] Other parameters may include one or more of the following: color space, transform size, quantization parameters, the number of transform coefficients in the first coded video data, or the number of bits per transform coefficient in the first coded video data.
[0139] The receiving device 104 can obtain second error correction data (506) from the transmitting device 102. The second error correction data provides error correction information about one or more blocks of a second image of the video data.
[0140] Furthermore, the image estimation unit 226 of the receiving device 104 can estimate a second image (508) based on one or more previously reconstructed images (such as a first image). The estimated second image includes predictions of blocks of the second image of the video data based at least in part on blocks of the first image. For example, the image estimation unit 226 can use a combination of inter-frame prediction, intra-frame prediction, and other video decoding tools to generate the prediction data, for example, as described elsewhere in this disclosure.
[0141] The video encoder 228 of the receiving device 104 can generate second coded video data (510) based on the estimated second picture. For example, the video encoder 228 of the receiving device 104 can generate residual data based on the prediction data. For example, the video encoder 228 can perform intra-frame prediction to generate second prediction data based on the estimated second picture. The video encoder 228 can then generate residual data by subtracting the second prediction data from the prediction data generated by the picture estimation unit 226. The video encoder 228 can then generate a transform block by applying one or more forward transforms to the residual data. The video encoder 228 can perform the same process as the video encoder 210 of the transmitting device 102, and therefore may need to use the same parameters as the video encoder 210.
[0142] The channel decoder 222 of the receiving device 104 can perform a channel decoding process to generate second error-corrected coded video data (512) based on the second error-corrected data and the second coded video data. The channel decoder 222 of the receiving device 104 can perform the same process as the channel decoder 222 when generating the first error-corrected coded video data to generate the second error-corrected coded data.
[0143] The video decoder 224 of the receiving device 104 can perform a second video decoding process to reconstruct blocks (514) of the second picture based on the second error-corrected coded video data. The second reconstruction operation is controlled by the value of a parameter. The video decoder 224 can perform the second reconstruction operation in the same manner as the first reconstruction operation. In this way, the receiving device 104 can reconstruct the video data of the picture (or block) without receiving all the coded video data for each picture (or block).
[0144] As described above, the puncturing unit 214 of the transmitting device 102 can perform a bit puncturing operation on the error-correcting data generated by the channel encoder 212. Bit puncturing involves selectively discarding some error-correcting data before the transmitting device 102 transmits the error-correcting data. The discarded bits are generally the least important for performing error correction. The depuncturing unit 220 of the receiving device 104 can perform an inverse bit puncturing operation (i.e., a bit depuncturing operation), which reverses the bit puncturing operation performed by the puncturing unit 214. The puncturing unit 214 can perform the bit puncturing operation according to a set of one or more puncturing parameters. In different examples, the puncturing parameters can be predefined, static, or semi-static.
[0145] According to one or more techniques of this disclosure, transmitting device 102 can perform a decimation process that can reduce memory bandwidth and enhance compression. For example, the video encoder 210 of transmitting device 102 can segment images of video data into a grid of blocks (e.g., MB, LCU, etc.) and can generate transform blocks for each block. Transmitting device 102 may need to store each transform block destined for transmission to receiving device 104 in a memory (e.g., memory 116). Transmitting device 102 can then retrieve the stored transform blocks from memory for channel coding and ultimately for transmission. These writes to and reads from memory can increase time and energy requirements. These time and energy requirements may be directly related to the amount of data to be written and read. Therefore, reducing the amount of data to be written to and read from memory may be advantageous.
[0146] Performing a decimation process can reduce the amount of data written to and read from memory in the transform block. Performing a decimation process can also reduce the amount of data sent from transmitting device 102 to receiving device 104. In some examples, while video encoder 210 is encoding the current block of the current image, video encoder 210 can generate a transform block for the current block. Furthermore, channel encoder 212 of transmitting device 102 can determine whether the current block is the target of decimation based on the decimation mode. If the current block is the target of decimation (i.e., the transform block is a "non-anchor transform block"), channel encoder 212 can reduce the number of bits in the non-anchor transform block before storing the transform block in memory. If the current block is not the target of decimation (i.e., the transform block is an "anchor block"), channel encoder 212 does not reduce the number of bits in the anchor block. Channel encoder 212 can perform the decimation process after generating error correction data. Therefore, the error correction data generated by channel encoder 212 for the non-anchor transform block (and potentially sent to receiving device 104) can be based on the complete bit set of the transform block rather than a reduced number of bits.
[0147] Figure 6 This is a conceptual diagram illustrating an example extraction pattern 600 according to the technology disclosed herein. Figure 6 The example illustrates a grid of DCT blocks. A DCT block is a block of transform coefficients generated by applying a DCT transform to video data such as residual data or sampled data. In other examples, a DCT block can be a transform block generated using other types of transforms. Figure 6 In extraction mode 600, the "X" marker indicates the DCT block (i.e., the non-anchor transform block) as the extraction target. Therefore, in Figure 6 In the example, decimation mode 600 decimates the DCT block by 2 in both the horizontal and vertical directions. In some examples, which may be referred to as "full" decimation, the video encoder 210 can reduce the number of bits in the non-anchor transform block to zero.
[0148] Therefore, in some examples, the decimation mode defines the mode of anchor transform blocks and non-anchor transform blocks in the image. The receiving device 104 may receive the systematic bits of the anchor transform blocks, rather than the systematic bits of the non-anchor transform blocks. The systematic bits of the anchor transform blocks may represent the transform coefficients in the anchor transform blocks. The systematic bits of the non-anchor transform blocks may represent a reduced bit-depth version of the original transform coefficients in the non-anchor transform blocks. Error correction data may include error correction data for both anchor and non-anchor transform blocks. The error correction data for the non-anchor transform blocks is based on the original transform coefficients in the non-anchor transform blocks. As part of generating error-corrected coded video data, the channel decoder 222 may use the error correction data for the anchor transform blocks to perform error correction on the systematic bits in the anchor transform blocks. The channel decoder 222 may use the error correction data for the non-anchor transform blocks to perform error correction on portions of the coded video data corresponding to the non-anchor transform blocks. In some examples, the receiving device 104 may determine the decimation mode and transmit the decimation mode to the transmitting device 102.
[0149] In some examples, transmitting device 102 stores encoded bits (encoded video data and error correction data) in a circular buffer. Transmitting device 102 uses two parameters to select which bits in the circular buffer to transmit. The first parameter is the start position, and the second parameter indicates the number of consecutive bits to transmit. The start position can have various values to support selective transmission and non-transmission of systematic bits. The start position can be selected to skip the transmission of specific systematic bits (i.e., bits of encoded video data) without skipping the transmission of error correction data. Therefore, the decimation of non-anchored transform blocks can be easily achieved by manipulating the first and second parameters, causing transmitting device 102 not to transmit bits of non-anchored transform blocks.
[0150] In some examples, the channel encoder 212 applies a hybrid decimation method that, instead of reducing the number of bits in any target “non-anchor” transform block to zero, reduces the number of bits in the transform coefficients of the non-anchor transform block. For example, in Figure 6 In the example, channel encoder 212 may reduce the number of bits in each transform coefficient in the DCT block marked "X" by a predetermined number (e.g., 2, 4, 5, etc.). Channel encoder 212 does not reduce the number of bits in transform coefficients that are not the target of the decimation mode.
[0151] The channel decoder 222 of the receiving device 104 can receive the remaining, reduced bits of the non-anchor transform block and the error correction data of the non-anchor transform block. As part of the channel decoding process, the channel decoder 222 can use the error correction data of the non-anchor transform block to perform an error correction process that recovers the bits of the removed non-anchor transform block. This error correction process can be the same as the error correction process used by the channel decoder 222 to correct errors introduced by noise in the channel 230. In summary, the channel encoder 212 generates error correction data because it is needed to correct unavoidable noise in the channel 230, but this same error correction data is used to recover bits as if the noise in the channel 230 had just corrupted the least significant bits of a specific transform coefficient in a specific transform block of a specific picture. Therefore, the number of bits transmitted in the channel 230 can be effectively reduced.
[0152] In some examples, the channel encoder 212 may generate a correlation matrix based on the set of transform blocks in the image before performing any decimation process on any non-anchor transform block in the set of transform blocks. The correlation matrix includes values indicating the level of correlation between transform coefficients at corresponding locations within a transform block. For example, the correlation matrix may include correlation values for the DC transform coefficients (i.e., the top-left transform coefficients) of the set of transform blocks. If the differences between the DC transform coefficients are relatively small, the correlation values of the DC transform coefficients are likely to be relatively high. Conversely, if the differences between the DC transform coefficients are relatively large, the correlation values of the DC transform coefficients are likely to be relatively low. Each correlation value can be a value between 0 and 1.
[0153] In some examples, the channel encoder 212 can use the following formula to calculate the correlation values of the DC transform coefficients:
[0154] (5)
[0155] In equation (5) above, l represents the spacing between transform blocks containing DC coefficients, and N represents the number of transform blocks to which the calculation is performed. The function y is the transform coefficient. The line above y represents conjugate. If each consecutive transform block is used, the spacing can be 1; if alternating transform blocks are used, the spacing can be 2, and so on. The channel encoder 212 can calculate the correlation values of the corresponding AC transform coefficients (i.e., non-DC transform coefficients) in the same manner. In this disclosure, the corresponding transform coefficients occupy the same positions within the transform block. Therefore, by calculating the autocorrelation value of each transform coefficient in the transform block, the channel encoder 212 can generate the correlation matrix of the transform block. The channel encoder 212 can repeat the process of generating the correlation matrix for each transform block because the channel encoder 212 will use transform coefficients from different transform blocks when calculating the correlation values.
[0156] Transmitting device 102 can send the correlation matrix along with the encoded video data and error correction data to receiving device 104. The channel decoder 222 of receiving device 104 can use the correlation matrix as part of the process of recovering the original bit width of the non-anchored transform block. For example, continuing with the example of DC transform coefficients, after the error correction process is applied, channel decoder 222 can obtain the values of the non-anchored DC transform coefficients (i.e., the DC transform coefficients in the non-anchored transform block).
[0157] The error correction process can use correlation values to estimate the non-anchor transform coefficients. For example, if we treat all even-numbered transform blocks as anchor transform blocks and non-even-numbered transform blocks as non-anchor transform blocks, the channel decoder 222 can estimate the values of the non-anchor transform coefficients in transform block n in the following way:
[0158] y[n] = (R yy [1]*y[n+1] + R yy [3]*y[n+3] + R yy [3]*y[n+5]…) / ( R yy [1] + R yy [3] + R yy [3]…) (6)
[0159] In equation (6) above, R yy [1] represents the correlation value in the correlation matrix of the transform block with index 1 (i.e., the non-even, non-anchor transform block), y[n+1] represents the corresponding transform coefficient in the anchor transform block with index n+1, R yy [2] represents the correlation value in the correlation matrix of the transform block with index 3, y[n+3] represents the corresponding transform coefficient in the anchor transform block with index n+3, and so on. The number of transform blocks used in equation (6) can be configurable. In this way, the estimated value of the non-anchor transform coefficient can be considered as a weighted average of the corresponding transform coefficients in the anchor block, weighted by the correlation values at corresponding positions in the correlation matrix. In other words, the channel decoder 222 can interpolate the values of the non-anchor transform coefficients according to the correlation matrix. The channel decoder 222 can output the calculated transform coefficient values to the video decoder 224. The channel decoder 222 can perform this process on other transform coefficients. Using the correlation matrix in this way can improve the quality of the reconstructed video data.
[0160] In some examples, the video encoder 210 may reduce the bit width of each transform coefficient in the target transform block by the same amount. In other examples, the video encoder 210 may reduce the bit width of different transform coefficients in the target transform block by different amounts. In some examples, the amount by which the video encoder 210 reduces the bit width of the transform coefficients is related to the distance of the transform coefficients from the anchor transform block. The anchor transform block is the transform block that is not the target of the decimation mode.
[0161] Figure 7 This is a flowchart illustrating an example operation of a transmitting device 102 for hybrid extraction of transform blocks according to the technology of this disclosure. Figure 7 In the example, the sending device 102 can receive data from the video source 120 (…). Figure 7 The video encoder 210 of the transmitting device 102 can generate transform blocks based on the video data (702). For example, the video encoder 210 can generate prediction blocks by performing intra-frame prediction on blocks of images of the video data. The video encoder 210 can use the prediction blocks to generate residual data. The video encoder 210 can generate transform blocks by applying transforms such as DCT, DST, or other transforms to the residual data. In other examples, the video encoder 210 can generate transform blocks by applying transforms directly to blocks of the video data.
[0162] Channel encoder 212 can then determine which transform blocks are anchor transform blocks (704) based on the decimation mode. For example, in channel encoder 212 using Figure 6 In the example of decimation mode 600, video encoder 210 can determine that every other transform block in the horizontal and vertical directions is an anchor transform block. In other examples, channel encoder 212 can use other decimation modes to determine which transform blocks are anchor transform blocks. Channel encoder 212 can store the anchor transform blocks in the memory of transmitting device 102 (e.g., memory 116). Figure 1 ))(706).
[0163] The channel encoder 212 can compute the correlation matrix (708) of the transform block set. Each transform block set includes one or more anchor transform blocks and one or more non-anchor transform blocks. For example, each transform block set can correspond to Figure 6Transform blocks are located in different rows within a transform block set. In another example, each set of transform blocks can correspond to a group of transform blocks consisting of 2 x 2 transform blocks. The number of values in the correlation matrix of each set of transform blocks is the same as the number of transform coefficients in each transform block. Each value in the correlation matrix corresponds to a different position within the transform coefficient block. For example, the value at position (0,0) in the correlation matrix corresponds to the transform coefficient at position (0,0) in each transform coefficient block of the set of transform blocks, the value at position (0,1) in the correlation matrix corresponds to the transform coefficient at position (0,1) in each transform coefficient block of the set of transform blocks, and so on. The transform coefficient at position (0,0) in the transform coefficient block can be called the DC coefficient, and all other transform coefficients can be called the AC coefficients.
[0164] Additionally, the channel encoder 212 can generate a bit-reduced non-anchor transform matrix (710). The non-anchor transform matrix is a transform matrix other than the anchor transform matrix. For example, refer to... Figure 6 The transformation matrix marked with X can be an anchorless transformation matrix. The transformation coefficients in the bit-reduced anchorless transformation matrix can include fewer bits than the original version of the anchorless transformation matrix. The transmitting device 102 can then transmit the anchor transformation block, the anchorless transformation block, the bit-reduced value, the correlation matrix, and the error correction data (712).
[0165] The channel encoder 212 can reduce the bits in the non-anchored transform coefficients in one of several ways. For example, the channel encoder 212 can determine a bit reduction value for each transform coefficient in the non-anchored transform coefficient block. In this example, to calculate the bit reduction value, the channel encoder 212 can calculate the interpolated value of the transform coefficient based on the correlation matrix. The channel encoder 212 can use the above equation (6) to calculate the interpolated value. The channel encoder 212 can then subtract the interpolated value of the transform coefficient from the original value of the transform coefficient to calculate a first distortion value. The channel encoder 212 can then reduce the number of bits of the original value of the transform coefficient by one. The channel encoder 212 can subtract the interpolated value from the reduced bit original value of the transform coefficient to calculate a second distortion value. The channel encoder 212 can determine whether the second distortion value is acceptable based on the first distortion value and the second distortion value. For example, the channel encoder 212 can calculate the average or maximum squared error based on the interpolated value. The channel encoder 212 can determine whether the second distortion value is acceptable by comparing the second distortion value with a predetermined threshold.
[0166] If the second distortion value is acceptable, the channel encoder 212 can reduce the number of bits of the original transform coefficient value and repeat the process. If the second distortion value is unacceptable, the channel encoder 212 can increase the number of bits of the original transform coefficient value. The number of bits obtained by reducing the original transform coefficient value is the bit reduction value.
[0167] In some examples, channel encoder 212 can determine the bit reduction value of the non-anchored transform coefficient block as a whole. In this example, to calculate the bit reduction value of the transform coefficient block, channel encoder 212 can calculate the interpolated value of each transform coefficient based on the correlation matrix, for example, as described above. Video encoder 210 can then subtract the interpolated value of the transform coefficient from the original value of the transform coefficient and use the resulting difference to calculate a first distortion value. For example, video encoder 210 can calculate the first distortion value as a mean square error. Video encoder 210 can then reduce the number of bits of the original value of each transform coefficient by one. Video encoder 210 can subtract the interpolated value from the reduced bit original value of the transform coefficient and use the resulting value to calculate a second distortion value (e.g., using mean square error). Channel encoder 212 can determine whether the second distortion value is acceptable based on the first and second distortion values. If the second distortion value is acceptable, video encoder 210 can again reduce the number of bits of the original value of the transform coefficient and repeat the process. If the second distortion value is unacceptable, video encoder 210 can increase the number of bits of the original value of the transform coefficient. The number of bits obtained by reducing the original values of the transform coefficients is the bit reduction value.
[0168] Therefore, in some examples, transmitting device 102 can obtain video data from a video source. Transmitting device 102 can generate transform blocks based on the video data. Transmitting device 102 can determine which of the transform blocks are anchor transform blocks. Transmitting device 102 can calculate the correlation matrix of the transform block set. Additionally, transmitting device 102 can generate a bit-reduced non-anchor transform matrix. Transmitting device 102 can send the anchor transform blocks, non-anchor transform blocks, and correlation matrix to the receiving device. In some examples, transmitting device 102 can receive an indication of the decimation mode from receiving device 104.
[0169] Figure 8 This is a flowchart illustrating an example operation of a receiving device 104 for hybrid decimation of a transform block according to the technology of this disclosure. Figure 8 In the example, receiving device 104 can receive anchor transform blocks, non-anchor transform blocks, one or more bit reduction values, and the correlation matrix (800) of the non-anchor blocks.
[0170] In addition, Figure 8In the example, the channel decoder 222 of the receiving device 104 can compute the interpolated value (802) of the current non-anchor transform coefficients. The current non-anchor transform coefficients are the transform coefficients of one non-anchor transform block within the non-anchor transform block. The channel decoder 222 can compute the interpolated value of the current non-anchor transform coefficients based on the correlation matrix of the non-anchor block. The channel decoder 222 can compute the interpolated value of the current non-anchor transform coefficients in one of several ways. For example, in some examples, the channel decoder 222 can apply a machine learning model that takes one or more non-anchor transform coefficients (including the current non-anchor transform coefficient group), one or more anchor transform coefficients, and the correlation matrix of the non-anchor transform coefficient block as input. In this example, the machine learning model can output the interpolated value of the current non-anchor transform coefficients. In this example, the machine learning model can be implemented as a neural network model, a support vector machine, a regression model, or another type of machine learning model.
[0171] In another example, the correlation matrix may include values indicating the correlation between the current non-anchor transform coefficient and each corresponding anchor transform coefficient in one or more blocks of anchor transform coefficients. The video decoder 224 may calculate the interpolated values of the current non-anchor transform coefficients as follows:
[0172] (7)
[0173] In the equation above, t int These are the current non-anchor transform coefficients, a i It is the anchor transformation coefficient, c i This indicates that the current non-anchor transform coefficients are related to a. i The correlation value between them, where n represents the number of anchor transform coefficients from which the current non-anchor transform coefficients are derived. In this example, c0 to c n The values can be added together to equal 1.
[0174] Furthermore, the video decoder 224 can calculate the reconstructed values (804) of the non-anchor transform coefficients. The video decoder 224 can calculate the reconstructed values of the non-anchor transform coefficients based on the interpolated values of the current non-anchor transform coefficients and the transmitted values of the non-anchor transform coefficients. The transmitted values of the non-anchor transform coefficients are included in the received non-anchor transform block. In some examples, the video decoder 224 calculates the reconstructed values of the non-anchor transform coefficients as the average of the interpolated values of the current non-anchor transform coefficients and the transmitted values of the non-anchor transform coefficients.
[0175] Video decoder 224 can determine whether any remaining non-anchor transform coefficients exist in the non-anchor transform block (806). If one or more remaining non-anchor transform coefficients exist in the non-anchor transform block (the "yes" branch of 806), video decoder 224 repeats steps 802-806 with another non-anchor transform coefficient. Video decoder 224 can continue doing this until no remaining non-anchor transform coefficients exist (the "no" branch of 806). In this way, video decoder 224 can compute the reconstructed value for each non-anchor transform coefficient.
[0176] In this manner, receiving device 104 can receive the systematic bits of the anchor transform block, the systematic bits of the non-anchor transform block, and the correlation matrix. The systematic bits of the anchor transform block can represent the transform coefficients in the anchor transform block. The systematic bits of the non-anchor transform block can represent a reduced bit-depth version of the original transform coefficients in the non-anchor transform block. As part of performing the reconstruction operation, receiving device 104 can calculate the interpolated value of the non-anchor transform coefficient for each non-anchor transform coefficient in the non-anchor transform block based on the correlation matrix and the corresponding anchor transform coefficient. Reconstructed values of the non-anchor transform coefficients can be calculated at the receiving device based on the interpolated values of the non-anchor transform coefficients and the values of the non-anchor transform coefficients in the error-corrected coded video data.
[0177] In some examples, receiving device 104 can adaptively select a decimation mode for reducing or eliminating bits in a particular transform block. In such an example, receiving device 104 can transmit the selected decimation mode back to transmitting device 102. Transmitting device 102 can then use the selected decimation mode in one or more frames of video data.
[0178] Figure 9 This is a conceptual diagram illustrating an example decimation mode 900 adaptively selected by receiving device 104 according to one or more techniques of this disclosure. Figure 6 In contrast to extraction mode 600, non-anchor blocks in extraction mode 900 do not necessarily appear at regular intervals or spacing.
[0179] Receiving device 104 can determine the decimation mode based on information about previous images of the video data. These previous images may or may not have been decimated. In some examples, receiving device 104 can send a request to sending device 102 for an undecimated version of the image. In response to this request, sending device 102 can send the undecimated version of the image to receiving device 104. After receiving the undecimated version of the image, receiving device 104 can determine the decimation mode based on it. For example, receiving device 104 can perform a rate distortion optimization process that evaluates multiple potential decimation modes to identify which of the decimation modes results in the optimal combination of bit rate and distortion.
[0180] In some examples, transmitting device 102 and receiving device 104 may continue using the selected extraction mode for a predetermined number of images, after which transmitting device 102 and / or receiving device 104 may adaptively select another extraction mode. In some examples, transmitting device 102 may send a message to receiving device 104 requesting transmitting device 102 to select another extraction mode. In some examples, receiving device 104 may determine that an event or condition has occurred that would make selecting another extraction mode advantageous. For example, when receiving device 104 determines that a scene change has occurred, motion in the video data has crossed one or more thresholds, or other characteristics of the video data have changed, receiving device 104 may determine that selecting another extraction mode may be advantageous.
[0181] The receiving device 104 can signal the selected decimation mode to the transmitting device 102 in one of a variety of ways. For example, the receiving device 104 can signal the selected decimation mode to the transmitting device 102 by indicating a difference from an existing decimation mode (such as the currently used decimation mode). For example, in this example, the receiving device 104 can indicate the selected decimation mode to the transmitting device 102 by specifying a change in downsampling or upsampling along a specific axis, specifying a change in a specific region of the image, eliminating a specific transform block, enabling a specific transform block, etc.
[0182] In some examples, there may be a predefined mapping of index values to predefined decimation modes. In such an example, receiving device 104 can select a decimation mode from the predefined decimation modes and signal the index value of the selected decimation mode to transmitting device 102.
[0183] Figure 10 This is a block diagram illustrating example components of a transmitting device and a receiving device according to the technology of this disclosure. Figure 10 In the example, the transmitting device 102 may include a... Figure 2 The same components are shown. However, in Figure 10 In the example, receiving device 104 may additionally include reliability unit 1002. Unless otherwise stated, Figure 2 and Figure 10 The similarly named components of the transmitting device 102 and the receiving device 104 perform the same function.
[0184] Generally, it may be easier to accurately predict the most significant bits (MSB) of transform coefficients compared to their least significant bits (LSB). This is because the MSB is converted to a higher Euclidean distance in the video picture domain. Furthermore, video picture blocks experiencing high motion may be more difficult to predict accurately than blocks in low-motion regions.
[0185] According to one or more techniques disclosed herein, transmitting device 102 and receiving device 104 can implement a system in which bit-level reliability values are used. The use of bit-level reliability values allows transmitting device 102 and receiving device 104 to correctly weight prior information. This can lead to improved system performance (e.g., reduced data transmission volume and / or improved video quality). For example, decoding performance can be improved if “soft” information is used. In other words, when receiving device 104 is “told” of the reliability of each bit (what the prior probability is of a bit value being “0” or “1”), receiving device 104 can utilize this information and can improve performance.
[0186] exist Figure 10 In the example, the reliability unit 1002 of the receiving device 104 receives encoded video data from the video encoder 228 of the receiving device 104. The encoded video data may include transform coefficients of transform blocks of video data. Furthermore, in some examples, the reliability unit 1002 receives prediction quality information from the image estimation unit 226 of the receiving device 104.
[0187] During constant scaling, for each bit position of the transform coefficients in the transform block, the prediction quality information includes a reliability value for that bit position. For example, the most significant bit of a transform coefficient has a first reliability value, the second most significant bit has a second reliability value, the third most significant bit has a third reliability value, and so on. The reliability value of a bit position is a measure of the probability that the bit at that position has an incorrect value. For example, the bit at that position may have an incorrect value when the bit value is predicted as "0" but actually is "1", or when the bit is predicted as "1" but actually is "0".
[0188] Image estimation unit 226 can determine the reliability value of a bit position by collecting statistics on the error rate occurring in bits at that bit position. For example, image estimation unit 226 can determine the probability that the most significant bit of the transform coefficient contains an error, the probability that the second most significant bit of the transform coefficient contains an error, the probability that the third most significant bit of the transform coefficient contains an error, and so on. Image estimation unit 226 can collect this statistical information by counting the number of times the predicted bits (i.e., the bits in the estimated image) are incorrect over multiple test events. For example, image estimation unit 226 can estimate an image, and video encoder 228 can encode the video data of the estimated image. Subsequently, video decoder 224 can decode the error-corrected coded video data of the image. Image estimation unit 226 can compare the bits in the transform coefficients of the coded video data of the estimated image and the error-corrected video data of the image to determine whether the bits in the coded video data of the estimated image are erroneous.
[0189] Reliability unit 1002 can convert probability values into LLR values. In some examples, reliability unit 1002 can convert probability values into LLR values using the following formula:
[0190] (8)
[0191] In the above formula, M represents the LLR value, and P error LLR represents the error probability value, and ln denotes the natural logarithm function. In some examples, the LLR value is the reliability value.
[0192] Figure 11 A graph showing example error probabilities and corresponding absolute values of log-likelihood ratios (LLRs) according to one or more techniques of this disclosure is illustrated. Figure 11 In the example, 150 bits are used to represent each transform block. Curve 1100 plots the error probability at each bit position within the transform block. As can be seen in curve 1100, bits at specific positions have a higher error probability. Curve 1102 shows the error probability converted to the absolute value of the LLR. The absolute value of the LLR can be scaled.
[0193] In some examples, receiving device 104 uses a dynamic scaling process. During dynamic scaling, image estimation unit 226 dynamically determines prediction quality information based on video data. For example, some regions of an image are more difficult to predict than easier-to-predict regions (e.g., regions with reduced prediction accuracy). Examples of difficult-to-predict image regions may include regions with high motion. For example, if the total magnitude of motion vectors in a region exceeds a threshold, image estimation unit 226 can determine that the region is difficult to predict. Therefore, image estimation unit 226 can identify such regions and generate a reliability value for the bits of the transform coefficients of the transform block based at least in part on whether the transform block is inside or outside such a region. In some examples, image estimation unit 226 determines the reliability value of the bits of the transform coefficients based on general statistics regarding errors in the position of bits modified based on whether the transform block containing the transform coefficients is in a difficult-to-predict region.
[0194] The reliability unit 1002 can use a reliability value to scale the bits of the transform coefficients in the encoded video data generated by the video encoder 228. For example, a bit in the encoded video data generated by the video encoder 228 can be considered a "hard" bit and can have a value that is exactly 0 or exactly 1. The reliability unit 1002 can use the prediction quality information generated by the image estimation unit 226 and the encoded video data generated by the video encoder 228 to determine a "soft" bit between 0 and 1. For example, if the value of a bit in the encoded video data is 1, the reliability unit 1002 can generate a "soft" value for that bit by multiplying the absolute value of the LLR of that bit by -1 (i.e., -1). If the value of a bit in the encoded video data is 0, the reliability unit 1002 can generate a "soft" value for that bit by multiplying the absolute value of the LLR of that bit by +1 (i.e., +1). Therefore, the "soft" or scaled value of the bits of the transform coefficients can be M or -M.
[0195] Therefore, in some examples, if there is greater confidence in a bit having a value of 0, each bit can be converted into a scaled value with a more positive value, and if there is greater confidence in a bit having a value of 1, it can be converted into a scaled value with a more negative value. The reliability unit 1002 provides the scaled values as prior information to the channel decoder 222.
[0196] Channel decoder 222 uses scaled values to perform the channel decoding process. For example, reliability unit 1002 and channel encoder 212 can encode the encoded video data into codewords (e.g., low-density parity check codes). The bits of the codeword generated by reliability unit 1002 can be scaled as described above. The bits of the codeword generated by channel encoder 212 can be changed during transmission through channel 230 such that these bits of the codeword can be received as values between -1 and 1. Channel decoder 222 can apply the LPDC decoding process to the codeword to correct errors in the codeword. The bit values in the corrected codeword are 0 or 1. The LPDC decoding process can then convert the codeword from the encoded video data back to the original data. In other examples, other decoding schemes can be used. In this example, the error-correcting data received by channel decoder 222 may include cyclic redundancy check (CRC) data that is not used in the LPDC decoding process. In this way, channel decoder 222 can determine the value of each bit of the transform coefficients.
[0197] In some examples, channel encoder 212 may use predictive quality feedback to sort the bits of the transform coefficients before applying inequality protection channel decoding (e.g., polarity code or ridge code). Inequality protection channel decoding involves allocating decoding redundancy based on the importance of the information bits. For example, channel encoder 212 may use reliability data to protect the different bits of receiver 104 based on their predictability.
[0198] In some examples, the reliability unit 1002 sends predicted quality feedback to the transmitting device 102. The video encoder 210 can adjust one or more encoding parameters applied to the video encoding process of the video data. For example, the transmitting device 102 can determine the compression rate of a limited video encoding process based on the predictability (reliability) of the receiving device 104. For example, if predictability is low, the transmitting device 102 can reduce the quality of the encoded video by reducing the number of transform coefficients or the bit width of each transform coefficient, thus transmitting less information.
[0199] In some examples, transmitting device 102 may adjust one or more channel decoding parameters used by channel encoder 212 based on prediction quality feedback. For example, channel encoder 212 may use unequal guard codes, where the guard depends on the prediction quality feedback. In some examples, channel encoder 212 may select different LDPC maps based on reliability. In some examples, channel encoder 212 may change the coding scheme used to generate error correction data to increase error correction capability for bit positions or regions of the image with lower reliability, or decrease error correction capability for bit positions or regions of the image with higher reliability.
[0200] In some examples, transmitting device 102 may update one or more bit puncturing parameters used by puncturing unit 214 based on predicted quality feedback. For example, puncturing unit 214 may change the puncturing pattern to allow more error correction data to be transmitted for bit locations and / or image regions with lower reliability. Therefore, transmitting device 102 may avoid puncturing bits with lower reliability. In some examples, puncturing unit 214 may change the puncturing pattern to allow less error correction data to be transmitted for bit locations and / or regions of the image with higher reliability.
[0201] In some examples, the predicted quality feedback sent by reliability unit 1002 to transmitting device 102 applies to the entire image. In some examples, reliability unit 1002 may send predicted quality feedback to transmitting device 102 based on each region. Each region may be a defined region within the image. Reliability unit 1002 may send predicted quality feedback for some regions of the image instead of other regions.
[0202] In some examples, reliability unit 1002 may periodically send predicted quality feedback to transmitting device 102. For example, reliability unit 1002 may send predicted quality feedback to transmitting device 102 every N images, where N is an integer value. In some examples, reliability unit 1002 sends predicted quality feedback to transmitting device 102 after completing a specific number of group of pictures (GOPs). In other examples, reliability unit 1002 may send predicted quality feedback non-periodically, such as in response to specific conditions or events.
[0203] In some examples, prediction quality feedback may include prediction quality data based on one or more noise models, such as a Gaussian noise model or a Laplace noise model. Noise model parameters can control one or more noise models. Reliability unit 1002 may send noise model parameters to transmitting device 102. Using a noise model is an alternative approach to collecting error statistics. In this mode, prediction errors (statistics of errors between the estimated image estimated by image estimation unit 226 and the actual image reconstructed by video decoder 224) can be modeled using several parameters describing the error distribution function. The parameters of the noise model may be easier for receiving device 104 to transmit to transmitting device 102 because these parameters may include less data compared to per-bit statistics. Transmitting device 102 may use the noise model in the same manner as transmitting device 102 uses other types of prediction quality feedback.
[0204] In some examples, instead of the reliability unit 1002 receiving prediction quality information from the image estimation unit 226 of the receiving device 104, the video encoder 210 can generate prediction quality information and send it to the reliability unit 1002. The video encoder 210 can determine the prediction quality information based on prior information about the video encoding process. For example, the video encoder 210 can evaluate the reliability per bit based on compression parameters (e.g., MSB is more reliable than LSB, low-frequency transform coefficients are more reliable than high-frequency transform coefficients, etc.). The video encoder 210 can perform the prediction process to evaluate statistics independently, having predefined statistics for different light compression parameter sets. Furthermore, the transmitting device 102 can use other sensors to evaluate instantaneous motion and adjust the reliability accordingly.
[0205] Figure 12 This is a flowchart illustrating an example operation of a transmitting device 102 using scaled bits according to the technology of this disclosure. Figure 12In some examples, transmitting device 102 may obtain video data (1200) from video source 120, for example. Furthermore, transmitting device 102 may obtain prediction quality feedback (1202). In some examples, transmitting device 102 may obtain prediction quality feedback from receiving device 104. The prediction quality feedback includes bit reliability information. In some examples, the prediction quality feedback is represented by noise model parameters, such as parameters of a Gaussian noise model or a Laplace noise model.
[0206] Transmitting device 102 can adjust one or more of video coding parameters, channel coding parameters, or bit puncturing parameters based on prediction quality feedback (1204). Video encoder 210 of transmitting device 102 can perform a video coding process to generate encoded video data (1206). The video coding process can be controlled by video coding parameters. For example, video coding parameters can control the number of transform coefficients included in a transform block, the number of bits included in the transform coefficients, and so on. In some examples, video coding parameters include quantization parameters, and video encoder 210 can adjust the quantization parameters based on prediction quality feedback. For example, if the prediction quality feedback indicates low reliability, the quantization parameters can be decreased to reduce the quantization level. As part of performing the video coding process, video encoder 210 can use the quantization parameters to quantize the transform coefficients of transform blocks of one or more images.
[0207] The channel encoder 212 of the transmitting device 102 can perform a channel coding process on scaled bits to generate channel-coded data (1208). The channel coding process can be controlled by channel coding parameters. For example, channel coding parameters may include control over which LDPC graph to use to generate codewords in the channel coding process, channel coding parameters may control error correction capabilities, and so on. For example, channel coding parameters include LDPC graphs, and the channel encoder 212 can adjust the LDPC graph and use the LDPC graph to generate codewords for transmission to the receiving device.
[0208] Furthermore, the puncturing unit 214 of the transmitting device 102 can perform a bit puncturing process (1210) on the error correction data generated by the channel encoder 212. The bit puncturing process can be controlled by bit puncturing parameters. For example, prediction quality feedback can indicate that certain parts of the encoded video data are unreliable. Therefore, the transmitting device 102 can adjust the bit puncturing parameters to reduce bit puncturing of the error correction data on the unreliable parts of the encoded video data. The transmitting device 102 can transmit the channel-coded data and the bit-punctured error correction data (1212) to the receiving device 104.
[0209] Figure 13 This is a flowchart illustrating an example operation of a receiving device 104 using scaled bits according to the technology of this disclosure. Figure 13In the example, receiving device 104 can obtain error correction data (1300) from the transmitting device. The error correction data provides error correction information about the images of the video data.
[0210] Image estimation unit 226 can generate prediction data for images (1302). The prediction data for images may include predictions of blocks of images based at least in part on one or more previously reconstructed images from video data. For example, image estimation unit 226 may use inter-frame prediction and / or intra-frame prediction to generate block predictions.
[0211] Furthermore, the receiving device 104 can generate encoded video data (1304) based on the predicted data of the image. For example, the video encoder 228 of the receiving device 104 can perform a video encoding process to generate the encoded video data. The encoded video data includes transform blocks, which include transform coefficients.
[0212] The receiving device 104 can scale the bits (1306) of the transform coefficients of the transform block based on the reliability value of the bit position. In some examples, the receiving device 104 generates the reliability value. For example, the receiving device 104 can generate the reliability value based on statistics regarding the occurrence of errors in the bit position. In some examples, the receiving device 104 can generate the reliability value based on the reliability characteristics of various regions of the image of the video data. In some examples, the receiving device 104 can generate the reliability value based on a noise model. Furthermore, in some examples, the receiving device 104 can transmit the reliability value to the transmitting device 102. In other examples, the receiving device 104 can receive the reliability value from the transmitting device 102.
[0213] In addition, the channel decoder 222 of the receiving device 104 can use the error correction data to generate error-corrected coded video data to perform error correction operations on the scaling bits of the transform coefficients of the transform block (1308).
[0214] The video decoder 224 of the channel decoder 222 can reconstruct the image based on the error-corrected coded video data (1310).
[0215] This disclosure describes techniques that can reduce the complexity of video encoding at transmitting devices such as extended reality (XR) headsets. The transmitting device can acquire multi-view video data. Multi-view video data can include images from two or more viewpoints. For example, an XR headset can include two cameras for a stereoscopic view of the scene being viewed by the user. In this example, the multi-view video data can include images from each camera.
[0216] Processing multi-view video data can consume considerable processing resources. For example, in augmented reality (AR) or mixed reality (MR) environments, significant processing resources may be needed to determine the location of virtual elements and how they should appear. Multi-view video data can aid in processing virtual elements. For instance, the same virtual element might need to be darker when positioned in a shadowed area of a scene, and brighter when positioned in a sunlit area. As another example, the system might need to analyze the scene's content to determine if virtual elements are to be occluded by physical elements in the scene, such as rocks or trees. Multi-view video data can be useful in determining the depth of objects in a scene. To keep XR headsets lightweight and conserve battery power, it may be necessary to minimize the processing of video data performed on the XR headset itself. Therefore, processing video data on another device, such as the user's smartphone or other nearby devices, can help reduce the processing resources required on the XR headset.
[0217] While multiview video data can be extremely useful in specific situations, simply transmitting unencoded multiview video data may be impractical because the amount of data required to concurrently transmit multiple parallel video streams can be enormous. However, there is often considerable redundancy between images from different viewpoints in multiview video data. For example, what a person sees with their left eye is often not significantly different from what they see with their right eye. Therefore, video compression techniques have been developed to reduce this redundancy in order to reduce the amount of data required to transmit multiview video data.
[0218] However, some techniques used for multi-view video decoding inherently require considerable computational resources. For example, a video encoder might determine the differences between a set of two or more concurrent images to determine a depth map of a scene. A depth map is an array of values indicating the depth / distance of objects shown in the images from the camera. In this example, one of the concurrent images could be an anchor image, while one or more of the concurrent images could be non-anchor images. The video encoder can use the depth map to compute disparity vectors for blocks in the non-anchor images. The disparity vector of a block indicates the lateral displacement between the block and a corresponding block in another concurrent image (e.g., the anchor image). Typically, blocks representing deeper objects have disparity vectors with lower magnitudes than blocks representing closer objects. The video encoder can use the disparity vectors of the blocks to determine prediction blocks, generate residual data based on the prediction blocks, apply a transform to the residual data, quantize the transform coefficients of the resulting transformed blocks, and send the quantized transform coefficients with a signal. In this example, generating the depth map can involve considerable computational resources.
[0219] In another example, the illumination levels may differ between concurrent images from different viewpoints. These differences in illumination can degrade the decoding efficiency of multi-view video decoding. Therefore, illumination compensation can be applied to the non-anchor images to temporarily modify their illumination levels according to an illumination compensation factor, thereby making the illumination levels of the non-anchor images more consistent with those of the anchor images during video encoding. The original illumination levels of the non-anchor images can be restored during video decoding. Determining the illumination compensation factor can be computationally resource-intensive.
[0220] The technology disclosed herein can offload some processing associated with multiview video coding from a transmitting device (e.g., an XR headset) to a receiving device (e.g., a mobile device). For example, the transmitting device can obtain a first set of multiview images of video data. The first set of multiview images includes a first image and a second image. The first image is from a first viewpoint, and the second image is from a second viewpoint. The transmitting device can send first encoded video data to the receiving device. The first encoded video data is based on the first set of multiview images. The transmitting device can receive multiview coding cues from the receiving device. Furthermore, the transmitting device can obtain a second set of multiview images of video data. The second set of multiview images includes a third image and a fourth image. The third image is from the first viewpoint, and the fourth image is from the second viewpoint. The transmitting device can perform a multiview coding process on the second set of multiview images based on the multiview coding cues received from the receiving device to generate second encoded video data. The multiview coding process reduces interview redundancy between the third and fourth images. The transmitting device can then send the second encoded video data to the receiving device.
[0221] Similarly, the receiving device can obtain first encoded video data from the transmitting device. The first encoded video data is based on a first set of multiview images of the video data. The first set of multiview images may include a first image and a second image. The first image is from a first viewpoint, and the second image is from a second viewpoint. The receiving device can determine multiview encoding cues based on the first encoded video data. The receiving device can send multiview encoding cues to the transmitting device. Additionally, the receiving device can obtain second encoded video data from the transmitting device. The second encoded video data is based on a second set of multiview images including a third image and a fourth image. A multiview encoding process is used to encode the second encoded video data, which reduces interview redundancy between the third and fourth images based on the multiview encoding cues.
[0222] Because the receiving device determines the multiview coding cue and sends it to the transmitting device, the burden of determining the multiview encoder cue can be shifted from the transmitting device to the receiving device. This reduces the resource requirements at the transmitting device.
[0223] refer to Figure 2The video encoder 210 can perform a multiview coding process based on a multiview coding cue obtained from the receiving device 104. For example, the multiview coding cue may include a depth map. In this example, the video encoder 210 can use the depth map to estimate the disparity vector of a block of images in a non-anchor view of the multiview video data. The video encoder 210 can use the disparity vector of the current block of the current image to determine a prediction block of the current block based on samples from a concurrent reference image. The concurrent reference image has the same Picture Order Count (POC) value as the current image. The video encoder 210 can determine the residual data of the current block based on the original samples of the current block and the prediction block of the current block. The video encoder 210 can apply one or more transforms to the residual data to generate one or more transform blocks. The video encoder 210 can quantize the transform coefficients in the transform blocks. The encoded video data generated by the video encoder 210 can be based on the quantized transform coefficients.
[0224] In some examples, multiview coding cues may include one or more illumination compensation factors. When encoding the current image of the multiview video data, the video encoder 210 may modify each sample of the current image based on one or more illumination compensation factors. In some examples, different illumination compensation factors may be applied to different regions of the current image. Modifying the samples of the current image in this way can make the illumination level of the current image more consistent with the illumination level of the concurrent reference image. After modifying the samples of the current image, the video encoder 210 may perform a multiview coding process (such as the process described in the previous paragraphs) to encode blocks of the current image.
[0225] Furthermore, according to some examples of this disclosure, video decoder 224 can perform multi-view decoding processes. For example, video decoder 224 can generate prediction blocks using the disparity vectors of blocks in the current image. Video decoder 224 can reconstruct samples of blocks in the current image using prediction blocks and residual data received from channel decoder 222. In some examples where video encoder 210 applies illumination compensation to the image, video decoder 224 can use illumination parameters to reverse the illumination compensation applied to the image. In other examples, video decoder 224 can apply other multi-view decoding operations.
[0226] Furthermore, according to one or more techniques disclosed herein, the video decoder 224 can determine multiview coding cues based on encoded video data received from the transmitting device 102. For example, the video decoder 224 can determine a depth map, lighting compensation parameters, and other information that can be used in the multiview coding operation. The receiving device 104 can send the multiview coding cues back to the transmitting device 102, enabling the transmitting device 102 to use the multiview coding cues to perform a multiview coding process on subsequent images.
[0227] Image estimation unit 226 can generate an estimate of the next image of the video data. In some examples, image estimation unit 226 can estimate the image based on one or more previously reconstructed reference images associated with different viewpoints. For example, image estimation unit 226 can use information from images at previous times, such as disparity vectors or depth maps, to extrapolate the content of the image from images at the same time. In another example, image estimation unit 226 can extrapolate the image based on one or more images associated with the same viewpoint, regardless of images associated with other views, in a manner substantially similar to that discussed elsewhere in this disclosure regarding image estimation unit 226 for estimating images of single-viewpoint video data.
[0228] Video encoder 228 can perform the same operations as video encoder 210 on the estimated next picture. For example, video encoder 228 can perform intra-frame prediction to predict blocks, using the predicted blocks and corresponding predicted blocks of the picture generated by picture estimation unit 226 to generate residual data. In some examples, video encoder 228 can use multi-view coding cues to perform the same multi-view coding process as video encoder 210. Video encoder 228 can apply a transform (e.g., DCT transform) to the residual data to generate transform coefficients. Video encoder 228 can apply quantization to the transformed coefficients.
[0229] Figure 14 This is a flowchart illustrating an example of data exchange between a transmitting device 102 and a receiving device 104 related to multi-view processing according to one or more technologies of this disclosure. Figure 14 In one example, transmitting device 102 may obtain a first multiview image set (1400). Transmitting device 102 may send first coded video data based on the first multiview image set to receiving device 104. In some examples, transmitting device 102 performs shallow compression on the first multiview image set to generate the first coded video data. In other examples, the first coded video data may include an uncoded version of the first multiview image set.
[0230] The receiving device 104 can perform multiview processing (1402) on the first multiview image set. For example, if needed, the receiving device 104 can decode the first multiview image set. Furthermore, the receiving device 104 can determine multiview encoding prompts, such as those described elsewhere in this disclosure. The receiving device 104 can send the multiview encoding prompts to the transmitting device 102.
[0231] In addition, Figure 14In the example, transmitting device 102 can obtain a second set of multiview images (1404). Transmitting device 102 can perform a multiview encoding process on the second set of multiview images to generate second encoded video data (1406). The second encoded video data may include encoded anchor images and auxiliary (non-anchor) images. Receiving device 104 can perform multiview decoding on the second encoded video data to reconstruct the second set of multiview images (1408). Receiving device 104 can also perform multiview processing on the second set of multiview images to determine updated multiview encoding hints (1410). Receiving device 104 can send the updated multiview encoding hints to transmitting device 102. Transmitting device 102 can use the updated multiview encoding hints to perform multiview encoding on subsequent sets of multiview images.
[0232] Figure 15 This is a flowchart illustrating an example operation of a transmitting device 102 for multi-view processing according to the technology of this disclosure. Figure 5 In the example, transmitting device 102 can obtain a first multi-view image set (1500) of video data. The first multi-view image set includes a first image and a second image. The first image comes from a first viewpoint, and the second image comes from a second viewpoint. The communication interface 118 of transmitting device 102 ( Figure 1 The transmitting device 102 can send first coded video data (1502) to the receiving device 104. The first coded video data is based on a first multiview image set. The transmitting device 102 can receive multiview coding cues (1504) from the receiving device 104. In some examples, the multiview coding cues include one or more of the following: relative shifts between blocks of images from the first and second views, brightness correction between images from the first view, inter-block shifts between anchor blocks and reconstructed blocks, or motion data used to reference the shifts. The transmitting device 102 can receive the multiview coding cues in one of a variety of ways. For example, the transmitting device 102 can receive the multiview coding cues via Uplink Control Information (UCI) / Media Access Control-Control Element (MAC-CE) messages, Radio Resource Control (RRC) messages, or another type of message.
[0233] Furthermore, transmitting device 102 can obtain a second multiview image set (1506) of video data. The second multiview image set includes a third image and a fourth image. The third image is from a first viewpoint, while the fourth image is from a second viewpoint. Video encoder 210 can perform a multiview encoding process on the second multiview image set based on multiview encoding cues received from receiving device to generate second encoded video data (1508). The multiview encoding process reduces interview redundancy between the third and fourth images. Transmitting device 102 can then send the second encoded video data (1510) to receiving device 104.
[0234] For subsequent multi-view image sets, this process can be executed multiple times. Figure 15 The operation is as follows: For example, after sending second encoded video data to a receiving device, sending device 102 can receive updated multiview encoding cues from receiving device 104. Sending device 102 can obtain a third set of multiview images of the video data. The third set of multiview images may include a fifth image and a sixth image, the fifth image being from a first viewpoint and the sixth image being from a second viewpoint. The video encoder 210 of sending device 102 can encode the third set of multiview images based on the updated multiview encoding cues received from the receiving device to generate third encoded video data. Sending device 102 can send the third encoded video data to receiving device 104.
[0235] Figure 16 This is a flowchart illustrating an example operation of a receiving device 104 for multi-view processing according to the technology of this disclosure. Figure 16 In the example, receiving device 104 can obtain first encoded video data (1600) from transmitting device 102. For example, receiving device 104 can obtain this data via communication interface 134. Figure 1 The receiving device 104 obtains first coded video data. The first coded video data is based on a first multi-view image set of the video data. The first multi-view image set may include a first image and a second image. The first image is from a first viewpoint, and the second image is from a second viewpoint. The receiving device 104 may determine multi-view coding cues (1602) based on the first coded video data.
[0236] The receiving device 104 may send a multiview coding prompt (1604) to the transmitting device 102. The receiving device 104 may send the multiview coding prompt in one of a variety of ways. For example, the receiving device 104 may send the decimation mode indication via an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, or another type of message.
[0237] Additionally, receiving device 104 can obtain second encoded video data (1606) from transmitting device 102. The second encoded video data is based on a second multi-view image set including a third image and a fourth image. The second encoded video data is encoded using a multi-view encoding process that reduces inter-view redundancy between the third and fourth images based on multi-view encoding cues. Video decoder 224 of receiving device 104 can decode the second encoded video data.
[0238] In some examples, the multiview coding cue includes a depth map indicating the depth of the object represented in the first and second images. As part of determining the multiview coding cue, the receiving device 104 may determine the depth map based on the first and second images. In some examples, the multiview coding cue includes one or more illumination compensation factors, and as part of determining the multiview coding cue, the receiving device 104 may determine the illumination compensation factors based on the first and second images.
[0239] Figure 16 The process can be repeated multiple times. For example, receiving device 104 can determine a second multiview coding cue based on the second coded video data. Receiving device 104 can send the second multiview coding cue to transmitting device 102. Subsequently, receiving device 104 can obtain third coded video data from transmitting device 102. The third coded video data is based on a third multiview image set including a fifth image and a sixth image, and the third coded video data is encoded using a multiview coding process that reduces interview redundancy between the fifth and sixth images based on the second multiview coding cue.
[0240] According to one or more techniques of this disclosure, transmitting device 102 may receive a decimation mode indication from receiving device 104. The decimation mode indication may indicate a decimation mode. As described in more detail elsewhere in this disclosure, receiving device 104 may determine the decimation mode. The decimation mode may be a mode in which encoded video data is not transmitted.
[0241] Transmitting device 102 can receive the decimation mode indication in one of a variety of ways. For example, transmitting device 102 can receive the decimation mode indication via uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, sidelink control information (SCI) messages, or another type of message.
[0242] Transmitting device 102 can apply a decimation mode to the encoded video data generated by video encoder 210, thereby generating decimated video data. For example, the decimation mode can indicate a mode that skips transmitting encoded video data for complete images. Therefore, in this example, depending on the indicated mode, transmitting device 102 (e.g., channel encoder 212 of transmitting device 102) can transmit encoded video data for some images without transmitting encoded video data for others. For example, transmitting device 102 can skip transmitting encoded video data for every other image. In another example, transmitting device 102 can transmit encoded video data for one image and then not transmit encoded video data for the next two or more images.
[0243] In another example, the extraction mode can indicate a pattern for skipping the transmission of encoded video data for specific regions within an image. Specific regions within a series of images may not vary significantly (if at all) between images. For example, the background of a scene viewed from a static viewpoint may not change significantly when the variation occurs in a more limited region of interest. Because regions outside the region of interest do not vary much, such regions may be easier to predict accurately. Therefore, according to the techniques of this disclosure, receiving device 104 can identify regions outside the region of interest. Thus, transmitting device 102 can transmit encoded video data for the region of interest but not for other regions.
[0244] In another example, the video data is multi-view video data, and the extraction mode can indicate a mode that skips transmitting encoded video data from a specific viewpoint. For example, two viewpoints may have very similar content, such as a viewpoint primarily showing distant objects. Therefore, in this example, as indicated by the extraction mode, transmitting device 102 can transmit encoded video data from one viewpoint without transmitting encoded video data from one or more other viewpoints.
[0245] In another example, the decimation mode can indicate the pattern of bits to be omitted from the syntax element indicating the transform coefficients. For example, the decimation mode can indicate the omission of a specific number of least significant bits from the transform coefficients. In some examples, the decimation mode can indicate the omission of specific transform coefficients (e.g., high-frequency transform coefficients).
[0246] Figure 17 This is a block diagram illustrating example components of a transmitting and receiving device that perform extraction of coded video data according to the techniques of this disclosure. Figure 17 In the example, transmitting device 102 includes a video encoder 210, a channel encoder 212, a punching unit 214, and a transmitter extraction unit 1700. Receiving device 104 includes a de-punching unit 220, a channel decoder 222, a video decoder 224, an image estimation unit 226, a video encoder 228, and a receiver extraction unit 1702. The video encoder 210, channel encoder 212, punching unit 214, de-punching unit 220, channel decoder 222, video decoder 224, image estimation unit 226, and video encoder 228 can operate in the same manner as described elsewhere in this disclosure.
[0247] However, in Figure 17In the example, after the channel encoder 212 generates error correction data for the encoded video data, the transmitter decimation unit 1700 can apply a decimation mode to the encoded video data. The decimation mode indicates the mode in which the encoded video data is not transmitted. For example, the transmitter decimation unit 1700 can prevent the transmitting device 102 from transmitting encoded video data for a specific picture, picture region, block pattern within a picture, specific viewpoint, etc. The receiver decimation unit 1702 can determine the decimation mode indication based on the picture reconstructed by the video decoder 224. The receiver decimation unit 1702 can send a decimation mode indication to the transmitting device 102. The transmitter decimation unit 1700 can apply the decimation mode indicated by the decimation mode indication.
[0248] although Figure 17 This description pertains to a DVC-based scheme, but the techniques disclosed herein related to sending a decimation mode indication from receiving device 104 to transmitting device 102 are not necessarily limited thereto. For example, in some examples, the image estimation unit 226 and the video encoder 228 may be omitted.
[0249] Figure 18 This is a conceptual diagram illustrating an example exchange of information including extraction pattern indications according to the technology of this disclosure. Figure 18 In the example, sending device 102 can send a first set of images (e.g., images n-n1, images nn) to receiving device 104. n1+1 The transmitting device 102 can also transmit error correction data for the first set of images (n).
[0250] Receiving device 104 can transmit and transmitting device 102 can receive a decimation mode indication, which indicates a decimation mode determined based on a first set of coded images. Figure 18 In the example, the extraction pattern extracts images at a 1:2 ratio. In other words, the encoded video data of one out of every two images will be sent.
[0251] Therefore, transmitting device 102 can send the encoded video data of the second image set to receiving device 104. According to the decimation mode indicated by the received decimation mode instruction, transmitting device 102 skips transmitting the encoded video data of every other image in the second image set. For example... Figure 18 As shown in the example, the index values of the pictures in the second picture set (e.g., n+2, n+4, n+n2) are increased by 2 instead of 1 as in the case of the first picture set.
[0252] Subsequently, receiving device 104 can determine, based on the second set of images, that a more suitable decimation mode would be a 1:1 decimation mode (i.e., the decimation mode in which transmitting device 102 transmits coded video data for each image). Therefore, in Figure 18 In the example, receiving device 104 can send and transmitting device 102 can receive a second decimation mode indication indicating a second decimation mode. Subsequently, transmitting device 102 can transmit encoded video data of the third picture set. According to the second decimation mode, transmitting device 102 does not skip transmitting encoded video data of any picture in the third picture set. Therefore, as... Figure 18 As shown in the example, the index value (e.g., n+n2+1, n+n2+2, etc.) is increased by 1 instead of 2.
[0253] Figure 19 This is a flowchart illustrating an example operation of a transmitting device 102 according to the technology of this disclosure, wherein the transmitting device 102 receives a decimation mode indication. Figure 19 In the example, the video encoder 210 of the transmitting device 102 can encode a first set of images of the video data to generate first encoded video data (1900). The transmitting device 102 can then send the first encoded video data (1902) to the receiving device 104.
[0254] Furthermore, transmitting device 102 can receive a decimation mode indication from receiving device 104, which indicates a decimation mode (1904) determined based on the first set of images. The decimation mode can be a mode in which encoded video data is not transmitted. For example, in some examples, the decimation mode indicates a mode that skips the transmission of encoded video data for the entire image. In other words, transmitting device 102 may not transmit any encoded video data for a particular image, and may transmit some or all of the encoded video data for other images. In some examples, the decimation mode indicates a mode that skips the transmission of encoded video data for a specific region within an image. For example, the decimation mode may indicate that transmitting device 102 skips the transmission of encoded video data associated with a specific block of an image, such as... Figure 6 and Figure 9 As shown. In some examples where the video data is multi-view video data, the decimation mode can indicate a mode that skips the transmission of encoded video data from a specific viewpoint. In such examples, the viewpoint may be associated with a sensor on the same user device (e.g., the same XR headset), or the viewpoint may be associated with a sensor or camera on different user devices (e.g., different XR headsets worn by different users). In some examples, different decimation modes may exist for different regions where no image is available. For example, decimation may not be applied to the region of interest, and decimation modes that restrict the transmission of blocks or low effective bits or higher frequency transform coefficients may be applied to image regions outside the region of interest.
[0255] The video encoder 210 can encode the second set of images of the video data to generate second encoded video data (1906). Additionally, the transmitter decimation unit 1700 can apply a decimation mode to the second encoded video data to generate decimated video data (1908). The transmitting device 102 can send the decimated video data to the receiving device (1910).
[0256] In some examples, the transmitter decimation unit 1700 can determine the decimation mode. Therefore, in Figure 19 In the context of [the above context], transmitting device 102 can encode a third set of images of video data to generate third encoded video data, determine a second decimation mode indicating that the encoded video data has not been transmitted, and apply the second decimation mode to the third encoded video data to generate second decimated video data. Transmitting device 102 can send the second decimated video data to receiving device 104. Transmitting device 102 can also send a second decimation mode indication to receiving device. The second decimation mode indication indicates that the second decimation mode is applied to the third encoded video data.
[0257] The transmitter decimation unit 1700 can determine decimation patterns in various ways. For example, the transmitter decimation unit 1700 can test various decimation patterns. When testing a decimation pattern, the transmitter decimation unit 1700 can apply the decimation pattern to an image and reconstruct the image based on the image's error correction data and one or more previous original images of video data. The transmitter decimation unit 1700 can compare the reconstructed image with the image to determine the distortion level. The transmitter decimation unit 1700 can compare the distortion levels associated with different decimation patterns to determine the decimation pattern.
[0258] In some examples, Figure 19 The operations are performed within the DVC context. Therefore, the channel encoder 212 of the transmitting device 102 can generate first error correction data based on the first coded video data. The transmitting device 102 can send the first error correction data to the receiving device 104. The channel encoder 212 can generate second error correction data based on the second coded video data. The transmitting device 102 can send the second error correction data to the receiving device.
[0259] Figure 20 This is a flowchart illustrating an example operation of a receiving device 104 according to the technology of this disclosure, wherein the receiving device 104 transmits a decimation mode indication. Figure 20 In the example, receiving device 104 can receive first coded video data (2000) from transmitting device 102. Video decoder 224 can perform a decoding process to reconstruct a first set of pictures based on the first error-corrected coded video data (2002).
[0260] Additionally, the receiver extraction unit 1702 can determine an extraction mode based on the first set of images, which indicates a mode in which encoded video data is not transmitted (2004). In some examples, the extraction mode indicates a mode in which encoded video data of the entire image is skipped. In some examples, the extraction mode indicates a mode in which encoded video data of a specified region or block within an image is skipped, such as in... Figure 6 and Figure 9 In the example, in some embodiments, the video data is multi-view video data and the extraction mode indicates a mode that skips the transmission of encoded video data of images from a particular viewpoint.
[0261] The receiver decimation unit 1702 can determine the decimation mode in one of several ways. For example, for one or more trial decimation modes, the receiver decimation unit 1702 can apply the trial decimation mode to error-corrected coded video data generated by the channel decoder 222 for the first image set to generate decimated coded video data. The receiver decimation unit 1702 can then cause the channel decoder 222 to apply an error correction process to modify the decimated coded video data based on the first error-corrected data to generate trial error-corrected video data. The receiver decimation unit 1702 can then cause the video decoder 224 to apply a decoding process to reconstruct the first image set based on the trial error-corrected video data. The receiver decimation unit 1702 can determine whether the decimation mode meets one or more criteria by comparing the first image set reconstructed based on the trial error-corrected video data with the first image set reconstructed based on the first error-corrected video data. For ease of explanation, the image reconstructed based on the trial error-corrected video data may be referred to as a "trial image," and the image reconstructed based on the first error-corrected video data may be referred to as a "baseline image." The receiver extraction unit 1702 can repeat the process with multiple trial extraction patterns until the receiver extraction unit 1702 identifies an extraction pattern that meets the criterion.
[0262] For example, receiver extraction unit 1702 can compare each test image with a corresponding baseline image to determine whether the test image meets a criterion. For example, if the sum of the differences between the test image and the corresponding baseline image is less than a certain amount, receiver extraction unit 1702 can determine that the test image meets the criterion. If at least a given number of test images exceeds a threshold, receiver extraction unit 1702 can select an extraction mode associated with the test images.
[0263] In a more general example, receiver extraction unit 1702 can apply a function to the test image and the corresponding baseline image to generate a value. If the value is less than a threshold, receiver extraction unit 1702 can select an extraction mode associated with the test image.
[0264] Furthermore, in some examples, the receiver decimation unit 1702 may deactivate the decimation mode (e.g., revert to a mode that transmits all encoded video data) or switch to a less aggressive decimation mode if specific conditions are met. For example, the receiver decimation unit 1702 may deactivate the decimation mode in response to determining that a given number of baseline images fail to meet a criterion. For instance, if the sum of the differences between the test image and the corresponding baseline image is greater than a certain amount, the receiver decimation unit 1702 may determine that the test image does not meet the criterion. If at least a given amount of the test images that do not meet the criterion exceeds a threshold, the receiver decimation unit 1702 may deactivate the decimation mode or revert to a less aggressive decimation mode. In a more general example, the receiver decimation unit 1702 may apply a function to the test image and the corresponding baseline image to generate a value. If this value is greater than a second threshold, the receiver decimation unit 1702 may deactivate the decimation mode or revert to a less aggressive decimation mode.
[0265] The receiver decimation unit 1702 can send a decimation mode indication (2006) to the transmitting device 102, indicating the decimation mode determined therein. The receiver decimation unit 1702 can send the decimation mode indication using uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, sidelink control information (SCI) messages, or another type of message.
[0266] Receiving device 104 can receive extracted video data (2008) from transmitting device 102. The extracted video data may include second encoded video data for which an extraction mode has been applied. The second encoded video data is generated based on a second set of images from the video data.
[0267] Video decoder 224 can perform a decoding process to reconstruct a second set of pictures based on the second error-corrected coded video data (2010). Video decoder 224 can perform the same decoding process as described elsewhere in this disclosure.
[0268] In some examples, receiving device 104 can receive and use a decimation mode indication from transmitting device 102. Therefore, in Figure 20In the example, receiving device 104 can receive a second decimation mode indication indicating that a second mode of encoded video data has not been transmitted. Receiving device 104 can receive third error-correcting data and second decimated video data from transmitting device 102. The second decimated video data may include third encoded video data for which the second decimation mode has been applied. Third encoded video data can be generated based on a third set of images of the video data. Channel decoder 222 can apply an error correction process based on the third encoded video data and the third error-correcting data to generate third error-corrected encoded video data. Video decoder 224 can apply a decoding process to reconstruct the third set of images based on the third error-corrected encoded video data.
[0269] In some examples, Figure 20 The process can be executed in a DVC-based implementation. Therefore, receiving device 104 can receive first error-corrected data from transmitting device 102. Receiving device 104 applies an error-correction process to modify first coded video data based on the first error-corrected data to generate first error-corrected coded video data. Receiving device 104 can perform a decoding process to reconstruct a first set of images based on the first error-corrected coded video data. Additionally, receiving device 104 can receive second error-corrected data from transmitting device 102. Receiving device 104 can apply an error-correction process to generate second error-corrected coded video data based on the second coded video data, predictive coded video data generated by video encoder 228, and the second error-corrected data. Receiving device 104 can perform a decoding process to reconstruct a second set of images based on the second error-corrected coded video data.
[0270] During the video encoding process, a video encoder typically analyzes multiple encoding options and selects the optimal one. For example, a video encoder might analyze various ways to divide a maximum decoding unit (LCU) or macroblock into decoding units (CUs) and / or prediction units (PUs). In another example, when performing intra-frame prediction to generate prediction blocks for a PU, a video encoder might analyze multiple intra-frame prediction modes. In yet another example, when performing inter-frame prediction to generate prediction blocks for a PU, a video encoder might analyze multiple reference frames and motion vectors. This analysis and selection can be resource-intensive. For example, to be efficient, a video encoder might need to process multiple options in parallel, which increases the hardware complexity and power requirements of the video encoder. Analysis and selection may also involve multiple requests to read and write data to memory, further increasing power demands.
[0271] According to one or more techniques of this disclosure, a large portion of the process of analyzing and selecting encoding operations is transferred from a transmitting device (e.g., transmitting device 102) to a receiving device (e.g., receiving device 104). For example, the transmitting device may encode a first image of video data to generate first encoded video data. The transmitting device may send the first encoded video data to the receiving device. The receiving device may receive the first encoded video data from the transmitting device and reconstruct the first image based on the first encoded video data. Additionally, the receiving device may estimate a second image of the video data based on the first image. The second image may be an image that appears after the first image in decoding order. The receiving device may generate encoding selection data for estimating the second image. The encoding selection data indicates encoding selections for encoding the estimated second image. The receiving device may send the encoding selection data for the second image. The transmitting device may receive the encoding selection data for the second image of the video data. The transmitting device may encode the second image based on the encoding selection data to generate second encoded video data. The transmitting device may send the second encoded video data to the receiving device. The receiving device may receive the second encoded video data from the transmitting device. The receiving device may reconstruct the second image based on the second encoded video data. In this way, because the transmitting device receives the encoding selection data from the receiving device, the transmitting device does not need to perform resource-intensive analysis and selection processes while encoding the second image, since the analysis and selection process for the second image has already occurred at the receiving device. This reduces the resource requirements of the transmitting device.
[0272] Transmitting and receiving devices can communicate over low-range, low-power links on ultra-wideband (e.g., high-bandwidth) communication links. In some examples, the transmitting and receiving devices can communicate using time-division duplex (TDD), subband non-overlapping full-duplex (SBFD), or SFFD schemes. The low latency associated with this type of communication allows the transmitting device to receive coded selection data quickly enough to continue transmitting coded video data at a predetermined picture rate.
[0273] Figure 21 This is a block diagram illustrating example components of a transmitting device 102 according to the technology of this disclosure and a receiving device 104 that transmits encoding selection data to the transmitting device. Figure 21 In the example, the transmitting device 102 may include a video encoder 210, a channel encoder 212, and a punching unit 214. The receiving device 104 may include a de-punching unit 220, a channel decoder 222, a video decoder 224, an image estimation unit 226, and a video encoder 228.
[0274] exist Figure 21In some examples, video encoder 210 can encode images of video data. Unlike some of the examples provided above, video encoder 210 can perform a complete video encoding process that may include intra-frame and inter-frame prediction. In some examples, video encoder 210 can encode video data using video codecs such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, Basic Video Decoding (EVC), AV1, etc. Channel encoder 212, puncturing unit 214, depuncturing unit 220, and channel decoder 222 can operate in the same manner as described elsewhere in this disclosure.
[0275] In addition, Figure 21 In the example, the video decoder 224 of the receiving device 104 can perform a video decoding process on the error-corrected coded video data generated by the channel decoder 222. The video decoder 224 can perform a complete video decoding process including intra-frame prediction and inter-frame prediction. The video decoder 224 can use the same video codec as the video encoder 210.
[0276] After the video decoder 224 reconstructs at least a portion of the images from the video data, the image estimation unit 226 can estimate the corresponding portion of subsequent images following the reconstructed images. The image estimation unit 226 can estimate subsequent images in the same manner as described elsewhere in this disclosure. Furthermore, the video encoder 228 can apply the video encoding process to subsequent images. In the example where both the video encoder 210 and the video decoder 224 use a video codec, the video encoder 228 can use the same codec.
[0277] However, according to one or more techniques of this disclosure, receiving device 104 may send encoding selection data 2100 to transmitting device 102. Encoding selection data 2100 indicates encoding selection for encoding the estimated subsequent images. For example, encoding selection data may include motion parameters of blocks in the estimated subsequent images. The block motion parameters may include motion vectors, reference image indicators, merge candidate indices, affine motion parameters, and other data for determining predicted blocks in one or more reference images. Therefore, in this example, video encoder 228 may use inter-frame prediction to encode the blocks, and encoding selection data 2100 may include motion parameters instructing video encoder 228 how to use inter-frame prediction to encode the blocks.
[0278] In some examples, the encoding selection data may include intra-prediction parameters for blocks in the estimated subsequent images. Intra-prediction parameters may include data indicating the intra-prediction mode (e.g., planar mode, DC mode, directional prediction mode, etc.) used by the video encoder 228 for intra-prediction of blocks. In some examples, the encoding selection data may include other information, such as information describing how the video encoder 228 divides the estimated subsequent images into blocks, whether residual prediction is used, whether and how intra-block copy (IBC) is used, and whether specific filters are used.
[0279] Video encoder 210 can use encoding selection data 2100 when encoding actual (non-estimated) subsequent images. That is, video encoder 210 can use the video encoding selection indicated by encoding selection data 2100 instead of searching for different possibilities during the video encoding process. For example, encoding selection data 2100 can indicate encoding a specific block of a subsequent image with a specific intra-prediction mode. Therefore, in this example, when encoding a subsequent image, video encoder 210 can utilize a specific intra-prediction mode to encode a specific block without analyzing different potential intra-prediction modes to select that specific intra-prediction mode. In another example, encoding selection data 2100 can indicate the motion vectors and reference image of a specific block of a subsequent image. Therefore, in this example, when encoding a subsequent image, video encoder 210 can use the motion vectors to determine the predicted block in the reference image without analyzing potential reference images and motion vectors. Video encoder 210 can then use the predicted block to encode the specific block.
[0280] Transmitting device 102 can process the encoded video data of subsequent images in the same way as other encoded video data. Similarly, receiving device 104 can process the encoded video data of subsequent images in the same way as other encoded video data. Therefore, after video decoder 224 reconstructs at least a portion of a subsequent image, image estimation unit 226 can predict the corresponding portion of the image following the subsequent image, video encoder 228 can encode the video data of the estimated subsequent image, and transmit encoding selection data for the estimated subsequent image; this cycle can be repeated. In this way, some of the burden of encoding video data can be transferred from video encoder 210 of transmitting device 102 to video encoder 228 of receiving device 104. This reduces the resource requirements of transmitting device 102.
[0281] about Figure 21The described process can be adapted for use with DVC technology. For example, transmitting device 102 can apply a decimation mode to the encoded video data (e.g., according to any example provided elsewhere in this disclosure), such that transmitting device 102 transmits only some encoded video data, but still transmits error-corrected data of the decimated video data. The channel decoder 222 of receiving device 104 can receive error-corrected data for a specific image from the transmitting device (e.g., via the de-puncturing unit 220). In this example, channel decoder 222 can apply an error-correction process to generate error-corrected encoded video data based on the error-corrected data for the specific image generated by the video encoder 228 of receiving device 104 and the encoded video data for that specific image. Video decoder 224 can decode the error-corrected encoded video data to reconstruct specific features.
[0282] In some cases, transmitting device 102 needs to transmit encoded video data according to a schedule. For example, to support a specific application, transmitting device 102 may need to transmit encoded video data to receiving device 104 at a predetermined pictures per minute rate. Therefore, it is possible that transmitting device 102 may not receive the image encoding selection data in time for encoding and transmitting the image's encoded video data. Thus, in some examples, based on the determination that the image encoding selection data has not been received from receiving device 104 before the time limit expires, video encoder 210 may encode the image without using the image encoding selection data. Video encoder 210 may use a finite video encoding process to encode the image. Furthermore, in some examples, the image's encoded video data may include encoding selection data generated by video encoder 210. In some examples, the image's encoded video data may include data indicating that the image's encoded video data is not generated based on the encoding selection data generated by receiving device 104. This time limit may be subject to or defined based on the capabilities of transmitting device 102.
[0283] Encoding selection data and time constraints can be defined on a per-image-fragment basis (e.g., slices, regions, etc.). Therefore, in this disclosure, the discussion of encoding selection data for images, encoded video data, or other types of data may apply only to individual image fragments.
[0284] Figure 22 This is a communication diagram illustrating an example of data exchange between a transmitting device 102 and a receiving device 104, including the transmission and reception of encoded selection data according to the technology of this disclosure. Figure 22In the example, transmitting device 102 sends encoded video data of image n-1 to receiving device 104. Receiving device 104 can reconstruct image n-1 based on the encoded video data of image n-1. Additionally, receiving device 104 can estimate image n and encode it based on image n-1. Receiving device 104 can send encoding selection data for image n to transmitting device 102. Transmitting device 102 can encode image n based on the encoding selection data and send the resulting encoded video data of image n to receiving device 104. This process can be repeated multiple times. Therefore, in Figure 22 In the example, receiving device 104 can reconstruct image n based on the encoded video data of image n, estimate image n+1 based on image n and / or one or more other previously reconstructed images, encode image n+1, and send the encoding selection data of image n+1 to transmitting device 102.
[0285] Figure 23 This is a flowchart illustrating an example operation of a transmitting device 102 according to the technology of this disclosure, wherein the transmitting device 102 receives encoded selection data. Figure 23 In the example, video encoder 210 encodes a first image of video data to generate first encoded video data (2300). In some examples, if transmitting device 102 has not yet received encoding selection data for the first image, transmitting device 102 may perform a finite video coding process on the first image. Compared to a full video coding process, a finite video coding process can use decoding tools with relatively lower computational intensity. For example, a finite video coding process can use intra-frame prediction instead of inter-frame prediction.
[0286] Transmitting device 102 can send first coded video data (2302) to receiving device 104. In some examples, transmitting device 102 can apply a channel coding process to the first coded video data to generate error correction data for the first coded video. Transmitting device 102 can send both the first coded video data and the error correction data to receiving device 104.
[0287] Subsequently, transmitting device 102 can receive encoding selection data (2304) of a second image of video data from receiving device 104. The encoding selection data can indicate encoding selections used to encode the estimate of the second image. The second image follows the first image in decoding order. In some examples, the second image may appear before or after the first image in the output decoder. In some examples, the encoding selection data is entropy-coded. Therefore, in such examples, transmitting device 102 can entropy-decode the encoding selection data. For example, transmitting device 102 can apply CABAC decoding, Golomb-Rice decoding, or another type of entropy decoding to the encoding selection data. In some examples, the encoding selection data is channel-coded. Therefore, transmitting device 102 can apply error correction operations to the encoding selection data based on error correction data used for the encoding selection data.
[0288] The video encoder 210 of transmitting device 102 can encode the second picture based on encoding selection data to generate second encoded video data (2306). For example, the encoding selection data may include data indicating how a particular macroblock is divided into CUs. In this example, the video encoder 210 may divide the macroblock into CUs in the manner indicated by the encoding selection data. In another example, the encoding selection data may indicate an intra-prediction mode for the block (e.g., CU or PU), and the video encoder 210 may use the indicated intra-prediction mode to encode the block. Therefore, in this example, the encoding selection data received from receiving device 104 may include intra-prediction parameters for the blocks of the second picture, and as part of encoding the second picture, the video encoder 210 may perform intra-prediction based on the intra-prediction parameters for the blocks of the second picture to generate predicted blocks. The second encoded video data may include encoded video data based on the predicted blocks.
[0289] In another example, the encoded selection data received from receiving device 104 includes motion parameters for blocks of the second image, and as part of encoding the second image, transmitting device 102 can perform motion compensation based on these motion parameters for the blocks of the second image to generate predicted blocks. The second encoded video data includes encoded video data based on the predicted blocks.
[0290] The transmitting device 102 may send second encoded video data (2308) to the receiving device. In some examples, the second encoded video data does not include encoding selection data, which indicates the encoding selection used by the transmitting device 102 when encoding the second picture or the encoding selection used by the receiving device 104 when encoding an estimate of the second picture. The second encoded video data may not need to include encoding selection data because the receiving device 104 has generated and therefore already has encoding selection data.
[0291] In some examples, Figure 23 The operation can be used in conjunction with DVC technology. For example, transmitting device 102 can receive encoding selection data for a third frame of video data. The encoding selection data for the third frame can indicate encoding selections used to encode an estimate of the third frame. Video encoder 210 can encode the third frame based on the encoding selection data for the third frame to generate third-coded video data. Channel encoder 212 can apply a channel coding process to generate error correction data for the third-coded video data. Transmitting device 102 can send the error correction data of the third-coded video data to the receiving device without sending at least a portion of the third-coded video data.
[0292] Figure 24 This is a flowchart illustrating an example operation of a receiving device 104 according to the technology of this disclosure, wherein the receiving device 104 transmits encoded selection data. Figure 24 In the example, receiving device 104 can receive first encoded video data (2400) from sending device.
[0293] The video decoder 224 of the receiving device 104 can reconstruct a first image (2402) of the video data based on the first encoded video data. The image estimation unit 226 of the receiving device 104 can estimate a second image (2404) of the video data based on the first image. The second image may be an image that appears after the first image in the decoding order.
[0294] The video encoder 228 of the receiving device 104 can generate encoding selection data (2406) for estimating the second image. The encoding selection data indicates encoding choices used to encode the estimated second image. For example, as part of encoding the estimated second image, the video encoder 228 can perform motion compensation based on motion parameters of blocks of the second image to generate prediction blocks. In this example, the encoding selection data may include motion parameters of blocks of the second image from one or more processors. In some examples, as part of encoding the second image, the video encoder 228 can perform intra-frame prediction based on intra-frame prediction parameters of blocks of the second image to generate prediction blocks. In this example, the encoding selection data may include intra-frame prediction parameters of blocks of the second image.
[0295] Receiver 104 can send encoding selection data (2408) for a second picture to transmitter 102. In some examples, receiver 104 can apply entropy coding (e.g., CABAC coding, Golomb-Rice decoding, etc.) to the encoding selection data for the second picture before sending it. In some examples, receiver 104 can perform a channel coding process on the encoding selection data to generate error correction data for the encoding selection data. Receiver 104 can send the encoding selection data and the error correction data for the encoding selection data to transmitter 102. In some examples, compared to other data transmissions in the data link (e.g., wireless sidelink channel 112, wireless uplink / downlink channel, etc.) between receiver 104 and transmitter 102, the communication interface 134 of receiver 104 ( Figure 1 The encoded selection data can be modulated at a lower modulation order. This increases the likelihood that the transmitting device 102 will correctly receive the encoded selection data.
[0296] Subsequently, receiving device 104 can receive second encoded video data (2410) from transmitting device. Video decoder 224 can reconstruct a second picture (2412) based on the second encoded video data. In some examples, the second encoded video data does not include encoding selection data. Video decoder 224 can apply a decoding process that includes reconstructing the second picture based on the second encoded video data using the encoding selection data.
[0297] Figure 24 The process can be used in conjunction with DVC technology. For example, image estimation unit 226 can estimate a third image of video data based on one or more of a first image or a second image. Video encoder 228 can encode the estimated third image to generate third encoded video data. Receiving device 104 can send third encoding selection data to transmitting device 102. The third encoding selection data can indicate encoding selections used to encode the estimated third image. Subsequently, receiving device 104 can receive error correction data for the third image. Channel decoder 222 can apply an error correction process based on the error correction data and the third encoded video data to generate error-corrected encoded video data for the third image. Video decoder 224 can apply a decoding process that reconstructs the third image based on the error-corrected encoded video data. In some examples where transmitting device 102 and receiving device 104 use DVC technology, the error-corrected video data for the third image does not include the third encoding selection data. However, the video decoder 224 can use the third encoding selection data generated by the video encoder 228 of the receiving device 104 to apply the decoding process to reconstruct the third picture based on the error-corrected encoded video data of the third picture.
[0298] Figure 25 This is a conceptual diagram illustrating an example hierarchical structure of encoded video data according to the technology of this disclosure. More specifically, Figure 25 The hierarchical structure of encoded video data generated using the H.264 / AVC video decoding standard is shown. For example... Figure 25 As shown in the example, the Network Abstraction Layer (NAL) is the highest layer of the hierarchy. In the NAL, data is organized into NAL units. In some examples, NAL units are assigned to different packets or decoding blocks for transmission. NAL units in the NAL can include Sequence Parameter Sets (SPS) and Picture Parameter Sets (PPS) containing high-level syntax. NAL units in the NAL can also include Video Decoding Layer (VCL) NAL units. VCL NAL units can include slice NAL units containing slice-level data. A slice can be a series of macroblocks within a picture. Slices can include Instantaneous Decoder Refresh (IDR) slices and regular slices. Decoding an IDR slice does not depend on any other slice. Regular slices may depend on other stripes.
[0299] Each slice NAL unit can include a slice header and slice data. The slice header of the slice NAL unit includes information for decoding the slice data of the slice NAL unit. The slice data of the slice NAL unit consists of a series of macroblocks (MBs). Skip indicators can be scattered between the MBs. Each MB contains encoded video data for a specific block used in the slice. Furthermore, as... Figure 25 As shown, a Block Markup (MB) can include a type indicator, prediction information, decoded block mode, quantization parameters (QP), and coded residual data. If intra-frame prediction is used to encode the MB, the prediction data can indicate one or more intra-frame modes used to encode the MB. If inter-frame prediction is used to encode the MB, the prediction data can indicate one or more reference images and one or more motion vectors. The coded residual data of the MB can include coded residual data for luma blocks within the MB, coded residual data for Cb blocks within the MB, and coded residual data for Cr blocks within the MB. Typically, the coded residual data constitutes the largest portion of the encoded video data.
[0300] According to the technology disclosed herein, everything in the macroblock layer hierarchy, except for the encoded residual data, can be encoded selection data. Therefore, in some examples, receiving device 104 can send type data, prediction data, decoded block mode, and QP to transmitting device 102 for each MB of the estimated image. Furthermore, in some examples, transmitting device 102 can send only the encoded residual data of the MB to receiving device 104, without sending the type data, prediction data, decoded block mode, or QP of the MB. In some examples, the encoded selection data sent by receiving device 104 may include slice header data, SPS data, and PPS data. Transmitting device 102 and receiving device 104 can exchange information, or information indicating encoder and decoder capabilities can be pre-configured.
[0301] In some examples, transmitting device 102 may also transmit data other than some images, some MB, or some slices of encoded residual data. Transmitting device 102 may send information (e.g., bits) at the network abstraction layer (e.g., as an image-level control field), which may indicate whether a DVC-based approach should be used or whether conventional compression should be used for a particular image.
[0302] Figure 26 This is a block diagram illustrating alternative example components of a transmitting device 102 according to one or more technologies of this disclosure. Figure 26 In the example, transmitting device 102 performs digital encoding and analog encoding on the video data. Transmitting device 102 transmits digitally encoded video data and analog-encoded video data to receiving device 104 via channel 230.
[0303] exist Figure 26 The example includes a video encoder 2600, a residual generation unit 2602, an analog encoder 2604, a reliability sorting unit 2606, an interleaving unit 2608, a channel encoder 2610, and a punching unit 2612. The video encoder 2600 can acquire video data and... Figure 2 The video encoder 210 operates in roughly the same way. For example... Figure 26 As shown in the example, the video encoder 2600 can receive values of encoding parameters (e.g., encoding selection parameters) sent by the receiving device 104. In some examples, the video encoder 2600 can send values of encoding parameters and / or encoding selection data to the receiving device 104. In this way, the video encoder 2600 of the transmitting device 102, the video encoder of the receiving device 104, and the video decoder of the receiving device 104 can operate based on the same encoding parameter values.
[0304] The video encoder 2600 can also output predicted data to the residual generation unit 2602. Furthermore, the video encoder 2600 can apply a higher level of quantization than the video encoder 210. The residual generation unit 2602 can generate residual data based on the predicted data and the video data.
[0305] The analog encoder 2604 can perform analog encoding operations on residual data. Examples of analog encoding operations can be found in the following documents: U.S. Patent 11,553,184, filed December 29, 2020, entitled "Hybrid Digital-Analog Modulation for Transmission of Video Data"; U.S. Patent 11,431,962, filed December 29, 2020, entitled "Analog Modulated Video Transmission with Variable Symbol Rate"; and U.S. Patent 11,457,224, filed December 29, 2020, entitled "Interlaced Coefficients in Hybrid Digital-Analog Modulation for Transmission of Video Data".
[0306] For example, in some examples, the analog encoder 2604 can generate coefficients based on residual data. For example, the analog encoder 2604 can binarynize the residual data to generate coefficients. The analog encoder 2604 can quantize the coefficients. In other examples of generating coefficients based on video data, the analog encoder 2604 can perform more, fewer, or different steps. For example, in some examples, the analog encoder 2604 does not perform a quantization step. In other examples, the analog encoder 2604 does not perform the step of binarynifying the residual data.
[0307] Furthermore, the analog encoder 2604 can generate coefficient vectors. Each coefficient vector includes n coefficients. The analog encoder 2604 can generate coefficient vectors in one of several ways. For example, in one example, the analog encoder 2604 can generate a coefficient vector as a set of n consecutive coefficients based on the coefficient decoding order. Various coefficient decoding orders can be used, such as raster scan order, sawtooth scan order, reverse raster scan order, and vertical scan order. In some examples, the coefficient vector may include one or more negative coefficients and one or more positive coefficients (i.e., signed coefficients). In some examples, the coefficient vector only includes non-negative coefficients (i.e., unsigned coefficients).
[0308] For each coefficient vector, the analog encoder 2604 can determine the amplitude value of the coefficient vector based on a mapping pattern. For each corresponding allowed coefficient vector among a plurality of allowed coefficient vectors, the mapping pattern maps the corresponding allowed coefficient vector to a corresponding amplitude value among a plurality of amplitude values. This corresponding amplitude value is adjacent to at least one other amplitude value among a plurality of amplitude values in n-dimensional space, and this at least one other amplitude value is adjacent to the corresponding amplitude value in the monotonic number line of the amplitude values.
[0309] In some examples, to determine the amplitude value of the coefficient vector, the analog encoder 2604 can determine the position in n-dimensional space. The coordinates of the position in n-dimensional space are based on the coefficients of the coefficient vector, and a mapping pattern maps different positions in n-dimensional space to different amplitude values among multiple amplitude values. The analog encoder 2604 can determine the amplitude value of the coefficient vector as the amplitude value corresponding to the determined position in n-dimensional space.
[0310] Analog encoder 2604 can modulate an analog signal based on the amplitude values of a coefficient vector. For example, analog encoder 2604 can determine an analog symbol based on a pair of amplitude values. The analog symbol can correspond to the phase shift and power of a point in the IQ plane with coordinates indicated by the amplitude value pair. Analog encoder 2604 can modulate the analog signal during the symbol sampling time based on the determined phase shift and power. The modem (e.g., communication interface 118) of transmitting device 102 can be configured to output an analog signal.
[0311] In addition, Figure 26In example A, the reliability sequencing unit 2606 can obtain encoded video data generated by the video encoder 2600. The reliability sequencing unit 2606 can obtain reliability side information from the video encoder 2600. In some examples, the reliability sequencing unit 2606 can receive values of channel and compression state feedback (CCSF) parameters. The values of the CCSF parameters can provide information about the conditions of channel 230 (e.g., signal-to-noise ratio, latency, network bandwidth congestion, etc.). In some examples, the values of the CCSF parameters provide information related to predicted reliability and quality. For example, the values of the CCSF parameters may include a decimation mode indicator. In some examples, the values of the CCSF parameters enable the channel encoder 2610 to determine the decimation mode.
[0312] Interleaving unit 2608 can perform an interleaving process that ensures reliable and unreliable bits are evenly distributed across code blocks. For example, encoded video data can be divided into code blocks. Channel encoder 2610 can generate a separate error correction dataset for each code block. Before channel encoder 2610 generates error correction data, interleaving unit 2608 can interleave encoded video data between code blocks according to a predefined interleaving pattern. For example, encoded video data representing different adjacent pixels can be interleaved into different code blocks. The deinterleaving process performed at receiving device 104 reverses the interleaving process after the channel decoding process is applied. Therefore, if one of the code blocks is corrupted during transmission, the pixels decoded from the corrupted code block can be spatially distributed within the image among the pixels decoded from the uncorrupted code blocks.
[0313] The channel encoder 2610 of the transmitting device 102 can perform a channel coding process on the video data obtained from the interleaving unit 2608. The channel encoder 2610 can, according to the information provided by the channel encoder 212 (…), Figure 2 The channel coding process can be performed using any example provided by the channel encoder 2610. The puncturing unit 2612 can perform bit puncturing operations on the error-correcting data generated ... according to the information provided by the puncturing unit 214 (…). Figure 2 Any example provided performs bit punching on the error correction data. Transmitting device 102 may transmit encoded video data and error correction data (e.g., bit-punched error correction data) to receiving device 104 via channel 230.
[0314] Figure 27 This is a block diagram illustrating example alternative components of a receiving device 104 according to one or more technologies of this disclosure. Figure 27 The version of the receiving device 104 shown can be used with Figure 26 The version of the transmitting device 102 shown is compatible. Figure 27 In the example, the receiving device 104 includes an analog decoder 2700, a de-puncturing unit 2702, a channel decoder 2704, a deinterleaving unit 2706, a video decoder 2708, a reconstruction unit 2710, an image estimation unit 2712, a video encoder 2714, a reliability unit 2716, and a feedback unit 2718.
[0315] The analog decoder 2700 can acquire analog encoded video data. The analog decoder 2700 can perform analog decoding operations to reconstruct residual data. Example details of the analog decoding operations can be found in U.S. Patents 11,553,184, 11,431,962, and 11,457,224.
[0316] For example, in some examples, the analog decoder 2700 can determine the amplitude values of multiple coefficient vectors based on the analog signal. For instance, the analog decoder 2700 can determine the phase shift and power at the symbol sampling time of the analog signal. The analog decoder 2700 can determine a point in the IQ plane indicated by the determined phase shift and power. The analog decoder 2700 can then determine the amplitude value pairs as coordinates of the point in the IQ plane.
[0317] For each coefficient vector, the analog decoder 2700 can determine the coefficients in the coefficient vector based on the amplitude values and mapping patterns of the coefficient vector. For each corresponding allowed coefficient vector among a plurality of allowed coefficient vectors, the mapping pattern can map the corresponding allowed coefficient vector to a corresponding amplitude value among a plurality of amplitude values. The corresponding amplitude value is adjacent to at least one other amplitude value among a plurality of amplitude values in n-dimensional space, and the at least one other amplitude value is adjacent to the corresponding amplitude value in the monotonic line of the amplitude values. Each coefficient vector can include n coefficients. The value n can be greater than or equal to 2. In some examples, the analog decoder 2700 can determine the coefficients in the coefficient vector as coordinates of positions in n-dimensional space corresponding to amplitude values. The mapping pattern maps different positions in n-dimensional space to different amplitude values among a plurality of amplitude values. In some examples, the coefficient vector includes one or more negative coefficients and one or more positive coefficients. In other examples, the coefficient vector may include only non-negative coefficients.
[0318] In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain a sign value, where the sign value indicates the positive or negative sign of the coefficients in the coefficient vector. In such examples, the analog decoder 2700 may determine the absolute value of the coefficients in the coefficient vector based on the magnitude value of the coefficient vector and the mapping pattern. The analog decoder 2700 may reconstruct the coefficients in the coefficient vector at least partially by applying the sign value to the absolute value of the coefficients in the coefficient vector. In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain data representing shift values. In such examples, the shift value indicates the most negative coefficient in the coefficient vector. Furthermore, in such examples, the analog decoder 2700 may determine the intermediate values of the coefficients in the coefficient vector based on the magnitude value of the coefficient vector and the mapping pattern. The analog decoder 2700 may reconstruct the coefficients in the coefficient vector at least partially by adding shift values to each of the intermediate values of the coefficients in the coefficient vector.
[0319] Furthermore, the analog decoder 2700 can generate residual data based on the coefficients in the coefficient vector. For example, in one example, the analog decoder 2700 can dequantize the coefficients of the coefficient vector. In this example, the analog decoder 2700 can perform a debinding process to convert the coefficients into digital sampled values. For example, the analog decoder 2700 can apply an inverse DCT to the coefficients to convert the coefficients into digital sampled values. In this way, the analog decoder 2700 can generate digital residual sampled values.
[0320] The de-puncturing unit 2702 can obtain encoded video data and bit-punctured error correction data. The de-puncturing unit 2702 can apply a de-puncturing process to the bit-punctured error correction data to reconstruct the error correction data. The de-puncturing unit 2702 can, according to other provisions of this disclosure... Figure 2 Any example of the de-drilling unit 220 to apply the de-drilling process.
[0321] Channel decoder 2704 can perform a channel decoding process that modifies encoded video data (e.g., encoded video data received via channel 230 or encoded video data generated by video encoder 2714 and modified by reliability unit 714 in some examples) based on error correction data. Channel decoder 2704 can be configured according to other parts of this disclosure. Figure 2 Any example of channel decoder 222 to perform the channel decoding process.
[0322] The deinterleaving unit 2706 can perform deinterleaving operations on the error-corrected coded video data generated by the channel decoder 2704. For example, the deinterleaving process can be the reverse of the interleaving process performed by the interleaving unit 2608 of the transmitting device 102. For example, the deinterleaving process can be performed according to the interleaving mode used by the interleaving unit 2608.
[0323] Video decoder 2708 can acquire encoded video data (e.g., deinterleaved encoded video data generated by deinterleaving unit 2706). Video decoder 2708 can perform a video decoding process on the encoded video data to reconstruct a picture of the video data. The video decoding process performed by video decoder 2708 can be the same as the process described in any other example provided elsewhere in this disclosure concerning video decoder 224. Reconstruction unit 2710 of receiving device 104 can add residual data generated by analog decoder 2700 to the corresponding sample of the reconstructed video data generated by video decoder 2708, thereby completely reconstructing a picture of the video data.
[0324] In addition, Figure 27 In the example, image estimation unit 2712 may estimate one or more images based on previously reconstructed images. As previously mentioned in this disclosure, the discussion of images can be applied to image segments, such as slices. Image estimation unit 2712 may estimate images based on any examples of image estimation unit 226 provided elsewhere in this disclosure. Video encoder 2714 may perform a video encoding process on the estimated images. As part of performing the video encoding process, video encoder 2714 may determine encoding selection data, such as encoding selection data 2100, as previously described. Receiving device 104 may send the encoding selection data to transmitting device 102. In some examples, video encoder 2714 may send encoding parameters to transmitting device 102, such as regarding... Figure 4 and Figure 5 The encoding parameters discussed enable the video encoder 2600 of transmitting device 102 to perform a finite encoding process in the same manner as the video encoder 2714 of receiving device 104. In some examples, the video encoder 2714 may determine a multiview encoding cue and send the multiview encoding cue to transmitting device 102.
[0325] Reliability unit 2716 can be used in conjunction with reliability unit 1002 ( Figure 10 It operates in roughly the same manner. The feedback unit 2718 can send predicted quality feedback (e.g., CCSF parameters) to the transmitting device 102 based on the output of the reliability unit 1002.
[0326] In some examples of this disclosure, the video encoder 2600 of the transmitting device 102 generates predicted data for a set of images, and the residual generation unit 2602 can generate residual data based on the first predicted data and the first set of images. The video encoder 2600 can apply a transform to the predicted data to generate transform blocks, transform coefficients for quantizing the transform blocks, and apply entropy coding to syntax elements representing the quantized transform coefficients to generate entropy-coded syntax elements. The first encoded video data may include entropy-coded syntax elements. The channel encoder 2610 can perform a channel coding process that generates error-correcting data for encoding the video data (including entropy-coded syntax elements). The analog encoder 2604 can perform analog modulation on the residual data to generate analog-modulated residual data. The communication interface of the transmitting device 102 can transmit the analog-modulated residual data, the error-correcting data, and the encoded video data.
[0327] The communication interface of receiving device 104 can receive analog modulated residual data and receive error correction data and decimated video data from transmitting device. Decimated video data may include coded video data with an applied decimation mode. The coded video data is generated from a set of images based on the video data. Channel decoder 2704 can apply an error correction process based on the coded video data and error correction data to generate error-corrected coded video data. Error-corrected coded video data includes entropy-coded syntax elements representing quantization transform coefficients. Video decoder 2708 can perform a decoding process to reconstruct a second set of images based on the error-corrected coded video data. As part of the decoding process to reconstruct the image set, video decoder 2708 can apply entropy decoding to the syntax elements to obtain quantization transform coefficients, inverse quantize the quantization transform coefficients to generate inverse quantization transform coefficients, and apply the inverse transform to the inverse quantization transform coefficients to generate prediction data. Analog decoder 2700 can demodulate the analog modulated residual data to obtain residual data. Reconstruction unit 2710 can reconstruct the image set based on the prediction data and residual data.
[0328] The following is a non-limiting list of one or more technologies under the terms of this disclosure.
[0329] Clause 1A. A method for decoding video data, the method comprising: obtaining error correction data at a receiving device and from a transmitting device, wherein the error correction data provides error correction information and is generated from coded video data of one or more blocks of a picture of the video data; generating prediction data of the picture at the receiving device using one or more decoding tools not used to generate the coded video data of one or more blocks, wherein the prediction data of the picture includes predictions of blocks of the picture based at least partially on one or more previously reconstructed blocks of the picture from the video data; generating coded video data at the receiving device based on the prediction data of the picture; generating error-corrected coded video data at the receiving device using the error correction data to perform an error correction operation on the coded video data; and performing a reconstruction operation at the receiving device, the reconstruction operation reconstructing blocks of the picture based on the error-corrected coded video data, wherein the reconstruction operation is controlled by values of one or more parameters.
[0330] Clause 2A. The method described in Clause 1A further includes receiving the value of the parameter at the receiving device and from the transmitting device.
[0331] Clause 3A. The method described in Clause 1A further includes determining the value of the parameter at the receiving device, without receiving the value of the parameter from the transmitting device.
[0332] Clause 4A. The method according to any one of Clauses 1A-3A, wherein: the parameters include one or more quantization parameters, generating coded video data includes using the quantization parameters to quantize transform coefficients generated from image-based prediction data, and performing the reconstruction operation includes using the quantization parameters to inverse quantize the transform coefficients of the error-correcting coded video data.
[0333] Clause 5A. The method of Clause 4A, wherein the method further comprises: calculating a quantization parameter based on the entropy ratio of the quantized transform coefficients to the unquantized transform coefficients.
[0334] Clause 6A. The method according to any one of Clauses 1A-5A, wherein: the parameters include a transform size parameter, generating coded video data includes applying a forward transform having a transform size indicated by the transform size parameter to sampled domain data of a picture, and performing a reconstruction operation includes applying an inverse transform having a transform size indicated by the transform size parameter to transform coefficients of the error-correcting coded video data.
[0335] Clause 7A. The method according to any one of Clauses 1A-6A, wherein: the parameters include parameters indicating the number of transform coefficients; generating coded video data includes including a set of transform coefficients in the coded video data, the set of transform coefficients including the number of transform coefficients indicated; and performing the reconstruction operation includes parsing the set of transform coefficients including the number of transform coefficients from the error-corrected coded video data.
[0336] Clause 8A. The method according to Clause 7A, wherein: obtaining coded video data and error correction data includes receiving coded video data and error correction data from a transmitting device via a communication channel at a receiving device, and the method further includes applying an optimization process that determines a plurality of transform coefficients based on the signal-to-noise ratio of the data transmitted over the communication channel.
[0337] Clause 9A. The method according to any one of Clauses 1A-8A, wherein: the parameters include a bit width parameter of a plurality of index values, and for each corresponding index value among the plurality of index values: performing a reconstruction operation includes parsing a first bit set from error-corrected coded video data, wherein the first bit set indicates transform coefficients having the corresponding index value, and the number of bits in the first bit set is equal to the bit width indicated by the bit width parameter of the corresponding index value; generating coded video data includes including a second bit set in the coded video data, wherein the second bit set indicates transform coefficients having the corresponding index value, and the number of bits in the second bit set is equal to the bit width indicated by the bit width parameter of the corresponding index value; and parsing a third bit set from the error-corrected coded video data, wherein the third bit set indicates transform coefficients having the corresponding index value, and the number of bits in the third bit set is equal to the bit width indicated by the bit width parameter of the corresponding index value.
[0338] Clause 10A. The method according to any one of Clauses 1A-9A, wherein the method further comprises: performing a bit de-puncturing operation on the error correction data before generating the error correction coded video data.
[0339] Clause 11A. The method according to any one of Clauses 1A-10A, wherein the parameters include one or more of the following: color space, transform size, quantization parameter, number of transform coefficients in the first coded video data, or number of bits per transform coefficient of the first coded video data.
[0340] Clause 12A. The method according to any one of Clauses 1A-11A, wherein: the extraction mode defines the patterns of anchor transform blocks and non-anchor transform blocks in the picture, the method further comprising: receiving at a receiving device systematic bits of the anchor transform blocks instead of systematic bits in the non-anchor transform blocks, the systematic bits in the anchor transform blocks representing transform coefficients in the anchor transform blocks, and the systematic bits in the non-anchor transform blocks representing reduced bit-depth versions of the original transform coefficients in the non-anchor transform blocks; the error correction data includes error correction data of the anchor transform blocks and error correction data of the non-anchor transform blocks, wherein the error correction data of the non-anchor transform blocks is based on the original transform coefficients in the non-anchor transform blocks, and generating error-corrected coded video data includes: using the error correction data of the anchor transform blocks to correct the systematic bits of the anchor transform blocks; and using the error correction data of the non-anchor transform blocks to perform error correction on the portion of the coded video data corresponding to the non-anchor transform blocks.
[0341] Clause 13A. The method described in Clause 12A further includes: determining a decimation mode at the receiving device; and sending the decimation mode to the transmitting device from the receiving device.
[0342] Clause 14A. The method according to any one of Clauses 1A-13A, wherein: the extracted pattern defines the patterns of anchor transform blocks and non-anchor transform blocks in an image, the transform coefficients in the non-anchor transform blocks having a reduced bit depth relative to the anchor transform blocks; the systematic bits of the anchor transform blocks, the systematic bits of the non-anchor transform blocks, and a correlation matrix are received at the receiving device, the systematic bits of the anchor transform blocks representing the transform coefficients in the anchor transform blocks, and the systematic bits of the non-anchor transform blocks representing a reduced bit depth version of the original transform coefficients in the non-anchor transform blocks; the reconstruction operation includes: for each non-anchor transform coefficient in the non-anchor transform blocks: calculating an interpolated value of the non-anchor transform coefficient at the receiving device based on the correlation matrix and the corresponding anchor transform coefficient; and calculating a reconstructed value of the non-anchor transform coefficient at the receiving device based on the interpolated value of the non-anchor transform coefficient and the value of the non-anchor transform coefficient in the error-corrected coded video data.
[0343] Clause 15A. A method for encoding video data, the method comprising: obtaining video data from a video source at a transmitting device; generating encoded video data of a first image of the video data and encoded video data of a second image of the video data at the transmitting device based on a set of parameters; performing channel coding on the encoded video data of the first image and the encoded video data of the second image at the transmitting device to generate error-corrected data of the first image and error-corrected data of the second image; and transmitting the encoded video data of the first image, the error-corrected data of the first image, and the error-corrected data of the second image at the transmitting device.
[0344] Clause 16A. The method described in Clause 15A further includes sending the value of the parameter from the transmitting device to the receiving device.
[0345] Clause 17A. The method according to Clause 16A, wherein: the parameters include one or more quantization parameters, and generating coded video data includes using the quantization parameters to quantize the transform coefficients of the first image and the transform coefficients of the second image.
[0346] Clause 18A. The method according to any one of Clauses 16A-17A, wherein: the parameters include a transform size parameter, and generating coded video data includes applying a forward transform to residual data blocks of a first picture and residual data blocks of a second picture, wherein the forward transform has a transform size indicated by the transform size parameter.
[0347] Clause 19A. The method according to any one of Clauses 16A-18A, wherein: the parameters include parameters indicating the number of transform coefficients, and generating coded video data of the first picture and coded video data of the second picture includes a set of transform coefficients in the coded video data of the first picture and the coded video data of the second picture, the set of transform coefficients including transform coefficients indicating the number.
[0348] Clause 20A. The method according to any one of Clauses 16A-10A, wherein the parameters include one or more of the following: color space, transform size, quantization parameters, the number of transform coefficients in the encoded video data, or the number of bits per transform coefficient of the encoded video data.
[0349] Clause 21A. A method for encoding video data, the method comprising: obtaining video data from a video source at a transmitting device; generating transform blocks based on the video data at the transmitting device; determining which of the transform blocks are anchor transform blocks at the transmitting device; calculating a correlation matrix of the set of transform blocks at the transmitting device; generating a bit-reduced non-anchor transform matrix at the transmitting device; and transmitting the anchor transform blocks, non-anchor transform blocks, and correlation matrix to a receiving device at the transmitting device.
[0350] Clause 22A. The method according to Clause 21A further includes receiving an indication of the extraction mode from the receiving device at the transmitting device.
[0351] Clause 23A. An apparatus comprising: a memory configured to store video data; a communication interface; and one or more processors implemented in a circuit and coupled to the memory, the one or more processors being configured to perform a method according to any one of Clauses 1A-22A.
[0352] Clause 24A. An apparatus comprising components for performing the method according to any one of Clauses 1A-22A.
[0353] Clause 25A. A computer-readable data storage medium having instructions stored thereon that, when executed, cause a device to perform the method according to any one of Clauses 1A-22A.
[0354] Clause 1B. An apparatus for processing video data, the apparatus comprising: a memory configured to store video data; and a communication interface configured to receive error correction data from a transmitting device, wherein the error correction data provides error correction information about pictures of the video data; one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: generate prediction data for pictures, wherein the prediction data for pictures includes predictions of blocks of pictures based at least in part on one or more previously reconstructed pictures of the video data; generate coded video data based on the prediction data for pictures, wherein the coded video data includes transform blocks, the transform blocks including transform coefficients; scale the bits of the transform coefficients of the transform blocks based on reliability values of bit positions; generate error-corrected coded video data using the error correction data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstruct pictures based on the error-corrected coded video data.
[0355] Clause 2B. The device as described in Clause 1B, wherein one or more processors are further configured to generate a reliability value at the receiving device.
[0356] Clause 3B. The device as described in Clause 2B, wherein one or more processors are configured to generate a reliability value based on statistics regarding the occurrence of errors in the said bit positions.
[0357] Clause 4B. An apparatus pursuant to any one of Clauses 2B or 3B, wherein one or more processors are configured to generate reliability values based on the reliability characteristics of various regions of a picture of video data.
[0358] Clause 5B. An apparatus pursuant to any one of Clauses 2B-4B, wherein one or more processors are configured to generate reliability values based on a noise model.
[0359] Clause 6B. The device according to any one of Clauses 1B-5B, wherein the communication interface is further configured to send a reliability value to the transmitting device.
[0360] Clause 7B. The device according to any one of Clauses 1B-5B, wherein the communication interface is further configured to receive reliability values from the transmitting device.
[0361] Clause 8B. An apparatus for processing video data, the method comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: acquire the video data; acquire predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated picture generated by a receiving device; adjust one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; perform a video coding process to generate coded video data based on one or more pictures of the acquired video data, wherein the video coding process is controlled by the video coding parameters; perform a channel coding process on the coded video data to generate channel-coded data, wherein the channel coding process is controlled by the channel coding parameters; and a communication interface configured to transmit the channel-coded data to a receiving device.
[0362] Clause 9B. The apparatus as described in Clause 8B, wherein: the video encoding parameters include quantization parameters, one or more processors are configured to adjust the quantization parameters as part of adjusting the video encoding / decoding parameters, and one or more processors are configured to use the quantization parameters to quantize the transform coefficients of transform blocks of one or more pictures as part of performing the video encoding process.
[0363] Clause 10B. An apparatus according to any one of Clauses 8B-9B, wherein: the channel coding parameters include a low-density parity-check (LDPC) graph, one or more processors are configured to adjust the LDPC graph as part of adjusting the channel encoder parameters, and one or more processors are configured to use the LDPC graph to generate codewords included in the channel coding data as part of performing the channel coding process.
[0364] Clause 11B. The apparatus according to any one of Clauses 8B-10B, wherein: the channel-coded data includes error-correcting data, one or more processors are further configured to adjust one or more bit-puncturing parameters based on prediction quality feedback, and one or more processors are configured to perform a bit-puncturing process on the error-correcting data, wherein the bit-puncturing process is controlled by one or more bit-puncturing parameters.
[0365] Clause 11B. A method of processing video data, the method comprising: obtaining error correction data at a receiving device and from a transmitting device, wherein the error correction data provides error correction information about a picture of the video data; generating prediction data of the picture at the receiving device, wherein the prediction data of the picture includes predictions of blocks of the picture based at least in part on one or more previously reconstructed pictures of the video data; generating coded video data at the receiving device based on the prediction data of the picture, wherein the coded video data includes transform blocks, the transform blocks including transform coefficients; scaling bits of the transform coefficients of the transform blocks at the receiving device based on reliability values of bit positions; generating error-corrected coded video data at the receiving device using the error correction data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstructing a picture at the receiving device based on the error-corrected coded video data.
[0366] Clause 12B. The method described in Clause 11B further includes generating a reliability value at the receiving device.
[0367] Clause 13B. The method according to Clause 12B, wherein generating a reliability value includes generating a reliability value at the receiving device based on statistics about the occurrence of errors at bit positions.
[0368] Clause 14B. The method according to any one of Clauses 12B or 13B, wherein generating a reliability value includes generating a reliability value at the receiving device based on the reliability characteristics of various regions of the image of the video data.
[0369] Clause 15B. The method according to any one of Clauses 12B-14B, wherein generating the reliability value includes generating the reliability value at the receiving device based on a noise model.
[0370] Clause 16B. The method according to any one of Clauses 11B-15B further includes transmitting a reliability value from the receiving device to the transmitting device.
[0371] Clause 17B. The method according to any one of Clauses 11B-15B further includes receiving a reliability value from the transmitting device at the receiving device.
[0372] Clause 18B. A method for processing video data, the method comprising: acquiring video data; acquiring predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated picture generated by a receiving device; adjusting one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process to generate coded video data based on one or more pictures of the acquired video data, wherein the video coding process is controlled by the video coding parameters; performing a channel coding process on the coded video data to generate channel-coded data, wherein the channel coding process is controlled by the channel coding parameters; and transmitting the channel-coded data to a receiving device.
[0373] Clause 19B. The method according to Clause 18B, wherein: the video coding parameters include quantization parameters, adjusting the video coding parameters includes adjusting the quantization parameters, and performing the video coding process includes using the quantization parameters to quantize the transform coefficients of one or more transform blocks of images.
[0374] Clause 20B. The method according to any one of Clauses 18B-19B, wherein: the channel coding parameters include a low-density parity-check (LDPC) graph, adjusting the channel encoder parameters includes adjusting the LPPC graph, and performing the channel coding process includes using the LDPC graph to generate codewords included in the channel coded data.
[0375] Clause 21B. The method according to any one of Clauses 18B-20B, wherein: the channel-coded data includes error-correcting data, and the method further comprises: adjusting one or more bit puncturing parameters based on prediction quality feedback, and performing a bit puncturing process on the error-correcting data, wherein the bit puncturing process is controlled by one or more bit puncturing parameters.
[0376] Clause 22B. An apparatus comprising components for performing the method according to any one of Clauses 11B-21B.
[0377] Clause 23B. A computer-readable data storage medium having instructions stored thereon that, when executed, cause a device to perform the method according to any one of Clauses 11B-21B.
[0378] Clause 1C. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: acquire a first multiview image set of the video data, wherein the first multiview image set includes a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; transmit first encoded video data to a receiving device, wherein the first encoded video data is based on the first multiview image set; receive a multiview encoding cue from the receiving device; acquire a second multiview image set of the video data, wherein the second multiview image set includes a third image and a fourth image, the third image being from the first viewpoint and the fourth image being from the second viewpoint; perform a multiview encoding process on the second multiview image set based on the multiview encoding cue received from the receiving device to generate second encoded video data, wherein the multiview encoding process reduces interview redundancy between the third and fourth images; and transmit the second encoded video data to the receiving device.
[0379] Clause 2C. The apparatus according to Clause 1, wherein one or more processors are further configured to: receive an updated multiview coding cue from the receiving device after sending second coded video data to the receiving device; obtain a third multiview image set of the video data, wherein the third multiview image set includes a fifth image and a sixth image, the fifth image being from a first viewpoint and the sixth image being from a second viewpoint; encode the third multiview image set based on the updated multiview coding cue received from the receiving device to generate third coded video data; and send the third coded video data to the receiving device.
[0380] Clause 3C. The device according to any one of Clauses 1C-2C, wherein the multi-view coding cue includes one or more of the following: relative shift between blocks of the first and second images, brightness correction between the first and second images, inter-block shift between anchor blocks and reconstructed blocks, or motion data for reference shift.
[0381] Clause 4C. A device according to any one of Clauses 1C-3C, wherein: the device is an extended reality (XR) headset, and one or more processors are configured to: receive virtual element data generated from a receiving device based on a first multiview image set and a second multiview image set; and output the virtual element data for display in an XR scene.
[0382] Clause 5C. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: obtain first encoded video data from a transmitting device, wherein the first encoded video data is based on a first multiview image set of the video data, the first multiview image set including a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; determine a multiview encoding cue based on the first encoded video data; transmit the multiview encoding cue to the transmitting device; and obtain second encoded video data from the transmitting device, wherein the second encoded video data is based on a second multiview image set including a third image and a fourth image, the second encoded video data being encoded using a multiview encoding process that reduces interview redundancy between the third and fourth images based on the multiview encoding cue.
[0383] Clause 6C. The apparatus described in Clause 5C, wherein one or more processors are further configured to decode second coded video data.
[0384] Clause 7C. The device according to any one of Clauses 5C-6C, wherein the multiview coding cue is a first multiview coding cue, and one or more processors are further configured to: determine a second multiview coding cue based on second coded video data; send the second multiview coding cue to a transmitting device; and obtain third coded video data from the transmitting device, wherein the third coded video data is based on a third multiview image set including a fifth image and a sixth image, and the third coded video data is encoded using a multiview coding process that reduces interview redundancy between the fifth and sixth images based on the second multiview coding cue.
[0385] Clause 8C. The device according to any one of Clauses 5C-7C, wherein: the multiview encoded cue includes a depth map indicating the depth of an object represented in a first picture and a second picture, and as part of determining the multiview encoded cue, one or more processors are configured to determine the depth map based on the first picture and the second picture.
[0386] Clause 9C. The apparatus according to Clauses 5C-8C, wherein the multiview coded cue includes one or more illumination compensation factors, and as part of determining the multiview coded cue, one or more processors are configured to determine the illumination compensation factors based on a first picture and a second picture.
[0387] Clause 10C. An apparatus pursuant to any one of Clauses 5C-9C, wherein: the transmitting device is an extended reality (XR) headset, and one or more processors are further configured to: process a second set of images to generate virtual element data; and transmit the virtual element data to the XR headset.
[0388] Clause 11C. A method for processing video data, the method comprising: obtaining a first multiview image set of the video data, wherein the first multiview image set includes a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first multiview image set; receiving a multiview encoding cue from the receiving device; obtaining a second multiview image set of the video data, wherein the second multiview image set includes a third image and a fourth image, the third image being from the first viewpoint and the fourth image being from the second viewpoint; performing a multiview encoding process on the second multiview image set based on the multiview encoding cue received from the receiving device to generate second encoded video data, wherein the multiview encoding process reduces interview redundancy between the third image and the fourth image; and transmitting the second encoded video data to the receiving device.
[0389] Clause 12C. The method according to Clause 11C further includes: receiving an updated multiview coding cue from the receiving device after sending the second coded video data to the receiving device; obtaining a third multiview image set of the video data, wherein the third multiview image set includes a fifth image and a sixth image, the fifth image being from a first viewpoint and the sixth image being from a second viewpoint; encoding the third multiview image set based on the updated multiview coding cue received from the receiving device to generate third coded video data; and sending the third coded video data to the receiving device.
[0390] Clause 13C. The method according to any one of Clauses 11C-12C, wherein the multi-view coding cue includes one or more of the following: relative shift between blocks of the first and second images, brightness correction between the first and second images, inter-block shift between anchor blocks and reconstructed blocks, or motion data for reference shift.
[0391] Clause 14C. The method according to any one of Clauses 11C-13C, wherein the method further comprises: receiving from a receiving device virtual element data generated based on the first multiview image set and the second multiview image set; and outputting virtual element data for display in an extended reality (XR) scene.
[0392] Clause 15C. A method for processing video data, the method comprising: obtaining first coded video data from a transmitting device, wherein the first coded video data is based on a first multiview image set of the video data, the first multiview image set including a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; determining a multiview coding cue based on the first coded video data; sending the multiview coding cue to the transmitting device; and obtaining second coded video data from the transmitting device, wherein the second coded video data is based on a second multiview image set including a third image and a fourth image, the second coded video data being encoded using a multiview coding process that reduces interview redundancy between the third image and the fourth image based on the multiview coding cue.
[0393] Clause 16C. The method described in Clause 15C further includes decoding the second coded video data.
[0394] Clause 17C. The method according to any one of Clauses 15C-16C, wherein the multiview coding cue is a first multiview coding cue, and the method further comprises: determining a second multiview coding cue based on second coded video data; sending the second multiview coding cue to a transmitting device; and obtaining third coded video data from the transmitting device, wherein the third coded video data is based on a third multiview image set including a fifth image and a sixth image, and the third coded video data is encoded using a multiview coding process that reduces interview redundancy between the fifth and sixth images based on the second multiview coding cue.
[0395] Clause 18C. The method according to any one of Clauses 15C-17C, wherein: the multiview encoded cue includes a depth map indicating the depth of an object represented in a first picture and a second picture, and determining the multiview encoded cue includes determining the depth map based on the first picture and the second picture.
[0396] Clause 19C. The method according to Clauses 15C-18C, wherein the multiview coding cue includes one or more lighting compensation factors, and determining the multiview coding cue includes determining the lighting compensation factors based on a first picture and a second picture.
[0397] Clause 20C. The method according to any one of Clauses 15C-19C, wherein: the transmitting device is an extended reality (XR) headset, and the method further comprises: processing a second set of images to generate virtual element data; and transmitting the virtual element data to the XR headset.
[0398] Clause 21C. An apparatus comprising: means for acquiring a first multiview image set of video data, wherein the first multiview image set includes a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; means for transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first multiview image set; means for receiving a multiview encoding cue from the receiving device; means for acquiring a second multiview image set of video data, wherein the second multiview image set includes a third image and a fourth image, the third image being from the first viewpoint and the fourth image being from the second viewpoint; means for performing a multiview encoding process on the second multiview image set based on the multiview encoding cue received from the receiving device to generate second encoded video data, wherein the multiview encoding process reduces interview redundancy between the third image and the fourth image; and means for transmitting the second encoded video data to the receiving device.
[0399] Clause 22C. An apparatus comprising: means for obtaining first coded video data from a transmitting device, wherein the first coded video data is based on a first multiview image set of the video data, the first multiview image set including a first image and a second image, the first image being from a first viewpoint and the second image being from a second viewpoint; means for determining a multiview coding cue based on the first coded video data; means for sending the multiview coding cue to the transmitting device; and means for obtaining second coded video data from the transmitting device, wherein the second coded video data is based on a second multiview image set including a third image and a fourth image, the second coded video data being encoded using a multiview coding process that reduces interview redundancy between the third image and the fourth image based on the multiview coding cue.
[0400] Clause 1D. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: encode a first set of images of the video data to generate first encoded video data; transmit the first encoded video data to a receiving device; receive from the receiving device a decimation mode indication indicating a decimation mode determined based on the first set of images, the decimation mode being a mode in which the encoded video data is not transmitted; encode a second set of images of the video data to generate second encoded video data; apply the decimation mode to the second encoded video data to generate decimated video data; and transmit the decimated video data to the receiving device.
[0401] Clause 2D. The apparatus according to Clause 1D, wherein one or more processors are configured to: generate first error correction data based on first coded video data; transmit the first error correction data to a receiving device; generate second error correction data based on second coded video data; and transmit the second error correction data to the receiving device.
[0402] Clause 3D. The device according to any one of Clauses 1D-2D, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for the complete picture.
[0403] Clause 4D. The device according to any one of Clauses 1D-3D, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for a specific region within the picture.
[0404] Clause 5D. The device according to any one of Clauses 1D-4D, wherein the video data is multi-view video data and the extraction mode indicates a mode that skips the transmission of encoded video data of pictures from a particular viewpoint.
[0405] Clause 6D. An apparatus according to any one of Clauses 1D-5D, wherein: a decimation mode indication is a first decimation mode indication, a mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, decimated video data is first decimated video data, and one or more processors are further configured to: encode a third set of images of video data to generate third encoded video data; determine a second decimation mode indicating a second mode in which encoded video data is not transmitted; apply the second decimation mode to the third encoded video data to generate second decimated video data; transmit the second decimated video data to a receiving device; and transmit a second decimation mode indication to the receiving device, the second decimation mode indication indicating that the second decimation mode is applied to the third encoded video data.
[0406] Clause 7D. A method according to any one of Clauses 1D-6D, wherein: as part of encoding a first set of images, one or more processors are configured to: generate first prediction data for the first set of images; generate residual data based on the first prediction data and the first set of images; apply a transform to the first prediction data to generate a transform block; quantize the transform coefficients of the transform block; apply entropy coding to a syntax element representing the quantized transform coefficients to generate a first entropy-coded syntax element, wherein the first encoded video data includes the first entropy-coded syntax element; the one or more processors are further configured to perform analog modulation on the residual data to generate first analog-modulated residual data; and the device further includes a communication interface configured to transmit the first analog-modulated residual data and the first encoded video data.
[0407] Clause 8D. A device according to any one of Clauses 1D-7D, wherein: the device is an extended reality (XR) headset and includes a display system, and one or more processors are further configured to: receive virtual element data from a receiving device; and the display system is configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0408] Clause 9D. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: receive first coded video data from a transmitting device; perform a decoding process to reconstruct a first set of images based on the first coded video data; determine a decimation mode based on the first set of images, the decimation mode indicating a mode in which the coded video data is not transmitted; send a decimation mode indication to the transmitting device indicating the determined decimation mode; receive decimated video data from the transmitting device, wherein the decimated video data includes second coded video data to which the decimation mode has been applied, wherein the second coded video data is generated based on a second set of images of the video data; and perform a decoding process to reconstruct a second set of images based on the second coded video data.
[0409] Clause 10D. The apparatus according to Clause 9D, wherein one or more processors are further configured to: receive first error-correcting data from a transmitting device; apply an error-correcting process to modify first coded video data based on the first error-correcting data to generate first error-corrected coded video data; wherein one or more processors are configured to perform a decoding process to reconstruct a first set of images based on the first error-corrected coded video data; wherein one or more processors are further configured to: receive second error-correcting data from a transmitting device; apply an error-correcting process to generate second error-corrected coded video data based on the second coded video data and the second error-correcting data; and wherein one or more processors are configured to perform a decoding process to reconstruct a second set of images based on the second error-corrected coded video data.
[0410] Clause 11D. The device according to any one of Clauses 9D-10D, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for the complete picture.
[0411] Clause 12D. The apparatus according to any one of Clauses 9D-11D, wherein, as part of determining an extraction mode, one or more processors are configured to: apply the extraction mode to first coded video data to generate decimated coded video data; apply an error correction process to modify the decimated coded video data based on the first error correction data to generate trial error-corrected video data; apply a decoding process to reconstruct a first set of images based on the trial error-corrected video data; and determine whether the extraction mode satisfies the standard by comparing the first set of images reconstructed based on the trial error-corrected video data with the first set of images reconstructed based on the first video data.
[0412] Clause 13D. The apparatus according to any one of Clauses 9D-12D, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for a specific region within a picture.
[0413] Clause 14D. The apparatus according to any one of Clauses 9D-13D, wherein the video data is multi-view video data and the extraction mode indicates a mode that skips the transmission of encoded video data of pictures from a particular viewpoint.
[0414] Clause 15D. An apparatus according to any one of Clauses 9D-14D, wherein: a decimation mode indication is a first decimation mode indication, a mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, decimated video data is first decimated video data, and one or more processors are further configured to: receive a second decimation mode indication indicating a second mode in which encoded video data is not transmitted; receive second decimated video data from a transmitting device, wherein the second decimated video data includes third encoded video data to which the second decimation mode has been applied, wherein the third encoded video data is generated based on a third set of pictures of the video data; and apply a decoding process to reconstruct the third set of pictures based on the third encoded video data.
[0415] Clause 16D. The apparatus according to any one of Clauses 9D-15D, wherein: the apparatus further includes a communication interface configured to receive analog modulation residual data, the second coded video data including entropy-coded syntax elements representing quantization transform coefficients; as part of an application decoding process to reconstruct a second set of pictures, one or more processors are configured to: apply entropy decoding to the syntax elements to obtain quantization transform coefficients; inverse quantize the quantization transform coefficients to generate inverse quantization transform coefficients; apply an inverse transform to the inverse quantization transform coefficients to generate prediction data; demodulate the analog modulation residual data to obtain residual data; and reconstruct the second set of pictures based on the prediction data and the residual data.
[0416] Clause 17D. The device according to any one of Clauses 9D-16D, wherein: one or more processors are further configured to process a second set of images to generate virtual element data, and the transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0417] Clause 18D. A method comprising: encoding a first set of images of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving a decimation mode indication from the receiving device, the decimation mode indication indicating a decimation mode determined based on the first set of images, the decimation mode being a mode in which the encoded video data is not transmitted; encoding a second set of images of video data to generate second encoded video data; applying the decimation mode to the second encoded video data to generate decimated video data; and transmitting the decimated video data to the receiving device.
[0418] Clause 19D. The method according to Clause 18D further includes: generating first error correction data based on first coded video data; sending the first error correction data to a receiving device; generating second error correction data based on second coded video data; and sending the second error correction data to the receiving device.
[0419] Clause 20D. The method according to any one of Clauses 18D-19D, wherein the extraction mode indicates a mode that skips the transmission of encoded video data of the complete picture.
[0420] Clause 21D. The method according to any one of Clauses 18D-20D, wherein the extraction mode indicates a mode for skipping the transmission of encoded video data of a specific region within the picture.
[0421] Clause 22D. The method according to any one of Clauses 18D-21D, wherein the video data is multi-view video data and the extraction mode indicates a mode that skips the transmission of encoded video data of pictures from a particular viewpoint.
[0422] Clause 23D. The method according to any one of Clauses 18D-22D, wherein: the decimation mode indication is a first decimation mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, the decimated video data is first decimated video data, and the method further comprises: encoding a third set of images of video data to generate third encoded video data; determining a second decimation mode indicating a second mode in which encoded video data is not transmitted; applying the second decimation mode to the third encoded video data to generate second decimated video data; transmitting the second decimated video data to a receiving device; and transmitting a second decimation mode indication to the receiving device, the second decimation mode indication indicating that the second decimation mode is applied to the third encoded video data.
[0423] Clause 24D. The method according to any one of Clauses 18D-23D, wherein: encoding the first set of images comprises: generating first prediction data for the first set of images; generating residual data based on the first prediction data and the first set of images; applying a transform to the first prediction data to generate a transform block; quantizing the transform coefficients of the transform block; applying entropy coding to a syntax element representing the quantized transform coefficients to generate a first entropy-coded syntax element, wherein the first encoded video data includes the first entropy-coded syntax element; the method further comprises: performing analog modulation on the residual data to generate first analog-modulated residual data; and transmitting the first analog-modulated residual data and the first encoded video data.
[0424] Clause 25D. The method according to any one of Clauses 18D-24D, wherein the method further comprises: receiving virtual element data from a receiving device; and displaying one or more virtual elements in an extended reality (XR) scene based on the virtual element data.
[0425] Clause 26D. A method comprising: receiving first coded video data from a transmitting device; applying a decoding process to reconstruct a first set of images based on the first coded video data; determining an extraction mode based on the first set of images, the extraction mode indicating a mode in which the coded video data is not transmitted; sending to the transmitting device an extraction mode indication indicating the determined extraction mode; receiving extracted video data from the transmitting device, wherein the extracted video data includes second coded video data to which the extraction mode has been applied, wherein the second coded video data is generated based on a second set of images of the video data; and performing a decoding process to reconstruct a second set of images based on the second coded video data.
[0426] Clause 27D. The method according to Clause 26D, further comprising: receiving first error correction data from a transmitting device; applying an error correction process to modify first coded video data based on the first error correction data to generate first error-corrected coded video data; wherein performing a decoding process to reconstruct the first set of images includes performing a decoding process to reconstruct the first set of images based on the first error-corrected coded video data; wherein the method further comprises: receiving second error correction data from a transmitting device; applying an error correction process to generate second error-corrected coded video data based on the second coded video data and the second error correction data; and wherein performing a decoding process to reconstruct the second set of images includes performing a decoding process to reconstruct the second set of images based on the second error-corrected coded video data.
[0427] Clause 28D. The method according to any one of Clauses 26D-27D, wherein the extraction mode indicates a mode that skips the transmission of encoded video data of the complete picture.
[0428] Clause 29D. The method according to any one of Clauses 26D-28D, wherein determining the extraction pattern comprises: applying the extraction pattern to first coded video data to generate decimated coded video data; applying an error correction process to modify the decimated coded video data based on the first error correction data to generate experimental error-corrected video data; applying a decoding process to reconstruct a first set of images based on the experimental error-corrected video data; and determining whether the extraction pattern satisfies the standard by comparing the first set of images reconstructed based on the experimental error-corrected video data and the first set of images reconstructed based on the first video data.
[0429] Clause 30D. The method according to any one of Clauses 26D-29D, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for a specific region within the picture.
[0430] Clause 31D. The method according to any one of Clauses 26D-30D, wherein the video data is multi-view video data and the extraction mode indicates a mode that skips the transmission of encoded video data of pictures from a particular viewpoint.
[0431] Clause 32D. The method according to any one of Clauses 26D-31D, wherein: the decimation mode indication is a first decimation mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, the decimated video data is first decimated video data, and the method further comprises: receiving a second decimation mode indication indicating a second mode in which encoded video data is not transmitted; receiving second decimated video data from a transmitting device, wherein the second decimated video data includes third encoded video data to which the second decimation mode has been applied, wherein the third encoded video data is generated based on a third set of images of the video data; and applying a decoding process to reconstruct the third set of images based on the third encoded video data.
[0432] Clause 33D. The method according to any one of Clauses 26D-32D, wherein: the method further includes receiving analog modulation residual data, the second coded video data including entropy-coded syntax elements representing quantization transform coefficients; applying a decoding process to reconstruct the second set of pictures including: applying entropy decoding to the syntax elements to obtain quantization transform coefficients; inverse quantizing the quantization transform coefficients to generate inverse quantization transform coefficients; applying an inverse transform to the inverse quantization transform coefficients to generate prediction data; demodulating the analog modulation residual data to obtain residual data; and reconstructing the second set of pictures based on the prediction data and the residual data.
[0433] Clause 34D. The method according to any one of Clauses 26D-33D, wherein: the method further comprises processing a second set of images to generate virtual element data, and the transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0434] Clause 35D. An apparatus comprising: means for encoding a first set of images of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device a decimation mode indication indicating a decimation mode determined based on the first set of images, the decimation mode being a mode in which the encoded video data is not transmitted; means for encoding a second set of images of video data to generate second encoded video data; means for applying the decimation mode to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to the receiving device.
[0435] Clause 36D. An apparatus comprising: means for receiving first coded video data from a transmitting device; means for performing a decoding process to reconstruct a first set of pictures based on the first coded video data; means for determining a decimation mode based on the first set of pictures, the decimation mode indicating a mode in which the coded video data is not transmitted; means for sending to the transmitting device a decimation mode indication indicating the determined decimation mode; means for receiving decimated video data from the transmitting device, wherein the decimated video data includes second coded video data to which the decimation mode has been applied, wherein the second coded video data is generated based on a second set of pictures of the video data; and means for performing a decoding process to reconstruct a second set of pictures based on the second coded video data.
[0436] Clause 1E. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: encode a first image of the video data to generate first coded video data; transmit the first coded video data to a receiving device; receive encoding selection data of a second image of the video data from the receiving device, wherein: the encoding selection data of the second image indicates encoding selections for encoding an estimate of the second image, and the second image follows the first image in decoding order; encode the second image based on the encoding selection data of the second image to generate second coded video data; and transmit the second coded video data to the receiving device.
[0437] Clause 2E. The apparatus according to Clause 1E, wherein: the encoded selection data received from the receiving device includes motion parameters of blocks of a second picture, and as part of encoding the second picture, one or more processors are configured to perform motion compensation based on the motion parameters of the blocks of the second picture to generate predicted blocks, and the second encoded video data includes encoded video data based on the predicted blocks.
[0438] Clause 3E. An apparatus pursuant to any one of Clauses 1E-2E, wherein: the encoding selection data received from the receiving apparatus includes intra-prediction parameters of blocks of the second picture, and as part of encoding the second picture, one or more processors are configured to perform intra-prediction based on the intra-prediction parameters of blocks of the second picture to generate prediction blocks, and the second coded video data includes coded video data based on the prediction blocks.
[0439] Clause 4E. The device pursuant to any one of Clauses 1E-3E, wherein the second encoded video data does not include encoding selection data.
[0440] Clause 5E. A device pursuant to any one of Clauses 1E-4E, wherein one or more processors are configured to entropy decode encoded selection data for a prior second picture.
[0441] Clause 6E. The device pursuant to any one of Clauses 1E-5E, wherein one or more processors are further configured to: generate first error correction data based on first coded video data; and transmit the first coded video data and the first error correction data to a receiving device.
[0442] Clause 7E. The device pursuant to any one of Clauses 1E-6E, wherein one or more processors are further configured to: encode the third picture without using the encoding selection data of the third picture, based on the determination that no encoding selection data of the third picture has been received from the receiving device before the time limit expires.
[0443] Clause 8E. The apparatus according to any one of Clauses 1E-7E, wherein one or more processors are further configured to: receive coding selection data of a third picture of video data, wherein the coding selection data of the third picture indicates coding selection for encoding an estimate of the third picture; encode the third picture based on the coding selection data of the third picture to generate third coded video data; apply a channel coding process for generating error correction data of the third coded video data; and transmit the error correction data of the third coded video data to a receiving device without transmitting at least a portion of the third coded video data.
[0444] Clause 9E. A device according to any one of Clauses 1E-8E, wherein: the device is an extended reality (XR) headset and includes a display system, and one or more processors are further configured to: receive virtual element data from a receiving device; and the display system is configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0445] Clause 10E. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the processors being configured to: receive first coded video data from a transmitting device; reconstruct a first picture of the video data based on the first coded video data; estimate a second picture of the video data based on the first picture, the second picture being a picture that appears after the first picture in decoding order; generate encoding selection data for the estimated second picture, wherein the encoding selection data indicates encoding selections for encoding the estimated second picture; transmit the encoding selection data of the second picture to the transmitting device; receive second coded video data from the transmitting device; and reconstruct the second picture based on the second coded video data.
[0446] Clause 11E. The apparatus according to Clause 10E, wherein: as part of encoding an estimated second picture, one or more processors are configured to perform motion compensation based on motion parameters of blocks of the second picture to generate predicted blocks, and encoded selection data includes motion parameters of blocks of the second picture, and second encoded video data includes encoded video data based on predicted blocks.
[0447] Clause 12E. An apparatus pursuant to any one of Clauses 10E-11E, wherein: as part of encoding a second picture, one or more processors are configured to perform intra-prediction based on intra-prediction parameters of blocks of the second picture to generate prediction blocks, encoding selection data includes intra-prediction parameters of blocks of the second picture, and second encoded video data includes encoded video data based on prediction blocks.
[0448] Clause 13E. The device pursuant to any one of Clauses 10E-12E, wherein: the second coded video data does not include coded selection data; and as part of the application decoding process, one or more processors are configured to use the coded selection data to reconstruct the second picture based on the second coded video data.
[0449] Clause 14E. The device according to any one of Clauses 10E-13E, wherein one or more processors are configured to entropy encode the encoding selection data of the second picture before transmitting the encoding selection data of the second picture.
[0450] Clause 15E. The device according to any one of Clauses 10E-14E, wherein: one or more processors are further configured to: estimate a third picture of video data based on one or more of a first picture or a second picture; perform an encoding process that encodes the estimated third picture to generate third coded video data, wherein third encoding selection data indicates encoding selections for encoding the estimated third picture; send the third encoding selection data to a transmitting device; receive error correction data of the third picture from the transmitting device; apply an error correction process to generate error-corrected coded video data of the third picture based on the error correction data of the third picture and the third coded video data; and apply a decoding process that reconstructs the third picture based on the error-corrected coded video data of the third picture.
[0451] Clause 16E. The apparatus as described in Clause 15E, wherein: the error-correcting coded video data of the third picture does not include third coding selection data, and as part of the application decoding process, one or more processors are configured to use the third coding selection data to reconstruct the third picture based on the error-correcting coded video data of the third picture.
[0452] Clause 17E. An apparatus pursuant to any one of Clauses 10E-16E, wherein one or more processors are configured to: apply a channel coding process to the coding selection data of the second picture to generate error-corrected data of the coding selection data of the second picture; and transmit the error-corrected data of the coding selection data of the second picture to a transmitting device.
[0453] Clause 18E. The device according to any one of Clauses 10E-17E, wherein the device includes a communication interface configured to select data with a lower modulation order modulation coding compared to other data transmissions in the data link between the device and the transmitting device.
[0454] Clause 19E. The device according to any one of Clauses 10E-18E, wherein: one or more processors are further configured to process a second set of images to generate virtual element data, and the transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0455] Clause 20E. A method for processing video data, the method comprising: encoding a first image of the video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving encoding selection data of a second image of the video data from the receiving device, wherein: the encoding selection data of the second image indicates encoding selection for encoding an estimate of the second image, and the second image follows the first image in decoding order; encoding the second image based on the encoding selection data of the second image to generate second encoded video data; and transmitting the second encoded video data to the receiving device.
[0456] Clause 21E. The method according to Clause 20E, wherein: the encoded selection data received from the receiving device includes motion parameters of blocks of the second picture, encoding the second picture includes performing motion compensation based on the motion parameters of the blocks of the second picture to generate predicted blocks, and the second encoded video data includes encoded video data based on the predicted blocks.
[0457] Clause 22E. The method according to any one of Clauses 20E-21E, wherein: the encoded selection data received from the receiving device includes intra-prediction parameters of blocks of the second picture, encoding the second picture includes performing intra-prediction based on the intra-prediction parameters of blocks of the second picture to generate prediction blocks, and the second encoded video data includes encoded video data based on the prediction blocks.
[0458] Clause 23E. The method described in any of Clauses 20E-22E, wherein the second encoded video data does not include encoding selection data.
[0459] Clause 24E. The method according to any one of Clauses 20E-23E further includes entropy decoding of the encoded selection data of the prior second picture.
[0460] Clause 25E. The method according to any one of Clauses 20E-24E further includes: generating first error correction data based on first coded video data; and transmitting the first coded video data and the first error correction data to a receiving device.
[0461] Clause 26E. The method according to any one of Clauses 20E-25E further includes: encoding the third image based on determining that no encoding selection data of the third image has been received from the receiving device before the time limit expires, without using the encoding selection data of the third image.
[0462] Clause 27E. The method according to any one of Clauses 20E-26E further includes: receiving coding selection data for a third picture of video data, wherein the coding selection data for the third picture indicates coding selection for encoding an estimate of the third picture; encoding the third picture based on the coding selection data for the third picture to generate third coded video data; applying a channel coding process to generate error correction data for the third coded video data; and transmitting the error correction data for the third coded video data to a receiving device without transmitting at least a portion of the third coded video data.
[0463] Clause 28E. The method according to any one of Clauses 20E-27E, wherein: the device is an extended reality (XR) headset and includes a display system, and the method further includes: receiving virtual element data from a receiving device; and displaying one or more virtual elements in an XR scene on the display system based on the virtual element data.
[0464] Clause 29E. A method for processing video data, the method comprising: receiving first coded video data from a transmitting device; reconstructing a first picture of the video data based on the first coded video data; estimating a second picture of the video data based on the first picture, the second picture being a picture that appears after the first picture in decoding order; generating encoding selection data for the estimated second picture, wherein the encoding selection data indicates encoding selections for encoding the estimated second picture; transmitting the encoding selection data of the second picture to the transmitting device; receiving second coded video data from the transmitting device; and reconstructing the second picture based on the second coded video data.
[0465] Clause 30E. The method according to Clause 29E, wherein: encoding the estimated second picture includes performing motion compensation based on motion parameters of blocks in the second picture to generate predicted blocks, and encoding selection data includes motion parameters of blocks in the second picture, and the second encoded video data includes encoded video data based on predicted blocks.
[0466] Clause 31E. The method according to any one of Clauses 29E-30E, wherein: encoding the second picture includes performing intra-prediction based on intra-prediction parameters of blocks of the second picture to generate prediction blocks, encoding selection data includes intra-prediction parameters of blocks of the second picture, and the second coded video data includes coded video data based on prediction blocks.
[0467] Clause 32E. The method according to any one of Clauses 29E-31E, wherein: the second coded video data does not include coded selection data; and the application of the decoding process includes using the coded selection data to reconstruct the second picture based on the second coded video data.
[0468] Clause 33E. The method according to any one of Clauses 29E-32E, wherein entropy encoding of the encoding selection data for the second picture occurs before the encoding selection data for the second picture is transmitted.
[0469] Clause 34E. The method according to any one of Clauses 29E-33E further includes: estimating a third picture of video data based on one or more of a first picture or a second picture; performing an encoding process that encodes the estimated third picture to generate third coded video data, wherein third encoding selection data indicates encoding selections for encoding the estimated third picture; sending the third encoding selection data to a transmitting device; receiving error correction data of the third picture from the transmitting device; applying an error correction process to generate error-corrected coded video data of the third picture based on the error correction data of the third picture and the third coded video data; and applying a decoding process that reconstructs the third picture based on the error-corrected coded video data of the third picture.
[0470] Clause 35E. The method according to any one of Clauses 34E, wherein: the error-corrected coded video data of the third picture does not include third coding selection data, and the application of the decoding process includes using the third coding selection data to reconstruct the third picture based on the error-corrected coded video data of the third picture.
[0471] Clause 36E. The method according to any one of Clauses 29E-35E further includes: applying a channel coding process to the coding selection data of the second picture to generate error-corrected data of the coding selection data of the second picture; and transmitting the error-corrected data of the coding selection data of the second picture to the transmitting device.
[0472] Clause 37E. The method according to any one of Clauses 29E-36E further includes selecting data with a modulation order modulation coding lower than that of other data transmissions in the data link between the device and the transmitting device.
[0473] Clause 38E. The method according to any one of Clauses 29E-37E, wherein: it further comprises processing a second set of images to generate virtual element data, and the transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0474] Clause 39E. An apparatus comprising: means for encoding a first image of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving encoding selection data of a second image of video data from the receiving device, wherein: the encoding selection data of the second image indicates encoding selections for encoding an estimate of the second image, and the second image follows the first image in decoding order; means for encoding the second image based on the encoding selection data of the second image to generate second encoded video data; and means for transmitting the second encoded video data to the receiving device.
[0475] Clause 40E. An apparatus comprising: means for receiving first coded video data from a transmitting device; means for reconstructing a first picture of the video data based on the first coded video data; means for estimating a second picture of the video data based on the first picture, the second picture being a picture that appears after the first picture in decoding order; means for generating encoding selection data for the second picture, wherein the encoding selection data for the second picture indicates encoding selections for encoding the second picture; means for transmitting the encoding selection data for the second picture to the transmitting device; means for receiving second coded video data from the transmitting device; and means for reconstructing the second picture based on the second coded video data.
[0476] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for technical practice). Furthermore, in some examples, actions or events may be performed concurrently, such as through multithreading, interrupt handling, or multiple processors, rather than sequentially.
[0477] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium, or a communication medium, including any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium (which is non-transitory) or (2) a communication medium (such as a signal or carrier wave). A data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0478] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave may be included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient, tangible storage media. As used herein, disks and optical discs include optical discs (CDs), laser discs, optical optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically magnetically copy data, while optical discs optically copy data using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0479] Instructions can be executed by one or more processors (e.g., programmable processors), such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuitry" as used herein can refer to any of the foregoing structures or any other structures suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be fully implemented in one or more circuit or logic elements.
[0480] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be provided in combination within a codec hardware unit or by a collection of interoperable hardware units, including one or more processors as described above, combined with suitable software and / or firmware.
[0481] Various examples have been described. These examples, as well as others, are within the scope of the following claims.
Claims
1. An apparatus comprising: A memory configured to store video data; as well as One or more processors, implemented in a circuit and coupled to the memory, are configured to: The first set of images in the video data is encoded to generate first encoded video data; Send the first encoded video data to the receiving device; The receiving device receives a extraction mode indication, which indicates an extraction mode determined based on the first image set, wherein the extraction mode is a mode in which encoded video data is not transmitted. The second set of images in the video data is encoded to generate second encoded video data; The extraction mode is applied to the second encoded video data to generate extracted video data; as well as The extracted video data is sent to the receiving device.
2. The device of claim 1, wherein the one or more processors are configured to: First error correction data is generated based on the first encoded video data; Send the first error correction data to the receiving device; Generate second error correction data based on the second encoded video data; and The second error correction data is sent to the receiving device.
3. The device of claim 1, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for the complete image.
4. The device of claim 1, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for a specific region within the image.
5. The apparatus of claim 1, wherein the video data is multi-view video data, and the extraction mode indicates a mode that skips the transmission of encoded video data containing images from a particular viewpoint.
6. The device according to claim 1, wherein: The extraction mode indication is a first extraction mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, and the extracted video data is the first extracted video data. The one or more processors are further configured to: The third set of images in the video data is encoded to generate third encoded video data; Determine the second extraction mode for the second mode in which the coded video data has not been transmitted; The second extraction mode is applied to the third encoded video data to generate the second extracted video data; The second extracted video data is sent to the receiving device; as well as A second extraction mode indication is sent to the receiving device, the second extraction mode indication indicating that the second extraction mode is applied to the third encoded video data.
7. The device according to claim 1, wherein: As part of encoding the first image set, the one or more processors are configured to: Generate the first prediction data for the first image set; Residual data is generated based on the first predicted data and the first image set; A transformation is applied to the first predicted data to generate a transform block; Quantize the transform coefficients of the transform block; Entropy coding is applied to the syntax elements representing the quantization transform coefficients to generate a first entropy-coded syntax element, wherein the first coded video data includes the first entropy-coded syntax element; The one or more processors are further configured to perform analog modulation on the residual data to generate first analog modulated residual data, and The device also includes a communication interface configured to transmit the first analog modulation residual data and the first coded video data.
8. The device according to claim 1, wherein: The device is an extended reality (XR) headset and includes a display system, and The one or more processors are further configured to: Receive virtual element data from the receiving device; and The display system is configured to display one or more virtual elements in an XR scene based on the virtual element data.
9. An apparatus comprising: A memory configured to store video data; as well as One or more processors, implemented in a circuit and coupled to the memory, are configured to: Receive first encoded video data from the transmitting device; Perform a decoding process to reconstruct the first set of images based on the first encoded video data; The extraction mode is determined based on the first image set, wherein the extraction mode indicates a mode in which encoded video data is not transmitted; Send a descent mode indication, which is determined by the indication, to the sending device; The device receives extracted video data, wherein the extracted video data includes second encoded video data to which the extraction mode has been applied, wherein the second encoded video data is generated based on a second set of images of the video data; as well as The decoding process is performed to reconstruct the second set of images based on the second encoded video data.
10. The device according to claim 9, The one or more processors are further configured to: Receive first error correction data from the transmitting device; The error correction process is applied to modify the first coded video data based on the first error correction data to generate the first error-corrected coded video data; The one or more processors are configured to perform the decoding process to reconstruct the first set of images based on the first error-corrected coded video data; The one or more processors are further configured to: Receive second error correction data from the transmitting device; The error correction process is applied to generate second error-corrected coded video data based on the second coded video data and the second error-corrected data; as well as The one or more processors are configured to perform the decoding process to reconstruct the second set of images based on the second error-corrected coded video data.
11. The apparatus of claim 9, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for the complete picture.
12. The device of claim 9, wherein, as part of determining the extraction mode, the one or more processors are configured to: The extraction mode is applied to the first encoded video data to generate extracted encoded video data; The error correction process is applied to modify the extracted coded video data based on the first error correction data to generate experimental error-corrected video data; The decoding process is applied to reconstruct the first set of images based on the experimental error-correcting video data; Whether the extraction mode meets the standard is determined by comparing the first set of images reconstructed based on the experimental error-correcting video data and the first set of images reconstructed based on the first video data.
13. The apparatus of claim 9, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for a specific region within the image.
14. The apparatus of claim 9, wherein the video data is multi-view video data, and the extraction mode indicates a mode that skips the transmission of encoded video data containing images from a particular viewpoint.
15. The device according to claim 9, wherein: The extraction mode indication is a first extraction mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, and the extracted video data is the first extracted video data. The one or more processors are further configured to: Receive a second extraction mode indication indicating that the second mode of encoded video data has not been transmitted; Receive second extracted video data from the transmitting device, wherein the second extracted video data includes third encoded video data to which the second extraction mode has been applied, wherein the third encoded video data is generated based on a third set of images of the video data; as well as The decoding process is applied to reconstruct the third set of images based on the third encoded video data.
16. The device according to claim 9, wherein: The device also includes a communication interface configured to receive analog modulated residual data. The second encoded video data includes entropy coding syntax elements representing quantization transform coefficients; As part of applying the decoding process to reconstruct the second image set, the one or more processors are configured to: Entropy decoding is applied to the syntax elements to obtain the quantization transform coefficients; as well as The quantization transform coefficients are dequantized to generate dequantized transform coefficients; The inverse transform coefficients are applied to generate prediction data; The analog modulation residual data is demodulated to obtain residual data; The second image set is reconstructed based on the predicted data and the residual data.
17. The device according to claim 9, wherein: The one or more processors are further configured to process the second image set to generate virtual element data, and The transmitting device is an extended reality (XR) headset, which is configured to display one or more virtual elements in an XR scene based on the virtual element data.
18. A method comprising: Encode the first set of images in the video data to generate the first encoded video data; Send the first encoded video data to the receiving device; The receiving device receives a extraction mode indication, which indicates an extraction mode determined based on the first image set, wherein the extraction mode is a mode in which encoded video data is not transmitted. The second set of images in the video data is encoded to generate second encoded video data; The extraction mode is applied to the second encoded video data to generate extracted video data; as well as The extracted video data is sent to the receiving device.
19. The method of claim 18, further comprising: First error correction data is generated based on the first encoded video data; Send the first error correction data to the receiving device; Generate second error correction data based on the second encoded video data; as well as The second error correction data is sent to the receiving device.
20. The method of claim 18, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for the complete picture.
21. The method of claim 18, wherein the video data is multi-view video data, and the extraction mode indicates a mode that skips the transmission of encoded video data containing images from a particular viewpoint.
22. The method of claim 18, wherein: The extraction mode indication is a first extraction mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, and the extracted video data is the first extracted video data. The method further includes: The third set of images in the video data is encoded to generate third encoded video data; Determine the second extraction mode for the second mode in which the coded video data has not been transmitted; The second extraction mode is applied to the third encoded video data to generate the second extracted video data; Send the second extracted video data to the receiving device; and A second extraction mode indication is sent to the receiving device, the second extraction mode indication indicating that the second extraction mode is applied to the third encoded video data.
23. The method of claim 18, wherein: Encoding the first image set includes: Generate the first prediction data for the first image set; Residual data is generated based on the first predicted data and the first image set; A transformation is applied to the first predicted data to generate a transform block; Quantize the transform coefficients of the transform block; Entropy coding is applied to the syntax elements representing the quantization transform coefficients to generate a first entropy-coded syntax element, wherein the first coded video data includes the first entropy-coded syntax element; The method further includes: The residual data is subjected to analog modulation to generate first analog modulated residual data, and Send the first analog modulation residual data and the first coded video data.
24. A method comprising: Receive first encoded video data from the transmitting device; The application decoding process reconstructs the first image set based on the first encoded video data; The extraction mode is determined based on the first image set, wherein the extraction mode indicates a mode in which encoded video data is not transmitted; Send a descent mode indication, which is determined by the indication, to the sending device; The device receives extracted video data, wherein the extracted video data includes second encoded video data to which the extraction mode has been applied, wherein the second encoded video data is generated based on a second set of images of the video data; as well as The decoding process is performed to reconstruct the second set of images based on the second encoded video data.
25. The method according to claim 24, The method further includes: Receive first error correction data from the transmitting device; The error correction process is applied to modify the first coded video data based on the first error correction data to generate the first error-corrected coded video data; The process of performing the decoding to reconstruct the first image set includes performing the decoding process to reconstruct the first image set based on the first error-corrected coded video data; The method further includes: Receive second error correction data from the transmitting device; The error correction process is applied to generate second error-corrected coded video data based on the second coded video data and the second error-corrected data; and The process of performing the decoding to reconstruct the second image set includes performing the decoding process to reconstruct the second image set based on the second error-corrected coded video data.
26. The method of claim 24, wherein the extraction mode indicates a mode that skips the transmission of encoded video data for the complete picture.
27. The method of claim 24, wherein determining the extraction mode comprises: The extraction mode is applied to the first encoded video data to generate extracted encoded video data; The error correction process is applied to modify the extracted coded video data based on the first error correction data to generate experimental error-corrected video data; The decoding process is applied to reconstruct the first set of images based on the experimental error-correcting video data; as well as Whether the extraction mode meets the standard is determined by comparing the first set of images reconstructed based on the experimental error-correcting video data and the first set of images reconstructed based on the first video data.
28. The method of claim 24, wherein the video data is multi-view video data, and the extraction mode indicates a mode that skips the transmission of encoded video data containing images from a particular viewpoint.
29. The method according to claim 24, wherein: The extraction mode indication is a first extraction mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, and the extracted video data is the first extracted video data. The method further includes: Receive a second extraction mode indication indicating that the second mode of encoded video data has not been transmitted; Receive second extracted video data from the transmitting device, wherein the second extracted video data includes third encoded video data to which the second extraction mode has been applied, wherein the third encoded video data is generated based on a third set of images of the video data; and The decoding process is applied to reconstruct the third set of images based on the third encoded video data.
30. The method of claim 24, wherein: The method further includes receiving analog-modulated residual data. The second encoded video data includes entropy coding syntax elements representing quantization transform coefficients; Reconstructing the second image set using the decoding process described above includes: Entropy decoding is applied to the syntax elements to obtain the quantization transform coefficients; The quantization transform coefficients are dequantized to generate dequantized transform coefficients; The inverse transform coefficients are applied to generate prediction data; The analog modulation residual data is demodulated to obtain residual data; The second image set is reconstructed based on the predicted data and the residual data.
Citation Information
Patent Citations
Analog modulated video transmission with variable symbol rate
US11431962B2
Interlaced coefficients in hybrid digital-analog modulation for transmission of video data
US11457224B2
Hybrid digital-analog modulation for transmission of video data
US11553184B2