Receiver-side prediction of coding selection data for video coding
By using distributed video decoding technology, the video encoding process is partially offloaded on the receiving device, solving the problems of high resource consumption and high latency in XR devices, and achieving efficient video data transmission and low-power video reconstruction.
Patent Information
- Application Number
- CN202480026345.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2024-04-08
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies suffer from high resource consumption, high latency, and high power consumption when processing high-quality video data, especially in extended reality (XR) devices. In particular, when transmitting unencoded video data in wireless communication systems, the bandwidth-limited and complex video encoding process increases the burden on the device.
By employing distributed video decoding (DVC) technology, the video encoding process is partially offloaded to the receiving device. The transmitting device generates limited video encoded data and sends error correction data. The receiving device estimates video data based on the error correction data and previously reconstructed images, and uses more sophisticated decoding tools for reconstruction and error correction, thereby reducing the amount of data sent and optimizing resource utilization.
It effectively reduces the resource consumption and power consumption of the transmitting device, reduces data transmission time, improves decoding efficiency, and maintains low-latency video data transmission quality.
Smart Images

Figure CN121002852A_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application Serial No. 18 / 306,131, filed April 24, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to video encoding and decoding. Background Technology
[0003] The popularity of virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies is growing rapidly and is expected to be widely used in applications beyond gaming, such as healthcare, education, social services, and retail. VR, AR, and MR can be collectively referred to as extended reality (XR). Due to this increasing popularity, there is a growing demand for XR devices with high 3D graphics quality, higher video resolution, and low latency response, such as XR goggles. Summary of the Invention
[0004] This disclosure describes techniques for processing video data in a transmitting and receiving device. The transmitting device may be an XR device or other type of device. The receiving device may be a user equipment (UE) device, such as a smartphone or tablet. The transmitting device may perform a limited video coding process on the video data to generate coded video data. The transmitting device may apply channel coding to the coded video data to generate error correction data. The transmitting device may transmit at least some of the error correction data and the coded video data to the receiving device. The receiving device may estimate the video data based on one or more previously reconstructed frames. The receiving device may then encode the estimated video data. The receiving device may use one or more decoding tools to encode the estimated video data that the transmitting device did not use when performing the limited video coding process on the video data. The receiving device may use the error correction data and the predicted video data to regenerate portions of the coded video data that the transmitting device did not transmit. This process avoids the need to transmit portions of the coded video data.
[0005] In one example, this disclosure describes a method for decoding video data, comprising: obtaining error correction data from a transmitting device at a receiving device, wherein the error correction data provides error correction information and is generated based on coded video data of one or more blocks of a frame of the video data; generating prediction data of the frame at the receiving device using one or more decoding tools not used to generate the coded video data of the one or more blocks, wherein the prediction data of the frame includes predictions of blocks of the frame based at least in part on one or more previously reconstructed frames of the video data; generating coded video data at the receiving device based on the prediction data of the frame; generating error-corrected coded video data at the receiving device using the error correction data to perform an error correction operation on the coded video data; and performing a reconstruction operation at the receiving device to reconstruct blocks of the frame based on the error-corrected coded video data, wherein the reconstruction operation is controlled by the values of one or more parameters.
[0006] In another example, this disclosure describes a method for encoding video data, comprising: obtaining video data from a video source at a transmitting device; generating encoded video data of a first frame and encoded video data of a second frame of the video data based on a parameter set at the transmitting device; performing channel coding on the encoded video data of the first frame and the encoded video data of the second frame at the transmitting device to generate error-corrected data of the first frame and error-corrected data of the second frame; and transmitting the encoded video data of the first frame, the error-corrected data of the first frame, and the error-corrected data of the second frame at the transmitting device.
[0007] In another example, this disclosure describes a method for encoding video data, comprising: obtaining video data from a video source at a transmitting device; generating transform blocks based on the video data at the transmitting device; determining which transform blocks in the transform blocks are anchor transform blocks at the transmitting device; calculating a correlation matrix of the transform block set at the transmitting device; generating a bit-reduced non-anchor transform matrix at the transmitting device; and transmitting the anchor transform blocks, non-anchor transform blocks, and correlation matrix to a receiving device at the transmitting device.
[0008] In another example, this disclosure describes an apparatus comprising: a memory configured to store video data; a communication interface; and one or more processes implemented in a circuit and coupled to the memory, the one or more processors being configured to perform the method according to any one of claims 1-22.
[0009] In another example, this disclosure describes an apparatus for processing video data, comprising: a memory configured to store video data; and a communication interface configured to obtain error correction data from a transmitting device, wherein the error correction data provides error correction information about frames of the video data; one or more processes implemented in circuitry and coupled to the memory, the one or more processors being configured to: generate prediction data for frames, wherein the prediction data for frames includes predictions of blocks of frames based at least in part on one or more previously reconstructed frames of the video data; generate coded video data based on the prediction data for frames, wherein the coded video data includes transform blocks, wherein the transform blocks include transform coefficients; scale the bits of the transform coefficients of the transform blocks based on reliability values of bit positions; use the error correction data to generate error-corrected coded video data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstruct a frame based on the error-corrected coded video data.
[0010] In another example, this disclosure describes an apparatus for processing video data, comprising: a memory configured to store video data; and one or more processes implemented in circuitry and coupled to the memory, the one or more processors being configured to: acquire the video data; acquire predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated frame generated by a receiving device; adjust one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; perform a video coding process to generate coded video data based on one or more frames of the acquired video data, wherein the video coding process is controlled by the video coding parameters; perform a channel coding process on the coded video data to generate channel-coded data, wherein the channel coding process is controlled by the channel coding parameters; and a communication interface configured to transmit the channel-coded data to the receiving device.
[0011] In another example, this disclosure describes a method for processing video data, comprising: obtaining error correction data from a transmitting device at a receiving device, wherein the error correction data provides error correction information about a frame of video data; generating prediction data of the frame at the receiving device, wherein the prediction data of the frame includes predictions of blocks of the frame based at least in part on one or more previously reconstructed frames of the video data; generating coded video data at the receiving device based on the prediction data of the frame, wherein the coded video data includes transform blocks, wherein the transform blocks include transform coefficients; scaling bits of the transform coefficients of the transform blocks at the receiving device based on reliability values of bit positions; generating error-corrected coded video data at the receiving device using the error correction data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstructing a frame at the receiving device based on the error-corrected coded video data.
[0012] In another example, this disclosure describes a method for processing video data, comprising: acquiring video data; acquiring predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated frame generated by a receiving device; adjusting one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process to generate coded video data based on one or more frames of the acquired video data, wherein the video coding process is controlled by the video coding parameters; performing a channel coding process on the coded video data to generate channel-coded data, wherein the channel coding process is controlled by the channel coding parameters; and transmitting the channel-coded data to a receiving device.
[0013] In another example, this disclosure describes an apparatus including: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: acquire a first set of multi-view frames of the video data, wherein the first set of multi-view frames includes a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; transmit first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multi-view frames; receive a multi-view encoding prompt from the receiving device; acquire a second set of multi-view frames of the video data, wherein the second set of multi-view frames includes a third frame and a fourth frame, the third frame being from the first viewpoint and the fourth frame being from the second viewpoint; perform a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompt received from the receiving device to generate second encoded video data, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame; and transmit the second encoded video data to the receiving device.
[0014] In another example, this disclosure describes an apparatus including: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: obtain first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multi-view frames of the video data, the first set of multi-view frames including a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; determine a multi-view encoding cue based on the first encoded video data; send the multi-view encoding cue to the transmitting device; and obtain second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multi-view frames including a third frame and a fourth frame, the second encoded video data being encoded using a multi-view encoding process, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame based on the multi-view encoding cue.
[0015] In another example, this disclosure describes a method for processing video data, comprising: obtaining a first set of multi-view frames of video data, wherein the first set of multi-view frames includes a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multi-view frames; receiving a multi-view encoding prompt from the receiving device; obtaining a second set of multi-view frames of video data, wherein the second set of multi-view frames includes a third frame and a fourth frame, the third frame being from the first viewpoint and the fourth frame being from the second viewpoint; performing a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompt received from the receiving device to generate second encoded video data, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame; and transmitting the second encoded video data to the receiving device.
[0016] In another example, this disclosure describes a method for processing video data, comprising: obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multi-view frames of the video data, the first set of multi-view frames including a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; determining a multi-view encoding cue based on the first encoded video data; sending the multi-view encoding cue to the transmitting device; and obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multi-view frames including a third frame and a fourth frame, the second encoded video data being encoded using a multi-view encoding process, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame based on the multi-view encoding cue.
[0017] In another example, this disclosure describes an apparatus comprising: components for acquiring a first set of multi-view frames of video data, wherein the first set of multi-view frames includes a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; components for transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multi-view frames; components for receiving a multi-view encoding prompt from the receiving device; components for acquiring a second set of multi-view frames of video data, wherein the second set of multi-view frames includes a third frame and a fourth frame, the third frame being from the first viewpoint and the fourth frame being from the second viewpoint; components for performing a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompt received from the receiving device to generate second encoded video data, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame; and components for transmitting the second encoded video data to the receiving device.
[0018] In another example, this disclosure describes an apparatus comprising: components for obtaining first encoded video data from a transmitting device, wherein the first encoded video data is a first set of multi-view frames based on video data, the first set of multi-view frames including a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; components for determining multi-view encoding prompts based on the first encoded video data; components for sending multi-view encoding prompts to the transmitting device; and components for obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multi-view frames including a third frame and a fourth frame, the second encoded video data being encoded using a multi-view encoding process, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame based on the multi-view encoding prompts.
[0019] In another example, this disclosure describes an apparatus including: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: encode a first set of frames of the video data to generate first encoded video data; transmit the first encoded video data to a receiving device; receive from the receiving device a decimation mode indication indicating a decimation mode determined based on the first set of frames, wherein the decimation mode is a mode in which the encoded video data is not transmitted; encode a second set of frames of the video data to generate second encoded video data; apply the decimation mode to the second encoded video data to generate decimated video data; and transmit the decimated video data to the receiving device.
[0020] In another example, this disclosure describes an apparatus including: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: receive first encoded video data from a transmitting device; perform a decoding process to reconstruct a first set of frames based on the first encoded video data; determine an extraction mode based on the first set of frames that indicates a mode in which the encoded video data has not been transmitted; send an extraction mode indication to the transmitting device that indicates the determined extraction mode; receive extracted video data from the transmitting device, wherein the extracted video data includes second encoded video data to which the extraction mode has been applied, wherein the second encoded video data is generated based on a second set of frames of the video data; and perform a decoding process to reconstruct a second set of frames based on the second encoded video data.
[0021] In another example, this disclosure describes a method comprising: encoding a first set of frames of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving from the receiving device a decimation mode indication indicating a decimation mode determined based on the first set of frames, wherein the decimation mode is a mode in which the encoded video data is not transmitted; encoding a second set of frames of video data to generate second encoded video data; applying the decimation mode to the second encoded video data to generate decimated video data; and transmitting the decimated video data to the receiving device.
[0022] In another example, this disclosure describes a method comprising: receiving first encoded video data from a transmitting device; applying a decoding process to reconstruct a first set of frames based on the first encoded video data; determining an extraction mode based on the first set of frames that indicates a mode in which the encoded video data has not been transmitted; sending an extraction mode indication to the transmitting device that indicates the determined extraction mode; receiving extracted video data from the transmitting device, wherein the extracted video data includes second encoded video data to which the extraction mode has been applied, wherein the second encoded video data is generated based on a second set of frames of the video data; and performing a decoding process to reconstruct a second set of frames based on the second encoded video data.
[0023] In another example, this disclosure describes an apparatus comprising: means for encoding a first set of frames of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device an indication of a decimation mode determined based on the first set of frames, wherein the decimation mode is a mode in which the encoded video data is not transmitted; means for encoding a second set of frames of video data to generate second encoded video data; means for applying the decimation mode to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to the receiving device.
[0024] In another example, this disclosure describes an apparatus comprising: means for receiving first encoded video data from a transmitting device; means for performing a decoding process to reconstruct a first set of frames based on the first encoded video data; means for determining an extraction mode based on the first set of frames that indicates a mode in which the encoded video data has not been transmitted; means for sending an extraction mode indication to the transmitting device that indicates the determined extraction mode; means for receiving extracted video data from the transmitting device, wherein the extracted video data includes second encoded video data to which the extraction mode has been applied, wherein the second encoded video data is generated based on a second set of frames of the video data; and means for performing a decoding process to reconstruct a second set of frames based on the second encoded video data.
[0025] In another example, this disclosure describes an apparatus including: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: encode a first frame of the video data to generate first encoded video data; transmit the first encoded video data to a receiving device; receive encoding selection data for a second frame of the video data from the receiving device, wherein: the encoding selection data for the second frame indicates encoding selections for encoding an estimate of the second frame, and the second frame follows the first frame in decoding order; encode the second frame based on the encoding selection data for the second frame to generate second encoded video data; and transmit the second encoded video data to the receiving device.
[0026] In another example, this disclosure describes an apparatus including: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: receive first coded video data from a transmitting device; reconstruct a first frame of the video data based on the first coded video data; estimate a second frame of the video data based on the first frame, wherein the second frame is a frame that appears after the first frame in decoding order; generate encoding selection data for the second frame, wherein the encoding selection data for the second frame indicates encoding selections for encoding the second frame; transmit the encoding selection data for the second frame to the transmitting device; receive second coded video data from the transmitting device; and reconstruct the second frame based on the second coded video data.
[0027] In another example, this disclosure describes a method for processing video data, comprising: encoding a first frame of the video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving encoding selection data for a second frame of the video data from the receiving device, wherein: the encoding selection data for the second frame indicates encoding selections for encoding an estimate of the second frame, and the second frame follows the first frame in a decoding order; encoding the second frame based on the encoding selection data for the second frame to generate second encoded video data; and transmitting the second encoded video data to the receiving device.
[0028] In another example, this disclosure describes a method for processing video data, comprising: receiving first encoded video data from a transmitting device; reconstructing a first frame of the video data based on the first encoded video data; estimating a second frame of the video data based on the first frame, wherein the second frame is a frame that appears after the first frame in decoding order; generating encoding selection data for the second frame, wherein the encoding selection data for the second frame indicates encoding selections for encoding the second frame; transmitting the encoding selection data for the second frame to the transmitting device; receiving second encoded video data from the transmitting device; and reconstructing the second frame based on the second encoded video data.
[0029] In another example, this disclosure describes an apparatus comprising: means for encoding a first frame of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving encoding selection data for a second frame of the video data from the receiving device, wherein: the encoding selection data for the second frame indicates encoding selections for encoding an estimate of the second frame, and the second frame follows the first frame in decoding order; means for encoding the second frame based on the encoding selection data for the second frame to generate second encoded video data; and means for transmitting the second encoded video data to the receiving device.
[0030] In another example, this disclosure describes an apparatus comprising: means for receiving first encoded video data from a transmitting device; means for reconstructing a first frame of the video data based on the first encoded video data; means for estimating a second frame of the video data based on the first frame, wherein the second frame is a frame that appears after the first frame in decoding order; means for generating encoding selection data for the second frame, wherein the encoding selection data for the second frame indicates encoding selections for encoding the second frame; means for transmitting the encoding selection data for the second frame to the transmitting device; means for receiving second encoded video data from the transmitting device; and means for reconstructing the second frame based on the second encoded video data.
[0031] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Attached Figure Description
[0032] Figure 1 This is a block diagram illustrating an example system according to the technology of this disclosure.
[0033] Figure 2 This is a block diagram illustrating example components of a transmitting device and a receiving device according to the technology of this disclosure.
[0034] Figure 3A This is a conceptual diagram illustrating an example channel coding process according to the technology disclosed herein.
[0035] Figure 3B This is a block diagram illustrating an example channel decoding process according to the technology of this disclosure.
[0036] Figure 4 This is a flowchart illustrating an example operation of a transmitting device according to the technology of this disclosure.
[0037] Figure 5 This is a flowchart illustrating an example operation of a receiving device according to the technology disclosed herein.
[0038] Figure 6 This is a conceptual diagram illustrating an example decimation pattern according to the technology disclosed herein.
[0039] Figure 7 This is a flowchart illustrating an example operation of a transmitting device for hybrid extraction of transform blocks according to the technology of this disclosure.
[0040] Figure 8 This is a flowchart illustrating an example operation of a receiving device for hybrid decimation of a transform block according to the technology of this disclosure.
[0041] Figure 9 This is a conceptual diagram illustrating an example extraction mode adaptively selected by a receiving device according to one or more techniques of this disclosure.
[0042] Figure 10 This is a block diagram illustrating example components of a transmitting device and a receiving device according to the technology of this disclosure.
[0043] Figure 11 A graph showing example error probabilities and corresponding absolute values of the covariance likelihood ratio (LLR) according to one or more techniques of this disclosure is provided.
[0044] Figure 12 This is a flowchart illustrating an example operation of a transmitting device using scaled bits according to the technology of this disclosure.
[0045] Figure 13 This is a flowchart illustrating an example operation of a receiving device using scaled bits according to the technology of this disclosure.
[0046] Figure 14 This is a flowchart illustrating an example data exchange between a transmitting device and a receiving device in relation to multi-view processing according to one or more techniques of this disclosure.
[0047] Figure 15 This is a flowchart illustrating an example operation of a transmitting device for multi-view processing according to the technology of this disclosure.
[0048] Figure 16 This is a flowchart illustrating an example operation of a receiving device for multi-view processing according to the technology of this disclosure.
[0049] Figure 17 This is a block diagram illustrating example components of a transmitting device and a receiving device that perform extraction of coded video data according to the technology of this disclosure.
[0050] Figure 18 This is a conceptual diagram illustrating an example exchange of information including an extraction pattern indication according to the technology of this disclosure.
[0051] Figure 19 This is a flowchart illustrating an example operation of a transmitting device according to the technology of this disclosure, wherein the transmitting device receives a decimation mode indication.
[0052] Figure 20 This is a flowchart illustrating an example operation of a receiving device according to the technology of this disclosure, wherein the receiving device sends a decimation mode indication.
[0053] Figure 21 This is a block diagram illustrating example components of a transmitting device according to the technology of this disclosure and a receiving device that transmits encoded selection data to the transmitting device.
[0054] Figure 22 This is a communication diagram illustrating an example exchange of data, including the transmission and reception of encoded selection data, between a transmitting device and a receiving device according to the technology of this disclosure.
[0055] Figure 23 This is a flowchart illustrating an example operation of a transmitting device according to the technology of this disclosure, wherein the transmitting device receives encoding selection data.
[0056] Figure 24 This is a flowchart illustrating an example operation of a receiving device according to the technology of this disclosure, wherein the receiving device transmits encoding selection data.
[0057] Figure 25 This is a conceptual diagram illustrating an example hierarchy of encoded video data according to the technology of this disclosure.
[0058] Figure 26 This is a block diagram illustrating alternative example components of a transmitting device according to one or more technologies of this disclosure.
[0059] Figure 27 This is a block diagram illustrating example alternative components of a receiving device according to one or more technologies of this disclosure. Detailed Implementation
[0060] While modern video coding processes can significantly reduce the amount of data required to represent video data, these processes are typically resource-intensive and can involve numerous memory operations. Therefore, modern video coding processes require sophisticated processors, fast memory, and consume considerable power. However, for some contemporary and future planned wireless communication systems, such as 5G and 6G wireless communication systems, wireless transmission bandwidth may be less constrained, especially when communicating over short distances (such as the distance between devices on a person).
[0061] This disclosure describes techniques for reducing the complexity of video coding at a transmitting device by using error correction performed as part of channel decoding that utilizes error correction data. The transmitting device can perform a finite video coding process that generates coded video data. Finite video coding processes typically use relatively less resource-intensive coding tools, such as intra-frame prediction. Because the video coding process uses less complex coding tools, the resulting coded video data can be larger than the video data encoded using more complex and resource-intensive coding tools. Error correction data is based on the coded video data. The transmitting device can send the error correction data to a receiving device. The transmitting device may not necessarily send all the coded video data for one or more frames to the receiving device.
[0062] The receiving device can estimate frames of video data based on one or more previously reconstructed frames. In some examples, to estimate frames, the receiving device can extrapolate the content of blocks from previously reconstructed frames. The receiving device can then perform a full video coding process on the estimated frames to generate estimated coded video data for the frames. When performing a full video coding process, the receiving device can use more complex decoding tools, such as inter-frame prediction, than the limited video coding process performed by the transmitting device. The receiving device can perform a channel decoding process, which generates error-corrected coded video data based on the estimated coded video data of the frames and the error-correcting data of the frames. In some cases, the channel decoding process can generate error-corrected coded video data based on the error-correcting data of the frames and a combination of the estimated coded video data of the frames and the coded video data of the frames transmitted by the transmitting device. The receiving device can reconstruct the frames based on the error-corrected coded video data. In this way, the receiving device can reconstruct each frame of the video data even if the transmitting device has not transmitted all the coded video data of the frames.
[0063] As further described in this disclosure, various techniques can be applied to specify which transform blocks of lightly coded video data are not signaled or to reduce bit depth, such as the application of decimation modes. Furthermore, in some examples of this disclosure, reliability values can be determined for bit positions, and these reliability values can be used to scale the bits of the transform coefficients of the transform blocks, and the scaling values can be used for channel coding and channel decoding.
[0064] As further described in this disclosure, the receiving device can determine a decimation mode based on a first set of frames. The decimation mode is a mode in which encoded video data is not transmitted. The receiving device can send a decimation mode indication, indicating the determined decimation mode, to the transmitting device. The transmitting device can receive the decimation mode indication from the receiving device and apply the indicated decimation mode to the encoded video data to generate decimated video data. The transmitting device can then send the decimated video data to the receiving device. In this way, the technology of this disclosure can further reduce resource consumption at the transmitting device while still avoiding the transmission of excessive data. This can further improve decoding efficiency.
[0065] Figure 1 This is a block diagram illustrating an example system 100 according to the technology of this disclosure. Figure 1 In this example, system 100 includes a transmitting device 102, a receiving device 104, and a base station 106. The transmitting device 102 can be a device configured to perform actions including extended reality (XR) devices (e.g., XR headsets), mobile devices, wearable devices, sensor devices, Internet of Things (IoT) devices, intermediate networking devices, or other types of devices. In some examples, the transmitting device 102 may be included in a robot or vehicle. The receiving device 104 can be a computing device, such as a mobile device (e.g., a mobile phone or tablet computer), a personal computer, a vehicle-based computing device, a wireless base station, a wearable computing device, an intermediate networking device, a dedicated device, an Internet of Things (IoT) device, or other types of devices. In some examples, the receiving device 104 is a device that a user of the transmitting device 102 may have in addition to the transmitting device 102.
[0066] Transmitting device 102 and receiving device 104 can communicate with base station 106. In some examples, transmitting device 102 and receiving device 104 can use fifth-generation (5G) wireless communication protocols, sixth-generation (6G) wireless communication protocols, WiFi protocols, Bluetooth protocols, or another type of wireless communication protocol to communicate with base station 106. Base station 106 can transmit data from network 115 to transmitting device 102 and receiving device 104 via wireless downlink channels 108A and 108B (collectively referred to as "wireless downlink channel 108"). Base station 106 can receive data from transmitting device 102 and receiving device 104 for transmission to other devices connected to network 115 via wireless uplink channels 110A and 110B (collectively referred to as "wireless uplink channel 110"). Transmitting device 102 and receiving device 104 can communicate directly with each other via wireless sidelink channel 112. In other examples, transmitting device 102 and receiving device 104 can communicate via other types of channels.
[0067] exist Figure 1 In the example, transmitting device 102 includes one or more processors 114, memory 116, communication interface 118, video source 120, and display system 122. Receiving device 104 includes one or more processors 130, memory 132, and communication interface 134. Processors 114 and 130 may include circuitry configured to perform various information processing tasks, including executing computer-readable instructions. Processors 114 and 130 may include microprocessors, digital signal processors, and other types of circuitry. Memory 116 and 132 may be configured to store data, such as computer-readable instructions, video data, and other types of data. Communication interfaces 118 and 134 may be configured to transmit and receive data, for example, via wireless downlink channel 108, wireless uplink channel 110, and wireless sidelink channel 112.
[0068] Generally, video source 120 refers to a source of video data (e.g., raw, unencoded video data). Video source 120 may include one or more video capture devices, such as cameras, video archives containing previously captured raw video, and / or video feed interfaces for receiving video from video content providers. Alternatively, video source 120 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video.
[0069] In the example where the transmitting device 102 is an XR device that presents MR and AR visuals to a user, it may be necessary to analyze video data from video source 120 so that the display system 122 of the transmitting device 102 can display virtual elements in the correct positions. Processing video data in this way may require significant computing resources. In other words, powerful processors and a large amount of energy can be used when processing video data. Because the transmitting device 102 may be designed to be worn on a user's head, minimizing the weight and power consumption of the transmitting device 102 while supporting high-quality, low-latency video is important.
[0070] Furthermore, in some examples, the transmitting device 102 is an XR headset, and the transmitting device 102 can be configured to process video data to generate virtual element data. The receiving device 104 can be configured to transmit (and the transmitting device 102 is configured to receive) the virtual element data. The transmitting device 102 may include a display system 122 configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0071] Therefore, it may be desirable to offload video data processing to a device other than transmitting device 102 (such as receiving device 104). Receiving device 104 may have more resources than transmitting device 102, either permanently or temporarily. For example, receiving device 104 may be equipped with a larger battery and a relatively powerful processor. However, in order for receiving device 104 to process video data, transmitting device 102 may need to transmit video data to receiving device 104 via wireless sidelink channel 112. Because a very large number of bits may be required to represent unencoded high-quality video data, transmitting unencoded high-quality video data to receiving device 104 will consume a significant amount of time and energy. The time required for transmission may undermine the goal of providing low-latency video to the user. The energy required for transmission may undermine the goal of minimizing power consumption. Encoding video data using a video decoding specification (such as H.264 / Advanced Video Decoding (AVC), H.265 / High-Efficiency Video Decoding (HEVC), or H.266 / Various Video Decoding (VVC)) can significantly reduce the amount of data required to represent video data. However, the encoding process itself may introduce its own latency and power consumption requirements.
[0072] This disclosure describes techniques that can solve these problems. According to the techniques of this disclosure, transmitting device 102 and receiving device 104 can use a distributed video decoding (DVC) process. The DVC process reduces the amount of encoding work performed by transmitting device 102 and offloads some of the encoding work to receiving device 104. Receiving device 104 may have more resources (e.g., computing power, power access, etc.) than transmitting device 102, and therefore can be better equipped to perform encoding work. In some examples, the DVC process can be used to load balance computational tasks between devices. For example, the system may determine that receiving device 104 is generally more efficient than transmitting device 102 in performing a particular video-related computational task.
[0073] In addition to the video encoding process, transmitting device 102 may also perform a channel coding process to prepare encoded video data for transmission to receiving device 104. The channel coding process can generate error correction data for the data sequence within the encoded video data. Typically, receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, according to the technology of this disclosure, transmitting device 102 may transmit error correction data for some encoded video data, but not the encoded video data corresponding to the error correction data. Receiving device 104 may estimate one or more subsequent frames. Receiving device 104 may perform a video encoding process on the subsequent frames to generate estimated encoded video data. The receiving device can use the estimated encoded video data and the received error correction data to generate error-corrected encoded video data. Receiving device 104 can then decode the error-corrected encoded video data to reconstruct the video data not transmitted by transmitting device 102.
[0074] Therefore, in some examples, receiving device 104 can obtain first coded video data and first error correction data from transmitting device 102. The first coded video data may represent one or more blocks of a first frame of video data. The first error correction data can provide error correction information about the blocks of the first frame. Receiving device 104 can use the first error correction data to generate first error-corrected coded video data to perform error correction operations on the first coded video data. Additionally, receiving device 104 can perform a first reconstruction operation to reconstruct blocks of the first frame based on the first coded video data. The first reconstruction operation can be controlled by the values of one or more parameters.
[0075] Furthermore, receiving device 104 can obtain second error correction data from transmitting device 102. The second error correction data can provide error correction information about one or more blocks of a second frame of video data. Receiving device 104 can generate prediction data for the second frame. The prediction data for the second frame can include predictions of blocks of the second frame of video data based at least in part on blocks of one or more previously reconstructed frames (such as the first frame). Receiving device 104 can use one or more decoding tools to generate prediction data for encoded video data not used to generate the second frame. Receiving device 104 can generate second encoded video data based on predictions of blocks of the second frame. Receiving device 104 can use the second error correction data to generate second error-corrected encoded video data to perform error correction operations on the second encoded video data. Receiving device 104 can perform a second reconstruction operation based on the second error-corrected encoded video data to reconstruct blocks of the second frame. The second reconstruction operation is controlled by the values of parameters.
[0076] Furthermore, according to one or more techniques of this disclosure, receiving device 104 can receive a decimation mode indication. Receiving device 104 can determine the decimation mode indication based on previously reconstructed frames. The decimation mode indication can indicate a mode in which encoded video data is not transmitted. For example, the decimation mode can indicate a mode in which the transmission of encoded video data for the entire frame is skipped. In some examples, the decimation mode indicates a mode in which the transmission of encoded video data for a specified region within a frame is skipped. In some examples where the video data is multi-view video data, the decimation mode can indicate a mode in which the transmission of encoded video data for frames from a specified view is skipped.
[0077] Transmitting device 102 can perform a video encoding process on the frame of video data. The video encoding process can compress the frame less than "heavy" or more complex compression operations (such as those described in the H.264, H.265, and H.266 video decoding standards). In addition to the video encoding process, transmitting device 102 can also perform a channel coding process to prepare encoded video data for transmission to receiving device 104. The channel coding process can generate error correction data for the data sequence within the encoded video data. Typically, receiving device 104 uses the error correction data to correct errors introduced into the encoded video data during transmission. However, receiving device 104 can also use the error correction data to recover information intentionally not sent to the receiving device. Therefore, transmitting device 102 can apply a decimation mode to the second encoded video data to generate decimated video data. Transmitting device 102 can then transmit the error correction data (which is generated based on the undecimated encoded video data) and the decimated video data to receiving device 104.
[0078] Receiving device 104 can obtain first coded video data and first error correction data from transmitting device 102. The first coded data can represent one or more blocks of a first frame of video data. The first error correction data can provide error correction information about the blocks of the first frame. Receiving device 104 can use the first error correction data to generate first error-corrected coded video data to perform error correction operations on the first coded video data. Additionally, receiving device 104 can perform a first reconstruction operation to reconstruct blocks of the first frame based on the first coded video data. The first reconstruction operation can be controlled by the values of one or more parameters.
[0079] Furthermore, receiving device 104 can obtain first error-correcting data and first encoded video data from transmitting device 102. Receiving device 104 can apply an error correction process to modify the first encoded video data based on the first error-correcting data to generate first error-corrected encoded video data. Receiving device 104 can also apply a decoding process to reconstruct a first set of frames based on the first error-corrected encoded video data. Receiving device 104 can determine an extraction mode based on the first set of frames, indicating a mode in which encoded video data has not been transmitted. Receiving device 104 can send an extraction mode indication to transmitting device 102 indicating the determined extraction mode. Receiving device 104 can receive second error-correcting data and extracted video data from transmitting device 102. Extracted video data may include second encoded video data for which an extraction mode has been applied. The second encoded video data is generated based on a second set of frames of video data. Receiving device 104 can apply an error correction process to modify the second encoded video data based on the second error-correcting data to generate second error-corrected encoded video data. Receiving device 104 can apply a decoding process to reconstruct a second set of frames based on the second error-corrected encoded video data.
[0080] Figure 2 This is a block diagram illustrating example components of a transmitting and receiving device according to the technology of this disclosure. System 200 includes a transmitting device 102 and a receiving device 104. The transmitting device 102 is configured to transmit encoded video data to the receiving device 104. Figure 2 In one example, transmitting device 102 includes a video encoder 210, a channel encoder 212, and a puncturing unit 214. Receiving device 104 includes a de-puncturing unit 220, a channel decoder 222, a video decoder 224, a frame estimation unit 226, and a video encoder 228. In other examples, transmitting device 102 and receiving device 104 may include more, fewer, or different units. The processor 114 of transmitting device 102 (… Figure 1The receiving device 104's processor 130 can implement a video encoder 210, a channel encoder 212, and a punching unit 214. The receiving device 104's processor 130 can implement a de-punching unit 220, a channel decoder 222, a video decoder 224, a frame estimation unit 226, and a video encoder 228. Communication interface 118 ( Figure 1 () can represent sending device 102 to send and receive data. Communication interface 134 ( Figure 1 () can represent receiving device 104 to send and receive data.
[0081] The video encoder 210 of the transmitting device 102 can transmit from a video source (e.g., video source 120). Figure 1 The transmitting device 102 receives video data. The video data may include, for example, raw, unencoded video footage from video source 120. In some examples, the transmitting device 102's memory (e.g., memory 116) receives video data. Figure 1 The video encoder 210 can store video data. It can perform a video encoding process on the video data to generate encoded video data. The video encoding process can be "limited" in the sense that it can be relatively fast and consume fewer resources than more robust video compression processes (such as H.264 / AVC, H.265 / HEVC, or H.266 / VVC). The video encoding process may not reduce the number of bits representing the video data to the same extent as a more robust or complete video encoding process.
[0082] The video encoder 210 can perform a finite video encoding process in one of several ways. For example, in some examples, the video encoder 210 can perform a prediction process (such as intra-frame prediction) on each frame of video data to produce prediction data. The video encoder 210 can generate residual data based on the prediction data. For example, the video encoder 210 can subtract samples of the prediction data from corresponding samples of the original frame to determine samples of the residual data. The samples can be values indicating color values (such as Y, Cb, or Cr values in the YCbCr color gamut or red, green, or blue values in the RGB color gamut).
[0083] Video encoder 210 may apply a transform (e.g., Discrete Cosine Transform (DCT)) to the residual data to produce a transform block including transform coefficients. Additionally, video encoder 210 may quantize the transform coefficients. Video encoder 210 may apply entropy coding (e.g., Context Adaptive Binary Arithmetic Decoding (CABAC) or Exponential Columbus-Rice Decoding) to the syntax elements representing the quantized transform coefficients. Encoded video data may include entropy-coded syntax elements. In some examples, video encoder 210 applies the transform and / or quantization directly to the video data without first using intra-frame prediction. In some examples where video encoder 210 does not apply entropy coding, the encoded video data includes syntax elements representing quantized transform coefficients, unquantized transform coefficients, or residual data.
[0084] In the example where the video encoder 210 does not use inter-frame prediction, fewer memory read requests are required compared to a more robust video compression process that would require reading data from memory about previously decoded frames. Such memory read requests can be relatively time- and energy-intensive.
[0085] In some examples where the video data is multi-view video data, the video encoder 210 can perform multi-view video encoding to generate prediction data. For example, the video encoder 210 can use inter-view prediction to generate prediction data for blocks of non-anchor frames (e.g., macroblocks, decoding units, etc.). In some cases, inter-view prediction may involve determining the disparity vector of a block, where the disparity vector indicates the lateral displacement between the block and a corresponding block in one or more reference views.
[0086] The channel encoder 212 of the transmitting device 102 can apply a channel coding process to encode video data. The channel coding process prepares the encoded video data for transmission over a wireless communication channel (e.g., channel 230). Channel 230 can be a wireless sidelink channel 112 (…). Figure 1 (This can be a direct translation of the original text, but the full context is missing.) Channel-coded video data may include error correction data. Channel encoder 212 can generate error correction data in various ways. For example, channel encoder 212 can generate error correction data as convolutional codes or turbo codes. Error correction data can help receiving device 104 determine whether the received coded video data has changed during transmission via channel 230, and can help receiving device 104 correct such changes. A more detailed discussion of channel coding and channel decoding is provided below with reference to Figure 3.
[0087] In addition, Figure 2In the example, the puncturing unit 214 of the transmitting device 202 can apply bit puncturing processing to the error-correcting data to generate bit-punctured error-correcting data. The bit puncturing process can reduce the number of bits in the error-correcting data. For example, the puncturing unit 214 can perform an operation to remove bits from the error-correcting data according to the puncturing pattern.
[0088] Transmitting device 102 can transmit data, such as encoded video data and error-corrected data (e.g., bit-punctured error-corrected data), to receiving device 104 via channel 230. Channel 230 can introduce noise into the transmitted data. In some examples, channel 230 is a multipath channel, and the data transmitted in channel 230 can be time-varying. Receiving device 104 can receive the noise-modified data. Receiving device 104 can store the noise-modified data, at least temporarily, in a memory such as memory 132. Figure 1 ) in the memory.
[0089] The de-puncturing unit 220 can perform a de-puncturing operation on the received bit-punctured error-corrected data to reconstruct the error-corrected data. The de-puncturing operation can replace puncture symbols with neutral values according to an indication of the puncturing pattern. The de-puncturing operation can generate erase bits that indicate the presence of neutral symbols in the error-corrected data.
[0090] Channel decoder 222 can apply a channel decoding process to generate error-corrected coded video data based on error-corrected data and coded video data (such as coded video data received from the transmitting device and / or coded video data generated by the receiving device 104). For example, channel decoder 222 can modify the value of bit-coded video data according to any of a variety of error correction schemes such as low-density parity-check (LDPC) decoding or forward error correction (FEC).
[0091] Video decoder 224 can perform a video decoding process to reconstruct the image based on error-corrected coded video data. For example, video decoder 224 can apply an entropy decoding process to the bits of error-corrected coded video data to obtain quantization transform coefficients. Video decoder 224 can apply an inverse quantization operation to the quantization transform coefficients, apply an inverse transform to the inverse quantization transform coefficients to generate residual data, generate prediction data, and use the prediction data and residual data to reconstruct the image of the video data. Video decoder 224 can generate prediction data in the same manner as video encoder 210.
[0092] The frame estimation unit 226 can generate an estimate of the next frame of the video data. For example, the frame estimation unit 226 can extrapolate the next frame from two or more previously reconstructed frames. For example, in this example, the frame estimation unit 226 can segment a first previously reconstructed frame into blocks. For each block of the first previously reconstructed frame, the frame estimation unit 226 can determine one or more corresponding blocks of one or more blocks attached to the blocks in the previous reconstructed frames. The corresponding blocks of a block can be the best available match for the block. The frame estimation unit 226 can generate a prediction for the block based on one or more corresponding blocks of the block. The frame estimation unit 226 can use one-way prediction or two-way prediction to generate the prediction. Therefore, by generating a prediction for each block of the next frame, the frame estimation unit 226 can generate an estimate of the next frame. In some examples, the frame estimation unit 226 generates the next frame by applying global motion to the previously reconstructed frames.
[0093] In some examples, the frame estimation unit 226 can re-encode the current frame that the video decoder 224 has already decoded. The next frame of the video data can be the frame that follows the frame just decoded by the video decoder 224 in decoding order. In this example, the frame estimation unit 226 can perform intra-frame prediction or inter-frame prediction on blocks of the current frame. When performing inter-frame prediction on blocks, the frame estimation unit 226 can determine one or more motion vectors for the block. For example, the frame estimation unit 226 can determine that a particular block of the current frame has a motion vector with a magnitude m relative to a reference block in a reference frame having a frame order count (POC) distance p1 from the current frame. In this example, the current frame and the next frame can have a POC distance p2. The frame estimation unit 226 can determine a scaling factor s as p2 / p1. The frame estimation unit 226 can then scale the motion vector of the particular block by s (e.g., s The frame estimation unit 226 can determine the position in the next frame indicated by the scaling motion vector and set the sample at the determined position as the sample of a specific block in the current frame. The frame estimation unit 226 can repeat this process for each inter-frame prediction block of the current frame.
[0094] In some examples, the image estimation unit 226 may apply one or more filters to the prediction data. For example, the image estimation unit 226 may apply one or more unblocking filters, smoothing filters, adaptive loop filters, or other types of filters to the prediction data.
[0095] Video encoder 228 can perform the same finite video encoding process as video encoder 210 on the video data generated by frame estimation unit 226. For example, video encoder 228 can perform intra-frame prediction to generate prediction data. Video encoder 228 can use the prediction data and corresponding blocks of video data generated by frame estimation unit 226 to generate residual data. Video encoder 228 can apply transforms (e.g., DCT transform, DST transform, etc.) to the residual data to generate transform coefficients. Video encoder 228 can apply quantization to the transform coefficients. In addition, video encoder 228 can apply entropy coding to the syntax elements representing the transform coefficients.
[0096] As briefly mentioned above, the channel decoder 222 can apply the channel decoding process to channel-coded video data. Figure 3A and Figure 3B More information is provided about the channel encoding process performed by the channel encoder 212 and the channel decoding process performed by the channel decoder 222.
[0097] Specifically Figure 3A This is a block diagram illustrating an example channel coding process according to the technology of this disclosure. For each frame of encoded video data, the channel encoder 212 of the transmitting device 102 may apply a system decoding operation (e.g., a low-density parity-check (LDPC) decoding operation) to the system bits of the frame to generate error correction data for the frame. The system bits of the frame may include the encoded video data of the frame generated by the video encoder 210.
[0098] exist Figure 3A In the example, error correction data is labeled "error correction bits". For frame n, channel encoder 212 can generate error correction data 300A based on system bits 302A. Similarly, for frame n+1, channel encoder 212 can generate error correction data 300B based on system bits 302B. Channel encoder 212 can classify video data frames into anchor video frames and non-anchor frames. Channel encoder 212 can classify frames such that anchor frames appear periodically between frames. In some examples, if channel decoder 222 (e.g., from receiving device 104) receives an indication that an error exists in a frame, channel encoder 212 can classify the frame as an anchor frame. For each anchor frame, transmitting device 102 can transmit the encoded anchor frame and its error correction data. However, for non-anchor frames, transmitting device 102 can only transmit the error correction data for the non-anchor frames.
[0099] For example, in Figure 3AIn the example, frame n can be an anchor frame and frame n+1 is a non-anchor frame. Therefore, the transmitting device 102 can transmit the system bit 302A of frame n, the error correction data 300A of frame n, and the error correction data 300B of frame n+1, but does not transmit the system bit 302B of frame n+1.
[0100] Figure 3B This is a block diagram illustrating an example channel decoding process according to the technology of this disclosure. As mentioned above, channel decoder 222 can perform a channel decoding process on encoded video data to reconstruct the encoded video data. When processing anchor frames (e.g., frame n), channel decoder 222 can obtain error correction data (in the form of error correction data) of the anchor frames from de-punching unit 220. Figure 3B The middle is represented as y n ) and anchor screen system bits (in Figure 3B The Chinese character is represented as SI. n The system bits of the anchor frame can represent the encoded video data of the anchor frame. The channel decoder 222 can use the error correction data of the anchor frame to detect and / or correct errors in the system bits of the anchor frame. The video decoder 224 can use the obtained error-corrected encoded video data of the anchor frame to reconstruct the anchor frame. The receiving device 104 can store the reconstructed frames (including the reconstructed anchor frame and the reconstructed non-anchor frame) in the decoded frame buffer 350.
[0101] When processing non-anchor frames, the channel decoder 222 can obtain the system bits (denoted as SI) representing the encoded video data of the non-anchor frames. n+1 The video encoder 228 of the receiving device 104 can generate coded video data of the non-anchor frame based on the video data generated by the frame estimation unit 226. The channel decoder 222 can obtain error correction data of the non-anchor frame from the de-punching unit 220 (in...). Figure 3B The middle is represented as y n+1 The channel decoder 222 can then perform the same channel decoding process that it applied when processing the anchor frame. Therefore, the channel decoder 222 can use error correction data from the non-anchor frame to detect and / or correct "errors" in the system bits of the non-anchor frame. However, the "errors" in the system bits of the non-anchor frame are not due to noise in the channel 230 (as is the case with errors in the system bits of the anchor frame). Instead, the "errors" in the system bits of the non-anchor frame may be due to differences between the predicted version of the non-anchor frame and the original version of the non-anchor frame. Therefore, the channel decoder 222 can use the error correction data from the transmitted non-anchor frame as a mechanism for "correcting" prediction errors.
[0102] The video encoder 210 of transmitting device 102, the video decoder 224 of receiving device 104, and the video encoder 228 of receiving device 104 can perform video encoding and video decoding processes based on the values of one or more sets of parameters. In other words, the values of the parameters can control various aspects of the video encoding process performed by the video encoder 210, video encoder 228, and video decoder 224. In some examples, the parameters may include one or more of the following:
[0103] • Parameters indicating the color space (e.g., red-green-blue, Y-Cb-Cr, etc.)
[0104] Pixel extraction parameters
[0105] Parameters indicating DCT size
[0106] • Transmitted DCT coefficients
[0107] • A parameter indicating the number of bits for each DCT coefficient
[0108] • Quantization parameters (e.g., parameters indicating the quantization scheme, such as linear, Max-Lloyd, etc.).
[0109] Each of the video encoder 210, video decoder 224, and video encoder 228 requires the same parameter value. Therefore, according to one or more techniques of this disclosure, the transmitting device 102 can send the parameter value to the receiving device 104. The receiving device 104 can receive the sent parameter value. The video decoder 224 and video encoder 228 can use the parameter value during video decoding and video encoding processes.
[0110] In some examples, the values of the transmitted parameters are static or semi-static. For instance, in an example where the transmitted parameter values are static, the transmitting device 102 may transmit the parameter value to the receiving device 104 once, and the receiving device 104 may use the parameter values to operate at indeterminate time intervals. In an example where the transmitted parameter values are semi-static, the transmitting device 102 may periodically update the parameter values and resend the updated parameter values to the receiving device 104.
[0111] Transmitting device 102 can transmit the value of the parameter in one of a variety of ways. For example, in some examples, transmitting device 102 can use uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, or another type of message to transmit the value of the parameter to receiving device 104.
[0112] In the example where transmitting device 102 sends parameter values to receiving device 104, video encoder 210 can segment each of the color components (e.g., R, G, and B components; Y, Cb, and Cr components) of the video data frame into uniformly sized (M×M) blocks. Examples of such blocks may include macroblocks (MBs) and maximum decoding units (LCUs). The video encoder 210 of transmitting device 102 can calculate a transform (e.g., 2D-DCT) for each block, obtaining M 2 There are N transform coefficients. The video encoder 210 can assign a sort to the transform coefficients of a block. For example, the video encoder 210 can sort the transform coefficients of a block according to a zigzag scan order, where the zigzag scan order starts with the most important transform coefficient (e.g., the lowest frequency) and ends with the least important transform coefficient (e.g., the highest frequency). The video encoder 210 can select the top N transform coefficients. c There are N transformation coefficients, where N c This is a parameter value that indicates the number of transform coefficients sent. The video encoder 210 can discard unselected transform coefficients.
[0113] Additionally, the video encoder 210 can quantize the selected transform coefficients. For example, the selected transform coefficients can have a range from 0 to N. c In the case of index i = -1, the parameter can include a bit width parameter corresponding to different index values (e.g., B). i i = 0, 1, ..., N c -1). For the selected transformation coefficients d i For each of the following, the video encoder 210 can quantize the selected transform coefficient d using the following equation. i .
[0114] c i = round ( α d i B i (1)
[0115] In the equation above, c i The transformation coefficient d i The quantized version, where α is the scaling constant, B i `i` is the bit width parameter for index `i`, and `round` is the function that rounds to the nearest integer. Therefore, in the example where B0 is 8, B1 is 4, and B2 is 4, the quantization transform coefficients can be, for example, c0=00100011, c1=0110, c2=1001, etc.
[0116] The video encoder 228 of the receiving device 104 can generate prediction data for the image, generate residual data based on the prediction data, and apply one or more transforms to the residual data to generate a transform block including transform coefficients. The video encoder 228 may need to use the same bit width parameters as the video encoder 210 so that the channel decoder 222 can correctly associate specific system bits with the corresponding error correction data received from the de-puncturing unit 220.
[0117] In some examples of this disclosure, receiving device 104 can determine the value of one or more of these parameters without sending the values of these parameters to receiving device 104 by transmitting device 102. Examples where receiving device 104 determines the value of one or more parameters without transmitting the values to receiving device 104 can achieve a better compression-distortion tradeoff with lower control signaling overhead. For example, receiving device 104 can determine the value of one or more parameters without transmitting the values to receiving device 102. c Determine the number of DCT coefficients (N) given the value of N. c For example, in this example, receiving device 104 can determine N to achieve the desired peak signal-to-noise ratio (P-SNR). c The minimum value. In other words, the receiving device 104 can achieve the desired P-SNR of N. c The minimum value is determined as follows:
[0118] (2)
[0119] In the equation above, c i The transform coefficients (e.g., DCT coefficients) with index i, SNR d It is the expected P-SNR, M 2 -1 represents the maximum number of transform coefficients. In some examples, receiving device 104 can evaluate the transform coefficients (N) for each frame (or other segment) at once based on robust prediction of the frame. c The number of ( ). Periodic resets of parameter values can be applied to prevent error propagation.
[0120] In another example, receiving device 104 can determine the number of quantized bits based on the predicted image, rather than receiving the number of quantized bits from transmitting device 102. For example, in this example, receiving device 104 can calculate the probability distribution of quantized and unquantized coefficients:
[0121] (3)
[0122] In the equation above, It is the probability distribution of the quantization coefficients, p i It is the probability distribution of the unquantized coefficients, c iThe transformation coefficient d i The quantified version.
[0123] The receiving device 104 can determine the quantized bit B based on the entropy ratio of the quantized coefficients to the unquantized coefficients. i The quantities are as follows:
[0124] (4)
[0125] Therefore, receiving device 104 can evaluate B once for each frame (or segment) based on the prediction of the frame generated by frame estimation unit 226. i .
[0126] Figure 4 This is a flowchart illustrating an example operation of the transmitting device 102 according to the technology of this disclosure. Figure 4 In the example, the video encoder 210 of the transmitting device 102 can obtain video data (400). For example, the video encoder 210 can obtain video data from the video source 120. Additionally, the video encoder 210 can perform video encoding on the video data to generate encoded video data (402). For example, the video encoder 210 can apply intra-frame prediction to generate prediction data, generate residual data based on the prediction data and the original video data, and apply a transform (e.g., DCT) to blocks of the residual data to generate transform blocks. The video encoder 210 can quantize the transform coefficients of the transform blocks. Furthermore, in some examples, the video encoder 210 can apply entropy encoding to syntax elements representing the quantized transform coefficients. In some examples, the video encoder 210 can implement a reconstruction loop that can apply entropy decoding, inverse quantization, and one or more inverse transforms to reconstruct the residual data. The video encoder 210 can apply the prediction data and the reconstructed residual data to reconstruct the video data. In some examples, the video encoder 210 applies one or more filters to the reconstructed video data, such as a deblocking filter, an adaptive loop filter, a sample adaptive offset filter, etc. The video encoder 210 can use reconstructed video data as reference data for intra-frame prediction.
[0127] The video encoder 210 can perform the video encoding process based on the values of one or more parameters. For example, the video encoder 210 quantizes transform coefficients according to specific quantization parameters, uses a specific color space, etc.
[0128] The channel encoder 212 of transmitting device 102 can perform channel coding on the encoded video data to generate error correction data (404). Transmitting device 102 can transmit the encoded video data and error correction data to receiving device 104 (406), for example, via channel 230. In some examples, transmitting device 102 can selectively transmit portions of the encoded video data and transmit other portions of the encoded video data. For example, transmitting device 102 transmits encoded video data for some frames but not for others. In another example, transmitting device 102 can transmit encoded video data for some transform blocks of a frame but not for others. In some examples, transmitting device 102 can transmit a certain number of the most significant bits of the transform coefficients but not the less significant bits.
[0129] In some examples, transmitting device 102 may also send the values of one or more parameters to receiving device 104. The values of these parameters can control how receiving device 104 reconstructs the video data. For example, parameters may include a transform size parameter, which indicates the size of transform blocks in the encoded video data generated by video encoder 210. In this example, receiving device 104 would need to interpret the received encoded video data according to the same transform block size in order to properly reconstruct the video data. In other examples, parameters may include parameters indicating the number of transform coefficients, bit width parameters, etc.
[0130] Therefore, in Figure 4 In the example, transmitting device 102 can obtain video data from a video source. Transmitting device 102 can generate encoded video data for a first frame and encoded video data for a second frame of the video data based on a parameter set. Transmitting device 102 can perform channel coding on the encoded video data for the first and second frames to generate error-corrected data for the first and second frames. Transmitting device 102 can transmit the encoded video data for the first frame, the error-corrected data for the first frame, and the error-corrected data for the second frame. In some examples, transmitting device 102 can transmit parameter values to receiving device 104.
[0131] Figure 5 This is a flowchart illustrating an example operation of the receiving device 104 according to the technology of this disclosure. Figure 5 In the example, receiving device 104 can obtain first coded video data and first error correction data (500) from transmitting device 102. The first coded data represents one or more blocks of a first frame of video data. The first error correction data can provide error correction information about the blocks of the first frame.
[0132] Receiver 104 can use the first error-correcting data to generate first error-corrected coded video data to perform error correction operations on the first coded video data (502). For example, channel decoder 222 of receiver 104 can use the first error-correcting data to perform low-density parity-check (LDPC) decoding on the first coded video data. In other examples, channel decoder 222 can use the error-correcting data in other error correction algorithms, such as forward error correction (FEC) or turbo decoding. Performing error correction operations on the first coded video data can remove errors introduced by noise in channel 230.
[0133] The video decoder 224 of the receiving device 104 can perform a first reconstruction operation (504) to reconstruct blocks of the first frame based on the first error-corrected coded video data. The first reconstruction operation is controlled by the values of one or more parameters. For example, the video decoder 224 can perform an inverse transform on the transformed blocks of the first error-corrected coded video data to obtain residual data. Alternatively, in this example, the video decoder 224 can, for example, use intra-frame prediction to generate prediction data. In this example, the video decoder 224 can use the prediction data and the residual data to reconstruct blocks of the first frame.
[0134] In some examples, video encoders 210 and 228 can use quantization parameters to quantize transform coefficients generated from scene-based prediction data to generate encoded video data. When performing a reconstruction operation, video decoder 224 can use quantization parameters to inversely quantize the transform coefficients of the error-correcting encoded video data. In some examples, transmitting device 102 and / or receiving device 104 can calculate quantization parameters based on the entropy ratio of quantized transform coefficients to unquantized transform coefficients, for example, as described above.
[0135] In some examples, the parameters include a transform size parameter. As part of generating encoded video data, video encoders 210 and 228 may apply a forward transform with a transform size indicated by the transform size parameter to the sample domain data of the frame (e.g., prediction sample data or residual data). As part of performing a reconstruction operation, video decoder 224 may apply an inverse transform with a transform size indicated by the transform size parameter to the transform coefficients of the error-correcting encoded video data.
[0136] In some examples, the parameters include a parameter indicating the number of transform coefficients. As part of generating coded video data, video encoders 210 and 228 may include a set of transform coefficients in the coded video data, wherein the set of transform coefficients includes a transform coefficient indicating the number. When performing a reconstruction operation, video decoder 224 can parse the set of transform coefficients including the transform coefficient indicating the number from the error-corrected coded video data. Furthermore, in some examples, receiving device 104 may receive coded video data and error-corrected data from transmitting device via a communication channel, and receiving device 104 may apply an optimization process, wherein the optimization process determines the number of transform coefficients based on the signal-to-noise ratio of the data transmitted on the communication channel, for example, as described above.
[0137] In some examples, the parameters include a bit width parameter for multiple index values. For each corresponding index value among the multiple index values, performing the reconstruction operation may include resolving a first set of bits from the error-corrected coded video data. The first set of bits may indicate transform coefficients with the corresponding index value, and the number of bits in the first set of bits is equal to the bit width indicated by the bit width parameter of the corresponding index value. As part of generating the coded video data, video encoders 210 and 228 may include a second set of bits in the coded video data. The second set of bits may indicate transform coefficients with the corresponding index values, and the number of bits in the second set of bits is equal to the bit width indicated by the bit width parameter of the corresponding index value. Video decoder 224 may resolve a third set of bits from the error-corrected coded video data. The third set of bits may indicate transform coefficients with the corresponding index values, and the number of bits in the third set of bits is equal to the bit width indicated by the bit width parameter of the corresponding index value.
[0138] Other parameters may include one or more of the following: color space, transform size, quantization parameter, number of transform coefficients in the first coded video data, or number of bits per transform coefficient in the first coded video data.
[0139] The receiving device 104 can obtain second error correction data (506) from the transmitting device 102. The second error correction data provides error correction information about one or more blocks of a second frame of video data.
[0140] Additionally, the frame estimation unit 226 of the receiving device 104 can estimate the second frame (508) based on one or more previously reconstructed frames (such as the first frame). Estimating the second frame includes at least in part a prediction of blocks of the second frame of video data based on blocks of the first frame. For example, the frame estimation unit 226 can use a combination of inter-frame prediction, intra-frame and inter-frame prediction, or other video decoding tools (e.g., as described elsewhere in this disclosure) to generate the prediction data.
[0141] The video encoder 228 of the receiving device 104 can generate second coded video data (510) based on an estimated second frame. For example, the video encoder 228 of the receiving device 104 can generate residual data based on prediction data. For example, the video encoder 228 can perform intra-frame prediction to generate second prediction data based on the estimated second frame. The video encoder 228 can then generate residual data by subtracting the second prediction data from the prediction data generated by the frame estimation unit 226. The video encoder 228 can then generate transform blocks by applying one or more forward transforms to the residual data. The video encoder 228 can perform the same process as the video encoder 210 of the transmitting device 102, and therefore will need to use the same parameters as the video encoder 210.
[0142] The channel decoder 222 of the receiving device 104 can perform a channel decoding process to generate second error-corrected coded video data (512) based on the second error-corrected data and the second coded video data. The channel decoder 222 of the receiving device 104 can perform the same process as the channel decoder 222 when generating the first error-corrected coded video data to generate the second error-corrected coded video data.
[0143] The video decoder 224 of the receiving device 104 can perform a second video decoding process to reconstruct blocks (514) of the second frame based on the second error-corrected coded video data. The second reconstruction operation is controlled by the value of a parameter. The video decoder 224 can perform the second reconstruction operation in the same manner as the first reconstruction operation. In this way, the receiving device 104 can reconstruct the video data of the frame (or block) without receiving all the coded video data of each of the frames (or blocks).
[0144] As described above, the puncturing unit 214 of the transmitting device 102 can perform a bit puncturing operation on the error-correcting data generated by the channel encoder 212. Bit puncturing involves selectively discarding some error-correcting data before the transmitting device 102 transmits the error-correcting data. The discarded bits are generally the least important for performing error correction. The depuncturing unit 220 of the receiving device 104 can perform an inverse bit puncturing operation (i.e., a bit depuncturing operation), where the inverse bit puncturing operation is the opposite of the bit puncturing operation performed by the puncturing unit 214. The puncturing unit 214 can perform the bit puncturing operation based on a set of one or more puncturing parameters. In different examples, the puncturing parameters can be predefined, static, or semi-static.
[0145] According to one or more techniques of this disclosure, transmitting device 102 can perform a decimation process, which can reduce memory bandwidth and enhance compression. For example, the video encoder 210 of transmitting device 102 can divide a frame of video data into a grid of blocks (e.g., MB, LCU, etc.) and can generate a transform block for each block. Transmitting device 102 will need to store each transform block intended for transmission to receiving device 104 into a memory (e.g., memory 116). Transmitting device 102 can then retrieve the stored transform blocks from the memory for channel coding and ultimately for transmission. These writes to and reads from the memory increase time and energy requirements. These time and energy requirements may be directly related to the amount of data to be written and read. Therefore, reducing the amount of data to be written to and read from the memory may be advantageous.
[0146] Performing the decimation process can reduce the amount of data written to and read from memory in the transform block. Performing the decimation process can also reduce the amount of data sent from transmitting device 102 to receiving device 104. In some examples, while video encoder 210 is encoding the current block of the current frame, video encoder 210 can generate a transform block for the current block. Additionally, channel encoder 212 of transmitting device 102 can determine whether the current block is decimation-targeted based on the decimation mode. If the current block is decimation-targeted (i.e., the transform block is a "non-anchor transform block"), channel encoder 212 can reduce the number of bits in the non-anchor transform block before storing the transform block in memory. If the current block is not decimation-targeted (i.e., the transform block is an "anchor block"), channel encoder 212 does not reduce the number of bits in the anchor block. Channel encoder 212 can perform the decimation process after generating error correction data. Therefore, the error correction data generated by channel encoder 212 for non-anchor transform blocks (and potentially sent to receiving device 104) can be based on the complete set of bits in the transform block rather than the reduced number of bits.
[0147] Figure 6This is a conceptual diagram illustrating an example extraction pattern 600 according to the technology disclosed herein. Figure 6 The example illustrates a grid of DCT blocks. A DCT block is a block of transform coefficients generated by applying a DCT transform to video data, such as residual or sample data. In other examples, a DCT block can be a transform block generated using other types of transforms. Figure 6 In extraction mode 600, the "X" marker indicates the DCT block (i.e., the non-anchor transform block) targeted for extraction. Therefore, in Figure 6 In the examples, decimation mode 600 decimates the DCT block by a factor of 2 in both the horizontal and vertical directions. In some examples that may be referred to herein as “full” decimation, the video encoder 210 can reduce the number of bits in the non-anchor transform block to zero.
[0148] Therefore, in some examples, the decimation mode defines the mode of anchor transform blocks and non-anchor transform blocks in the frame. Receiving device 104 may receive the system bits of anchor transform blocks but not the system bits of non-anchor transform blocks. The system bits of anchor transform blocks may represent the transform coefficients in the anchor transform block. The system bits of non-anchor transform blocks may represent a reduced-bit-depth version of the original transform coefficients in the non-anchor transform block. Error correction data may include error correction data for both anchor and non-anchor transform blocks. The error correction data for non-anchor transform blocks is based on the original transform coefficients in the non-anchor transform block. As part of generating error-corrected coded video data, channel decoder 222 may use the error correction data for anchor transform blocks to perform error correction on the system bits of the anchor transform blocks. Channel decoder 222 may use the error correction data for non-anchor transform blocks to perform error correction on portions of the coded video data corresponding to the non-anchor transform blocks. In some examples, receiving device 104 may determine the decimation mode and send the decimation mode to transmitting device 102.
[0149] In some examples, transmitting device 102 stores encoded bits (encoded video data and error correction data) in a circular buffer. Transmitting device 102 uses two parameters to select which bits in the circular buffer to transmit. The first parameter is a start position, and the second parameter indicates the number of consecutive bits to transmit. The start position can have various values to support selective transmission and non-transmission of system bits. The start position can be selected to skip the transmission of specific system bits (i.e., bits of encoded video data) without skipping the transmission of error correction data. Therefore, the decimation of non-anchored transform blocks can be easily accomplished by manipulating the first and second parameters, causing transmitting device 102 not to transmit bits of the non-anchored transform block.
[0150] In some examples, the channel encoder 212 applies a hybrid decimation method, wherein the hybrid decimation method does not reduce the number of bits in any target “non-anchor” transform block to zero, but rather reduces the number of bits in the transform coefficients of the non-anchor transform block. For example, in Figure 6 In the example, channel encoder 212 may reduce the number of bits in each transform coefficient in the DCT block marked with "X" by a predetermined number (e.g., 2, 4, 5, etc.). Channel encoder 212 does not reduce the number of bits in transform coefficients that are not the target of the decimation mode.
[0151] The channel decoder 222 of the receiving device 104 can receive the remaining reduced bits of the non-anchor transform block and the error correction data of the non-anchor transform block. As part of the channel decoding process, the channel decoder 222 can use the error correction data of the non-anchor transform block to perform an error correction process, wherein the error correction process recovers the bits of the removed non-anchor transform block. This error correction process can be the same error correction process used by the channel decoder 222 to correct errors introduced by noise in the channel 230. In summary, the channel encoder 212 generates error correction data because error correction data will be needed to correct unavoidable noise in the channel 230, but this same error correction data is used to recover bits as if the noise in the channel 230 just happened to corrupt the least significant bits of a specific transform coefficient in a specific transform block in a specific frame. Therefore, the number of bits transmitted in the channel 230 can be effectively reduced.
[0152] In some examples, the channel encoder 212 may generate a correlation matrix based on the set of transform blocks before performing any decimation process on any non-anchor transform block in the set of transform blocks of the image. The correlation matrix includes values indicating the level of correlation between transform coefficients at corresponding locations within a transform block. For example, the correlation matrix may include correlation values for the DC transform coefficients (i.e., the top-left transform coefficients) of the set of transform blocks. If the differences between the DC transform coefficients are relatively small, the correlation values of the DC transform coefficients can be relatively high. Conversely, if the differences between the DC transform coefficients are relatively large, the correlation values of the DC transform coefficients can be relatively small. Each correlation value can be a value between 0 and 1.
[0153] In some examples, the channel encoder 212 can use the following formula to calculate the correlation values of the DC transform coefficients:
[0154] (5)
[0155] In equation (5) above, l represents the interval between transform blocks containing DC coefficients, and N represents the number of transform blocks to which the calculation is performed. The function y is the transform coefficient. The line above y represents conjugate. If each consecutive transform block is used, the interval can be 1; if alternating transform blocks are used, the interval can be 2, and so on. The channel encoder 212 can calculate the correlation values of the corresponding AC transform coefficients (i.e., non-DC transform coefficients) in the same manner. In this disclosure, the corresponding transform coefficients occupy the same positions within the transform block. Therefore, by calculating the autocorrelation value of each transform coefficient in the transform block, the channel encoder 212 can generate the correlation matrix of the transform block. The channel encoder 212 can repeat the process of generating the correlation matrix of each transform block because the channel encoder 212 will use transform coefficients from different transform blocks when calculating the correlation values.
[0156] Transmitting device 102 can send the correlation matrix along with encoded video data and error correction data to receiving device 104. The channel decoder 222 of receiving device 104 can use the correlation matrix as part of the process of recovering the original bit width of the non-anchored transform block. For example, continuing with the example of DC transform coefficients, after the error correction process is applied, channel decoder 222 can obtain the values of the non-anchored DC transform coefficients (i.e., the DC transform coefficients in the non-anchored transform block).
[0157] The error correction process can use correlation values to estimate the non-anchor transform coefficients. For example, if we treat all even-numbered transform blocks as anchor transform blocks and non-even-numbered transform blocks as non-anchor transform blocks, the channel decoder 222 can estimate the values of the non-anchor transform coefficients in transform block n in the following way:
[0158] y[n]=(Ryy[1] y[n+1]+Ryy[3] y[n+3]+Ryy[3] y[n+5]…) / (Ryy[1]+Ryy[3]+Ryy[3]…) (6)
[0159] In equation (6) above, R yy [1] indicates the correlation value in the correlation matrix of the transform block with index 1 (i.e., the non-even, non-anchor transform block), y[n+1] indicates the corresponding transform coefficient in the anchor transform block with index n+1, R yy[2] indicates the correlation value in the correlation matrix of the transform block with index 3, y[n+3] indicates that the corresponding transform coefficient in the anchor transform block is index n+3, and so on. The number of transform blocks used in equation (6) can be configurable. In this way, the estimated value of the non-anchor transform coefficient can be considered as a weighted average of the corresponding transform coefficients in the anchor block, weighted by the correlation values at corresponding positions in the correlation matrix. In other words, the channel decoder 222 can interpolate the value of the non-anchor transform coefficient based on the correlation matrix. The channel decoder 222 can output the calculated transform coefficient value to the video decoder 224. The channel decoder 222 can perform this process for other transform coefficients. Using the correlation matrix in this way can improve the quality of the reconstructed video data.
[0160] In some examples, the video encoder 210 may reduce the bit width of each transform coefficient in the target transform block by the same amount. In other examples, the video encoder 210 may reduce the bit width of different transform coefficients in the target transform block by different amounts. In some examples, the amount by which the video encoder 210 reduces the bit width of the transform coefficients is related to the distance of the transform coefficient from the anchor transform block. The anchor transform block is a transform block that is not the target of the decimation pattern.
[0161] Figure 7 This is a flowchart illustrating an example operation of a transmitting device 102 for hybrid extraction of transform blocks according to the technology of this disclosure. Figure 7 In the example, the sending device 102 can receive data from the video source 120 (…). Figure 7 The video encoder 210 of the transmitting device 102 can generate transform blocks based on the video data (702). For example, the video encoder 210 can generate prediction blocks by performing intra-frame prediction on blocks of video data. The video encoder 210 can use the prediction blocks to generate residual data. The video encoder 210 can generate transform blocks by applying a transform (such as DCT, DST, or other transforms) to the residual data. In other examples, the video encoder 210 can generate transform blocks by applying a transform directly to blocks of video data.
[0162] Then, channel encoder 212 can determine which transform blocks are anchor transform blocks (704) based on the decimation mode. For example, in channel encoder 212 using Figure 6 In the example of decimation mode 600, video encoder 210 can determine that every other transform block in the horizontal and vertical directions is an anchor transform block. In other examples, channel encoder 212 can use other decimation modes to determine which transform blocks are anchor transform blocks. Channel encoder 212 can store the anchor transform blocks in the memory of transmitting device 102 (e.g., memory 116). Figure 1 ))(706).
[0163] The channel encoder 212 can compute the correlation matrix (708) of the transform block set. Each transform block set includes one or more anchor transform blocks and one or more non-anchor transform blocks. For example, each transform block set may correspond to Figure 6 The different rows of the transform blocks in the matrix. In another example, each set of transform blocks can correspond to a group of 2 transform blocks multiplied by 2 transform blocks. The number of values in the correlation matrix of the transform block set is the same as the number of transform coefficients in each transform block. Each value in the correlation matrix corresponds to a different position within the transform coefficient block. For example, the value at position (0,0) in the correlation matrix corresponds to the transform coefficient at position (0,0) of each transform coefficient block in the transform block set, the value at position (0,1) in the correlation matrix corresponds to the transform coefficient at position (0,1) of each transform coefficient block in the transform block set, and so on. The transform coefficient at position (0,0) of the transform coefficient block can be called the DC coefficient, and all other transform coefficients can be called the AC coefficient.
[0164] Additionally, the channel encoder 212 can generate a bit-reduced non-anchor transform matrix (710). The non-anchor transform matrix is a transform matrix other than the anchor transform matrix. For example, refer to... Figure 6 The transformation matrix marked with X can be an anchorless transformation matrix. The transformation coefficients in the bit-reduced anchorless transformation matrix can include fewer bits than in the original version of the anchorless transformation matrix. The transmitting device 102 can then transmit the anchor transformation block, the anchorless transformation block, the bit reduction value, the correlation matrix, and the error correction data (712).
[0165] Channel encoder 212 can reduce the bits in the non-anchored transform coefficients in one of a variety of ways. For example, channel encoder 212 can determine the bit reduction value for each transform coefficient in the non-anchored transform coefficient block. In this example, to calculate the bit reduction value, channel encoder 212 can calculate the interpolated value of the transform coefficient based on the correlation matrix. Channel encoder 212 can use the above equation (6) to calculate the interpolated value. Channel encoder 212 can then subtract the interpolated value of the transform coefficient from the original value of the transform coefficient to calculate a first distortion value. Channel encoder 212 can then reduce the number of bits of the original value of the transform coefficient by 1. Channel encoder 212 can subtract the interpolated value from the reduced bit original value of the transform coefficient to calculate a second distortion value. Channel encoder 212 can determine whether the second distortion value is acceptable based on the first distortion value and the second distortion value. For example, channel encoder 212 can calculate the mean or maximum squared error from the interpolated value. Channel encoder 212 can determine whether the second distortion value is acceptable by comparing the second distortion value with a predefined threshold.
[0166] If the second distortion value is acceptable, the channel encoder 212 can reduce the number of bits of the original transform coefficient value and repeat the process. If the second distortion value is not acceptable, the channel encoder 212 can increase the number of bits of the original transform coefficient value. The number of bits obtained by reducing the original transform coefficient value is the bit reduction value.
[0167] In some examples, channel encoder 212 can determine the bit reduction value of the non-anchored transform coefficient block as a whole. In this example, to calculate the bit reduction value of the transform coefficient block, channel encoder 212 can calculate the interpolated value of each transform coefficient based on the correlation matrix, for example, as described above. Then, video encoder 210 can subtract the interpolated value of the transform coefficient from the original value of the transform coefficient and use the resulting difference to calculate a first distortion value. For example, video encoder 210 can calculate the first distortion value as a mean square error. Video encoder 210 can then reduce the number of bits of the original value of each transform coefficient by one. Video encoder 210 can subtract the interpolated value from the reduced bit original value of the transform coefficient and use the resulting value to calculate a second distortion value (e.g., using mean square error). Channel encoder 212 can determine whether the second distortion value is acceptable based on the first and second distortion values. If the second distortion value is acceptable, video encoder 210 can again reduce the number of bits of the original value of the transform coefficient and repeat the process. If the second distortion value is not acceptable, the video encoder 210 may increase the number of bits of the original value of the transform coefficient. The resulting number of bits by which the original value of the transform coefficient is reduced is the bit reduction value.
[0168] Therefore, in some examples, transmitting device 102 can obtain video data from a video source. Transmitting device 102 can generate transform blocks based on the video data. Transmitting device 102 can determine which transform blocks are anchor transform blocks. Transmitting device 102 can calculate the correlation matrix of the transform block set. Additionally, transmitting device 102 can generate a bit-reduced non-anchor transform matrix. Transmitting device 102 can send the anchor transform blocks, non-anchor transform blocks, and correlation matrix to the receiving device. In some examples, transmitting device 102 can receive an indication of the decimation mode from receiving device 104.
[0169] Figure 8 This is a flowchart illustrating an example operation of a receiving device 104 for hybrid decimation of a transform block according to the technology of this disclosure. Figure 8 In the example, receiving device 104 can receive anchor transform blocks, non-anchor transform blocks, one or more bit reduction values, and the correlation matrix (800) of the non-anchor blocks.
[0170] In addition, Figure 8In the example, the channel decoder 222 of the receiving device 104 can compute the interpolated value (802) of the current non-anchor transform coefficient. The current non-anchor transform coefficient is the transform coefficient of one of the non-anchor transform blocks. The channel decoder 222 can compute the interpolated value of the current non-anchor transform coefficient based on the correlation matrix of the non-anchor blocks. The channel decoder 222 can compute the interpolated value of the current non-anchor transform coefficient in one of various ways. For example, in some examples, the channel decoder 222 can apply a machine learning model, wherein the machine learning model takes one or more non-anchor transform coefficients (including the current non-anchor transform coefficient), one or more anchor transform coefficients, and the correlation matrix of the non-anchor transform coefficient blocks as input. In this example, the machine learning model can output the interpolated value of the current non-anchor transform coefficient. In this example, the machine learning model can be implemented as a neural network model, a support vector machine, a regression model, or another type of machine learning model.
[0171] In another example, the correlation matrix may include values indicating the correlation between the current non-anchor transform coefficient and each corresponding anchor transform coefficient in one or more blocks of anchor transform coefficients. The video decoder 224 may calculate the interpolated values of the current non-anchor transform coefficients as follows:
[0172] (7)
[0173] In the equation above, t int These are the current non-anchor transform coefficients, a i It is the anchor transformation coefficient, c i It indicates the relationship between the current non-anchor transformation coefficient and a. i The correlation value between them, and n indicates the number of anchor transform coefficients from which the current non-anchor transform coefficients are derived. In this example, c0 to c n The values can be added together to equal 1.
[0174] Additionally, the video decoder 224 can calculate the reconstructed values (804) of the non-anchor transform coefficients. The video decoder 224 can calculate the reconstructed values of the non-anchor transform coefficients based on the interpolated values of the current non-anchor transform coefficients and the transmitted values of the non-anchor transform coefficients. The transmitted values of the non-anchor transform coefficients are included in the received non-anchor transform block. In some examples, the video decoder 224 calculates the reconstructed values of the non-anchor transform coefficients as the average of the interpolated values of the current non-anchor transform coefficients and the transmitted values of the non-anchor transform coefficients.
[0175] Video decoder 224 can determine whether any remaining non-anchor transform coefficients exist in the non-anchor transform block (806). If one or more remaining non-anchor transform coefficients exist in the non-anchor transform block (the "yes" branch of 806), video decoder 224 can repeat steps 802-806 with another non-anchor transform coefficient. Video decoder 224 can continue doing this until there are no remaining non-anchor transform coefficients (the "no" branch of 806). In this way, video decoder 224 can compute the reconstructed value for each non-anchor transform coefficient.
[0176] In this manner, receiving device 104 can receive the system bits of the anchor transform block, the system bits of the non-anchor transform block, and the correlation matrix. The system bits of the anchor transform block can represent the transform coefficients in the anchor transform block. The system bits of the non-anchor transform block can represent a reduced-bit-depth version of the original transform coefficients in the non-anchor transform block. As part of performing the reconstruction operation, receiving device 104 can calculate an interpolated value of the non-anchor transform coefficient for each non-anchor transform coefficient in the non-anchor transform block, based on the correlation matrix and the corresponding anchor transform coefficient. Reconstructed values of the non-anchor transform coefficients can be calculated at the receiving device based on the interpolated values and values of the non-anchor transform coefficients in the error-corrected coded video data.
[0177] In some examples, receiving device 104 can adaptively select a decimation mode for reducing or eliminating bits in a particular transform block. In such an example, receiving device 104 can transmit the selected decimation mode back to transmitting device 102. Transmitting device 102 can then use the selected decimation mode in one or more frames of video data.
[0178] Figure 9 This is a conceptual diagram illustrating an example decimation mode 900 adaptively selected by receiving device 104 according to one or more techniques of this disclosure. Figure 6 In contrast to extraction mode 600, non-anchor blocks in extraction mode 900 do not necessarily occur at regular intervals or intermittently.
[0179] Receiving device 104 can determine the decimation mode based on information about previous frames of video data. The previous frames may or may not have been decimated. In some examples, receiving device 104 can send a request to sending device 102 for an undecimated version of the frame. Sending device 102 can respond to the request by sending the undecimated version of the frame to receiving device 104. After receiving the undecimated version of the frame, receiving device 104 can determine the decimation mode based on the undecimated version. For example, receiving device 104 can perform a rate-distortion optimization process, wherein the process evaluates multiple potential decimation modes to identify which decimation mode yields the optimal combination of bit rate and distortion.
[0180] In some examples, transmitting device 102 and receiving device 104 may continue to use the selected decimation mode for a predetermined number of frames, after which transmitting device 102 and / or receiving device 104 may adaptively select another decimation mode. In some examples, transmitting device 102 may send a message to receiving device 104 requesting transmitting device 102 to select another decimation mode. In some examples, receiving device 104 may determine that an event or condition has occurred that would make selecting another decimation mode advantageous. For example, when receiving device 104 determines that a scene change has occurred, motion in the video data has crossed one or more thresholds, or other characteristics of the video data have changed, receiving device 104 may determine that selecting another decimation mode would be advantageous.
[0181] The receiving device 104 can signal the selected decimation mode to the transmitting device 102 in one of a variety of ways. For example, the receiving device 104 can signal the selected decimation mode to the transmitting device 102 by indicating a difference from an existing decimation mode, such as the currently used decimation mode. For example, in this example, the receiving device 104 can indicate the selected decimation mode to the transmitting device 102 by specifying a change in downsampling or upsampling along a specific axis, specifying a change in a specific area of the image, eliminating a specific transform block, enabling a specific transform block, etc.
[0182] In some examples, there may be a predefined mapping from index values to predefined decimation patterns. In such an example, receiving device 104 can select a decimation pattern from the predefined decimation patterns and send the index value of the selected decimation pattern to transmitting device 102 via a signal.
[0183] Figure 10 This is a block diagram illustrating example components of a transmitting device and a receiving device according to the technology of this disclosure. Figure 10 In the example, the transmitting device 102 may include a... Figure 2 The same components shown. However, in Figure 10 In the example, receiving device 104 may additionally include reliability unit 1002. Unless otherwise stated, Figure 2 and Figure 10 The similarly named components of the transmitting device 102 and the receiving device 104 perform the same function.
[0184] Generally, the more significant bits (MSB) of a transform coefficient are easier to predict accurately than the less significant bits (LSB). This is because the MSB translates to a higher Euclidean distance in the video frame domain. Additionally, blocks of video frames experiencing high motion are more difficult to predict accurately than blocks in low-motion regions.
[0185] According to one or more techniques disclosed herein, transmitting device 102 and receiving device 104 can implement a system in which bit-level reliability values are used. The use of bit-level reliability values allows transmitting device 102 and receiving device 104 to correctly weight prior information. This can lead to improved system performance (e.g., reduced data transmission volume and / or improved video quality). For example, if “soft” information is used, decoding performance can be improved. In other words, when receiving device 104 is “informed” of the reliability of each bit (what the prior probability is that the bit value is “0” or “1”), receiving device 104 can utilize this information and performance can be improved.
[0186] exist Figure 10 In one example, the reliability unit 1002 of the receiving device 104 receives encoded video data from the video encoder 228 of the receiving device 104. The encoded video data may include transform coefficients of transform blocks of video data. Additionally, in some examples, the reliability unit 1002 receives prediction quality information from the image estimation unit 226 of the receiving device 104.
[0187] During constant scaling, for each bit position of the transform coefficients in the transform block, the prediction quality information includes a reliability value for that bit position. For example, the most significant bit of a transform coefficient has a first reliability value, the second most significant bit has a second reliability value, the third most significant bit has a third reliability value, and so on. The reliability value of a bit position is a measure of the probability that the bit at that position has an incorrect value. For example, a bit at a bit position may have an incorrect value when the bit value is predicted as "0" but actually is "1" or when the bit is predicted as "1" but actually is "0".
[0188] The image estimation unit 226 can determine the reliability value of a bit position by collecting statistics on the error rate occurring in bits at that bit position. For example, the image estimation unit 226 can determine the probability that the most significant bit of the transform coefficient contains an error, the probability that the second most significant bit of the transform coefficient contains an error, the probability that the third most significant bit of the transform coefficient contains an error, and so on. The image estimation unit 226 can collect these statistics by counting the number of times the predicted bits (i.e., the bits in the estimated image) are incorrect in multiple test events. For example, the image estimation unit 226 can estimate the image, and the video encoder 228 can encode the video data of the estimated image. Subsequently, the video decoder 224 can decode the error-corrected encoded video data of the image. The image estimation unit 226 can compare the bits in the transform coefficients of the encoded video data of the estimated image and the error-corrected video data of the image to determine whether the bits in the encoded video data of the estimated image are erroneous.
[0189] Reliability unit 1002 can convert probability values into LLR values. In some examples, reliability unit 1002 can convert probability values into LLR values using the following formula:
[0190] (8)
[0191] In the formula above, M indicates the LLR value, P error The LLR value indicates the probability of error, and ln represents the natural logarithm function. In some examples, the LLR value is a reliability value.
[0192] Figure 11 A graph showing example error probabilities and corresponding absolute values of the correspondence likelihood ratio (LLR) according to one or more techniques of this disclosure is provided. Figure 11 In the example, 150 bits are used to represent each transform block. Graph 1100 plots the error probability at each bit position within the transform block. As can be seen in Graph 1100, bits at specific positions have a higher error probability. Graph 1102 shows the error probability converted to the absolute value of the LLR. The absolute value of the LLR can be scaled.
[0193] In some examples, receiving device 104 uses a dynamic scaling process. During dynamic scaling, image estimation unit 226 dynamically determines prediction quality information based on video data. For example, some regions of the image are more difficult to predict than easier-to-predict regions (e.g., regions with reduced prediction accuracy). Examples of more difficult-to-predict image regions may include regions with high motion. For example, if the total value of motion vectors in a region crosses a threshold, image estimation unit 226 can determine that the region is a difficult-to-predict region. Therefore, image estimation unit 226 can identify such regions and generate a reliability value for the bits of the transform coefficients of the transform block based at least in part on whether the transform block is inside or outside such regions. In some examples, image estimation unit 226 determines the reliability value of the bits of the transform coefficients based on general statistics regarding the error in the position of bits modified based on whether the transform block containing the transform coefficients is in a difficult-to-predict region.
[0194] The reliability unit 1002 can use a reliability value to scale the bits of the transform coefficients in the encoded video data generated by the video encoder 228. For example, bits of the encoded video data generated by the video encoder 228 can be considered "hard" bits and can have values of exactly 0 or exactly 1. The reliability unit 1002 can use the prediction quality information generated by the image estimation unit 226 and the encoded video data generated by the video encoder 228 to determine "soft" bits between 0 and 1. For example, if the value of a bit in the encoded video data is 1, the reliability unit 1002 can generate a "soft" value for the bit by multiplying the absolute value of the bit's LLR by -1 (i.e., -1). If the value of a bit in the encoded video data is 0, the reliability unit 1002 can generate a "soft" value for the bit by multiplying the absolute value of the bit's LLR by +1 (i.e., +1). Therefore, the "soft" or scaled value of the bits of the transform coefficients can be M or -M.
[0195] Therefore, in some examples, each bit can be transformed into a scaling value, wherein the scaling value has a more positive value when there is a bit with a higher confidence level of 0 and a more negative value when there is a bit with a higher confidence level of 1. The reliability unit 1002 provides the scaling value as prior information to the channel decoder 222.
[0196] Channel decoder 222 uses scaling values to perform channel decoding processing. For example, reliability unit 1002 and channel encoder 212 can encode the encoded video data into codewords (e.g., low-density parity check codes). The bits of the codeword generated by reliability unit 1002 can be scaled as described above. The bits of the codeword generated by channel encoder 212 can be changed during transmission through channel 230, such that the bits of the codeword can be received as values between -1 and 1. Channel decoder 222 can apply the LPDC decoding process to the codeword to correct errors in the codeword. The bit values in the corrected codeword are 0 or 1. The LPDC decoding process can then convert the codeword from the encoded video data back to the original data. In other examples, other decoding schemes can be used. In this example, the error-correcting data received by channel decoder 222 may include cyclic redundancy check (CRC) data that was not used in the LPDC decoding process. In this way, channel decoder 222 can determine the value of each bit of the transform coefficients.
[0197] In some examples, channel encoder 212 may use predictive quality feedback to sort the bits of the transform coefficients before applying unequal protection channel decoding (such as polar code or spinal code). Unequal protection channel decoding involves allocating decoding redundancy based on the importance of the information bits. For example, channel encoder 212 may use reliability data to protect different bits based on the predictability of different bits by the receiving device 104.
[0198] In some examples, the reliability unit 1002 sends predicted quality feedback to the transmitting device 102. The video encoder 210 can adjust one or more encoding parameters applied to the video encoding process of the video data. For example, the transmitting device 102 can determine the compression rate of a limited video encoding process based on the predictability (reliability) of the receiving device 104. For example, if predictability is low, the transmitting device 102 can reduce the quality of the encoded video by reducing the number of transform coefficients or the bit width of each transform coefficient, thus transmitting less information.
[0199] In some examples, transmitting device 102 may adjust one or more channel decoding parameters used by channel encoder 212 based on prediction quality feedback. For example, channel encoder 212 may use unequal protection codes, where protection depends on prediction quality feedback. In some examples, channel encoder 212 may select different LDPC diagrams based on reliability. In some examples, channel encoder 212 may change the coding scheme used to generate error correction data to increase the error correction capability of bit positions or regions of a frame with lower reliability, or decrease the error correction capability of bit positions or regions of a frame with higher reliability.
[0200] In some examples, transmitting device 102 may update one or more bit puncturing parameters used by puncturing unit 214 based on predicted quality feedback. For example, puncturing unit 214 may change the puncturing pattern to allow more error correction data to be transmitted for bit positions and / or image regions with lower reliability. Therefore, transmitting device 102 may avoid puncturing bits with lower reliability. In some examples, puncturing unit 214 may change the puncturing pattern to allow less error correction data to be transmitted for bit positions and / or image regions with higher reliability.
[0201] In some examples, the predicted quality feedback sent by the reliability unit 1002 to the transmitting device 102 applies to the entire frame. In some examples, the reliability unit 1002 may send predicted quality feedback to the transmitting device 102 on a per-region basis. Each region may be a defined region within the frame. The reliability unit 1002 may send predicted quality feedback for some regions of the frame instead of others.
[0202] In some examples, reliability unit 1002 may send predicted quality feedback to transmitting device 102 on a periodic basis. For example, reliability unit 1002 may send predicted quality feedback to transmitting device 102 every N frames, where N is an integer value. In some examples, reliability unit 1002 sends predicted quality feedback to transmitting device 102 after a certain number of group of frames (GOPs) have been completed. In other examples, reliability unit 1002 may send predicted quality feedback on a non-periodic basis, such as in response to specific conditions or events.
[0203] In some examples, prediction quality feedback may include prediction quality data based on one or more noise models, such as a Gaussian noise model or a Laplace noise model. Noise model parameters can control one or more noise models. Reliability unit 1002 may send noise model parameters to transmitting device 102. The use of a noise model is an alternative for collecting error statistics. In this mode, the prediction error (the error statistics between the estimated frame estimated by frame estimation unit 226 and the actual frame reconstructed by video decoder 224) can be modeled using several parameters describing the error distribution function. Receiving device 104 can more easily transmit noise model parameters to transmitting device 102 because the noise model parameters may include less data compared to per-bit statistics. Transmitting device 102 may use the noise model in the same way as other types of prediction quality feedback.
[0204] In some examples, instead of the reliability unit 1002 receiving prediction quality information from the image estimation unit 226 of the receiving device 104, the video encoder 210 can generate prediction quality information and send it to the reliability unit 1002. The video encoder 210 can determine the prediction quality information based on prior information about the video encoding process. For example, the video encoder 210 can evaluate the reliability of each bit based on compression parameters (e.g., MSB is more reliable than LSB, low-frequency transform coefficients are more reliable than high-frequency transform coefficients, etc.). The video encoder 210 can perform the prediction process to evaluate statistics on its own, having predefined statistics for different sets of light compression parameters. Additionally, the transmitting device 102 can use other sensors to evaluate instantaneous motion and adjust the reliability accordingly.
[0205] Figure 12 This is a flowchart illustrating an example operation of a transmitting device 102 using scaled bits according to the technology of this disclosure. Figure 12In some examples, transmitting device 102 may obtain video data (1200) from video source 120, for example. Furthermore, transmitting device 102 may obtain prediction quality feedback (1202). In some examples, transmitting device 102 may obtain prediction quality feedback from receiving device 104. The prediction quality feedback includes bit reliability information. In some examples, the prediction quality feedback is represented based on noise model parameters (such as parameters of a Gaussian noise model or a Laplace noise model).
[0206] Transmitting device 102 can adjust one or more of video coding parameters, channel coding parameters, or bit puncturing parameters based on predictive quality feedback (1204). Video encoder 210 of transmitting device 102 can perform a video coding process to generate encoded video data (1206). The video coding process can be controlled by video coding parameters. For example, video coding parameters can control the number of transform coefficients included in a transform block, the number of bits included in the transform coefficients, etc. In some examples, video coding parameters include quantization parameters, and video encoder 210 can adjust the quantization parameters based on predictive quality feedback. For example, if the predictive quality feedback indicates low reliability, the quantization parameters can be reduced to lower the quantization level. As part of performing the video coding process, video encoder 210 can use the quantization parameters to quantize the transform coefficients of transform blocks of one or more frames.
[0207] The channel encoder 212 of the transmitting device 102 can perform a channel coding process on the scaling bits to generate channel-coded data (1208). The channel coding process can be controlled by channel coding parameters. For example, channel coding parameters may include controlling which LDPC graph to use to generate codewords in the channel coding process, and channel coding parameters may control error correction capabilities, etc. For example, if the channel coding parameters include an LDPC graph, the channel encoder 212 can adjust the LDPC graph and use the LDPC graph to generate codewords for transmission to the receiving device.
[0208] Furthermore, the puncturing unit 214 of the transmitting device 102 can perform a bit puncturing process (1210) on the error correction data generated by the channel encoder 212. The bit puncturing process can be controlled by bit puncturing parameters. For example, prediction quality feedback can indicate that certain parts of the encoded video data are less reliable. Therefore, the transmitting device 102 can adjust the bit puncturing parameters to reduce bit puncturing on the error correction data of the less reliable parts of the encoded video data. The transmitting device 102 can transmit the channel-coded data and the bit-punctured error correction data (1212) to the receiving device 104.
[0209] Figure 13 This is a flowchart illustrating an example operation of a receiving device 104 using scaled bits according to the technology of this disclosure. Figure 13In the example, receiving device 104 can obtain error correction data (1300) from the transmitting device. The error correction data provides error correction information about the frame of the video data.
[0210] The frame estimation unit 226 can generate prediction data for the frame (1302). The prediction data for the frame can include predictions of blocks of the frame based at least in part on one or more previously reconstructed frames of video data. For example, the frame estimation unit 226 can use inter-frame prediction and / or intra-frame prediction to generate block predictions.
[0211] Furthermore, the receiving device 104 can generate encoded video data (1304) based on the predicted data of the image. For example, the video encoder 228 of the receiving device 104 can perform a video encoding process to generate the encoded video data. The encoded video data includes transform blocks, wherein the transform blocks include transform coefficients.
[0212] The receiving device 104 can scale the bits (1306) of the transform coefficients of the transform block based on the reliability value of the bit position. In some examples, the receiving device 104 generates the reliability value. For example, the receiving device 104 can generate the reliability value based on statistics regarding the occurrence of errors in the bit position. In some examples, the receiving device 104 can generate the reliability value based on the reliability characteristics of individual regions of the video data frame. In some examples, the receiving device 104 can generate the reliability value based on a noise model. Furthermore, in some examples, the receiving device 104 can transmit the reliability value to the transmitting device 102. In other examples, the receiving device 104 can receive the reliability value from the transmitting device 102.
[0213] Additionally, the channel decoder 222 of the receiving device 104 can use error correction data to perform error correction operations on the scaling bits of the transform coefficients of the transform block to generate error-corrected coded video data (1308).
[0214] The video decoder 224 of the channel decoder 222 can reconstruct the picture based on the error-corrected coded video data (1310).
[0215] This disclosure describes techniques that can reduce the complexity of video encoding at a transmitting device, such as an extended reality (XR) headset. The transmitting device can acquire multi-view video data. The multi-view video data can include images from two or more viewpoints. For example, an XR headset can include two cameras for a stereoscopic view of a scene being viewed by a user. In this example, the multi-view video data can include images from each of the cameras.
[0216] Processing multiview video data consumes considerable processing resources. For example, in the context of augmented reality (AR) or mixed reality (MR), significant processing resources may be required to determine where virtual elements should be positioned and how they should appear. Multiview video data can aid in processing virtual elements. As an example, the same virtual element may need to be darker when positioned in a shadowed area of a scene, and brighter when positioned in a sunny area. As another example, the system may need to analyze the scene's content to determine if virtual elements will be occluded by physical elements in the scene, such as rocks or trees. Multiview video data can be used to determine the depth of objects in a scene. To keep XR headsets illuminated and maintain battery power at the XR headset, it may be desirable to minimize the processing of video data performed at the XR headset. Therefore, processing video data on another device, such as the user's smartphone or other nearby devices, can help reduce the processing resource requirements at the XR headset.
[0217] While multiview video data can be very useful in specific situations, simply sending unencoded multiview video data may be impractical because the amount of data required to send multiple parallel video streams simultaneously can be enormous. However, there is often considerable redundancy between the frames of different views in multiview video data. For example, what a person sees with their left eye is usually not that different from what their right eye sees. Therefore, video compression techniques have been developed to reduce this redundancy, thereby reducing the amount of data required to send multiview video data.
[0218] However, some techniques used for multi-view video decoding require considerable computational resources. For example, a video encoder might determine the differences between sets of two or more concurrent frames to determine a depth map of a scene. A depth map is an array of values indicating the depth / distance of objects shown in a frame from the camera. In this example, one of the concurrent frames might be an anchor frame, and one or more of the concurrent frames might be non-anchor frames. The video encoder can use the depth map to compute disparity vectors for blocks in the non-anchor frames. A block's disparity vector indicates the lateral displacement between the block and a corresponding block in another concurrent frame, such as an anchor frame. Typically, blocks representing deeper objects have disparity vectors with lower values than blocks representing closer objects. The video encoder can use the disparity vectors of the blocks to determine prediction blocks, generate residual data based on the prediction blocks, apply a transform to the residual data, quantize the transform coefficients of the resulting transformed blocks, and send the quantized transform coefficients via a signal. In this example, generating the depth map can involve considerable computational resources.
[0219] In another example, illumination levels can vary between concurrent frames from different viewpoints. These differences in illumination can degrade the decoding efficiency of multi-view video decoding. Therefore, illumination compensation can be applied to non-anchor frames to temporarily modify their illumination levels based on an illumination compensation factor, making the illumination levels of non-anchor frames more consistent with those of the anchor frames during video encoding. The original illumination levels of the non-anchor frames can be restored during video decoding. Determining the illumination compensation factor consumes computational resources.
[0220] The technology disclosed herein can offload some of the processing associated with multi-view video encoding from a transmitting device (e.g., an XR headset) to a receiving device (e.g., a mobile device). For example, the transmitting device can obtain a first set of multi-view frames of video data. The first set of multi-view frames includes a first frame and a second frame. The first frame is from a first viewpoint, and the second frame is from a second viewpoint. The transmitting device can transmit first encoded video data to the receiving device. The first encoded video data is based on the first set of multi-view frames. The transmitting device can receive multi-view encoding prompts from the receiving device. Furthermore, the transmitting device can obtain a second set of multi-view frames of video data. The second set of multi-view frames includes a third frame and a fourth frame. The third frame is from the first viewpoint, and the fourth frame is from the second viewpoint. The transmitting device can perform a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompts received from the receiving device to generate second encoded video data. The multi-view encoding process reduces inter-view redundancy between the third and fourth frames. The transmitting device can then transmit the second encoded video data to the receiving device.
[0221] Similarly, the receiving device can obtain first encoded video data from the transmitting device. The first encoded video data is based on a first set of multi-view frames of the video data. The first set of multi-view frames may include a first frame and a second frame. The first frame is from a first viewpoint, and the second frame is from a second viewpoint. The receiving device can determine multi-view encoding cues based on the first encoded video data. The receiving device can send the multi-view encoding cues to the transmitting device. Additionally, the receiving device can obtain second encoded video data from the transmitting device. The second encoded video data is based on a second set of multi-view frames including a third frame and a fourth frame. The second encoded video data is encoded using a multi-view encoding process based on multi-view encoding cues to reduce inter-view redundancy between the third and fourth frames.
[0222] Because the receiving device determines the multiview encoding hint and sends it to the transmitting device, the burden of determining the multiview encoding hint can be shifted from the transmitting device to the receiving device. This reduces the resource requirements at the transmitting device.
[0223] refer to Figure 2The video encoder 210 can perform a multi-view encoding process based on a multi-view encoding cue obtained from the receiving device 104. For example, the multi-view encoding cue may include a depth map. In this example, the video encoder 210 can use the depth map to estimate the disparity vector of a frame patch in a non-anchor view of the multi-view video data. The video encoder 210 can use the disparity vector of the current patch in the current frame to determine a prediction block for the current patch based on samples from a concurrent reference frame. The concurrent reference frame has the same Frame Order Count (POC) value as the current frame. The video encoder 210 can determine the residual data for the current patch based on the original samples and the prediction block of the current patch. The video encoder 210 can apply one or more transforms to the residual data to generate one or more transform blocks. The video encoder 210 can quantize the transform coefficients in the transform blocks. The encoded video data generated by the video encoder 210 can be based on the quantized transform coefficients.
[0224] In some examples, multi-view encoding cues may include one or more illumination compensation factors. When encoding the current frame of multi-view video data, the video encoder 210 may modify each sample of the current frame based on one or more illumination compensation factors. In some examples, different illumination compensation factors may be applied to different areas of the current frame. Modifying the samples of the current frame in this way can make the illumination level of the current frame more consistent with the illumination level of the concurrent reference frame. After modifying the samples of the current frame, the video encoder 210 may perform a multi-view encoding process, such as that described in the previous paragraphs, to encode blocks of the current frame.
[0225] Furthermore, according to some examples of this disclosure, video decoder 224 can perform multi-view decoding processes. For example, video decoder 224 can generate prediction blocks using the disparity vectors of blocks in the current frame. Video decoder 224 can reconstruct samples of blocks in the current frame using prediction blocks and residual data received from channel decoder 222. In some examples where video encoder 210 applies illumination compensation to the frame, video decoder 224 can use illumination parameters to invert the illumination compensation applied to the frame. In other examples, video decoder 224 can apply other multi-view decoding operations.
[0226] Furthermore, according to one or more techniques disclosed herein, the video decoder 224 can determine multi-view encoding cues based on encoded video data received from the transmitting device 102. For example, the video decoder 224 can determine a depth map, illumination compensation parameters, and other information that can be used in the multi-view encoding operation. The receiving device 104 can send the multi-view encoding cues back to the transmitting device 102, enabling the transmitting device 102 to use the multi-view encoding cues to perform a multi-view encoding process on subsequent frames.
[0227] The frame estimation unit 226 can generate an estimate of the next frame of the video data. In some examples, the frame estimation unit 226 can estimate the frame based on one or more previously reconstructed reference frames associated with different views. For example, the frame estimation unit 226 can use information from frames at previous times (such as disparity vectors or depth maps) to extrapolate the content of the frame from frames at the same time. In another example, the frame estimation unit 226 can extrapolate the frame based on one or more frames associated with the same view, regardless of frames associated with other views, in a manner substantially similar to that described elsewhere in this disclosure regarding the frame estimation unit 226 estimating frames of video data for a single view.
[0228] Video encoder 228 can perform the same operations as video encoder 210 on the estimated next frame. For example, video encoder 228 can perform intra-frame prediction to predict blocks, using the predicted blocks and corresponding predicted blocks of the frame generated by frame estimation unit 226 to generate residual data. In some examples, video encoder 228 can use multi-view coding cues to perform the same multi-view coding process as video encoder 210. Video encoder 228 can apply a transform (e.g., DCT transform) to the residual data to generate transform coefficients. Video encoder 228 can apply quantization to the transform coefficients.
[0229] Figure 14 This is a flowchart illustrating an example data exchange between a transmitting device 102 and a receiving device 104 related to multi-view processing according to one or more technologies of this disclosure. Figure 14 In one example, transmitting device 102 may obtain a first set of multi-view images (1400). Transmitting device 102 may transmit first encoded video data based on the first set of multi-view images to receiving device 104. In some examples, transmitting device 102 performs light compression on the first set of multi-view images to generate the first encoded video data. In other examples, the first encoded video data may include an unencoded version of the first set of multi-view images.
[0230] Receiving device 104 can perform multi-view processing (1402) on a first set of multi-view images. For example, receiving device 104 can decode the first set of multi-view images if necessary. Additionally, receiving device 104 can determine multi-view encoding prompts, for example, as described elsewhere in this disclosure. Receiving device 104 can send the multi-view encoding prompts to sending device 102.
[0231] In addition, Figure 14In the example, transmitting device 102 can obtain a second set of multiview frames (1404). Transmitting device 102 can perform a multiview encoding process on the second set of multiview frames to generate second encoded video data (1406). The second encoded video data may include encoded anchors and secondary (non-anchor) frames. Receiving device 104 can perform multiview decoding on the second encoded video data to reconstruct the second set of multiview frames (1408). Receiving device 104 can also perform multiview processing on the second set of multiview frames to determine updated multiview encoding hints (1410). Receiving device 104 can send the updated multiview encoding hints to transmitting device 102. Transmitting device 102 can use the updated multiview encoding hints for multiview encoding of subsequent sets of multiview frames.
[0232] Figure 15 This is a flowchart illustrating an example operation of a transmitting device 102 for multi-view processing according to the technology of this disclosure. Figure 5 In the example, transmitting device 102 can obtain a first set (1500) of multi-view frames of video data. The first set of multi-view frames includes a first frame and a second frame. The first frame comes from a first viewpoint, and the second frame comes from a second viewpoint. The communication interface 118 of transmitting device 102 ( Figure 1 The transmitting device 102 can send first encoded video data (1502) to the receiving device 104. The first encoded video data is based on a first set of multi-view frames. The transmitting device 102 can receive multi-view encoding prompts (1504) from the receiving device 104. In some examples, the multi-view encoding prompts include one or more of the following: relative shifts between blocks of the first frame (i.e., the frame of the first viewpoint) and the second frame (i.e., the frame of the second viewpoint), brightness correction between the first frame and the second frame, inter-block shifts between anchor blocks and reconstructed blocks, or motion data for reference shifts. The transmitting device 102 can receive the multi-view encoding prompts in one of a variety of ways. For example, the transmitting device 102 can receive the multi-view encoding prompts via an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, or another type of message.
[0233] Furthermore, the transmitting device 102 can obtain a second set of multi-view frames (1506) of video data. The second set of multi-view frames includes a third frame and a fourth frame. The third frame is from a first viewpoint, and the fourth frame is from a second viewpoint. The video encoder 210 can perform a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompt received from the receiving device to generate second encoded video data (1508). The multi-view encoding process reduces inter-view redundancy between the third and fourth frames. The transmitting device 102 can then send the second encoded video data to the receiving device 104 (1510).
[0234] For subsequent sets of multi-view screens, execution can be performed multiple times. Figure 15 The operation is as follows: For example, after sending second encoded video data to the receiving device, the sending device 102 can receive updated multi-view encoding prompts from the receiving device 104. The sending device 102 can obtain a third set of multi-view frames of the video data. The third set of multi-view frames may include a fifth frame and a sixth frame, the fifth frame being from a first viewpoint and the sixth frame being from a second viewpoint. The video encoder 210 of the sending device 102 can encode the third set of multi-view frames based on the updated multi-view encoding prompts received from the receiving device to generate third encoded video data. The sending device 102 can send the third encoded video data to the receiving device 104.
[0235] Figure 16 This is a flowchart illustrating an example operation of a receiving device 104 for multi-view processing according to the technology of this disclosure. Figure 16 In the example, receiving device 104 can obtain first encoded video data (1600) from transmitting device 102. For example, receiving device 104 can obtain this data via communication interface 134. Figure 1 The receiving device 104 obtains first encoded video data. The first encoded video data is based on a first set of multi-view frames of the video data. The first set of multi-view frames may include a first frame and a second frame. The first frame is from a first viewpoint, and the second frame is from a second viewpoint. The receiving device 104 may determine multi-view encoding cues (1602) based on the first encoded video data.
[0236] The receiving device 104 may send a multi-view encoding prompt (1604) to the transmitting device 102. The receiving device 104 may send the multi-view encoding prompt in one of a variety of ways. For example, the receiving device 104 may send the decimation mode indication via an uplink control information (UCI) / media access control-control element (MAC-CE) message, a radio resource control (RRC) message, or another type of message.
[0237] Additionally, receiving device 104 can obtain second encoded video data (1606) from transmitting device 102. The second encoded video data is based on a second set of multi-view images including a third and a fourth view. The second encoded video data is encoded using a multi-view encoding process that reduces inter-view redundancy between the third and fourth views based on multi-view encoding cues. Video decoder 224 of receiving device 104 can decode the second encoded video data.
[0238] In some examples, the multi-view encoding prompt includes a depth map indicating the depth of an object represented in a first and second view. The receiving device 104 may, as part of determining the multi-view encoding prompt, determine the depth map based on the first and second views. In some examples, the multi-view encoding prompt includes one or more illuminance compensation factors, and the receiving device 104 may, as part of determining the multi-view encoding prompt, determine the illuminance compensation factors based on the first and second views.
[0239] Figure 16 The process can be repeated multiple times. For example, receiving device 104 can determine a second multi-view encoding cue based on the second encoded video data. Receiving device 104 can send the second multi-view encoding cue to transmitting device 102. Subsequently, receiving device 104 can obtain third encoded video data from transmitting device 102. The third encoded video data is based on a third set of multi-view frames including a fifth frame and a sixth frame, and the third encoded video data is encoded using a multi-view encoding process that reduces inter-view redundancy between the fifth and sixth frames based on the second multi-view encoding cue.
[0240] According to one or more techniques disclosed herein, transmitting device 102 may receive a decimation mode indication from receiving device 104. The decimation mode indication may indicate a decimation mode. As described in more detail elsewhere in this disclosure, receiving device 104 may determine the decimation mode. The decimation mode may be a mode in which encoded video data is not transmitted.
[0241] Transmitting device 102 may receive the decimation mode indication in one of a variety of ways. For example, transmitting device 102 may receive the decimation mode indication via uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, sidelink control information (SCI) messages, or another type of message.
[0242] Transmitting device 102 can apply a decimation mode to the encoded video data generated by video encoder 210, thereby generating decimated video data. For example, the decimation mode can indicate a mode that skips the transmission of encoded video data for entire frames. Therefore, in this example, transmitting device 102 (e.g., the channel encoder 212 of transmitting device 102) can transmit encoded video data for some frames according to the indicated mode, without transmitting encoded video data for other frames. For example, transmitting device 102 can skip the transmission of encoded video data every other frame. In another example, transmitting device 102 can transmit encoded video data for one frame, and then not transmit the next two or more frames of encoded video data.
[0243] In another example, the extraction pattern can indicate a pattern for skipping the transmission of encoded video data in a specified region within a frame. A specific region of a series of frames may not change much (if at all) from frame to frame. For example, the background of a scene from a static viewpoint may not change significantly, while the change occurs in the more restricted region of interest. Because the region outside the region of interest changes little, such regions can be more easily predicted accurately. Therefore, according to the techniques of this disclosure, receiving device 104 can identify regions outside the region of interest. Therefore, transmitting device 102 can transmit encoded video data in the region of interest but not encoded video data in other regions.
[0244] In another example, the video data is multi-view video data, and the extraction mode can indicate a mode for skipping the transmission of encoded video data from a specified view. For example, two views may have very similar content, such as a view that primarily shows distant objects. Therefore, in this example, transmitting device 102 can transmit encoded video data from one of the views without transmitting encoded video data from one or more other views, as indicated by the extraction mode.
[0245] In another example, the decimation pattern can indicate the pattern of bits to be omitted from the syntax element indicating the transform coefficients. For example, the decimation pattern can indicate that a specific number of least significant bits will be omitted from the transform coefficients. In some examples, the decimation pattern can indicate that specific transform coefficients (e.g., high-frequency transform coefficients) will be omitted.
[0246] Figure 17 This is a block diagram illustrating example components of a transmitting and receiving device that performs extraction of coded video data according to the technology of this disclosure. Figure 17In the example, transmitting device 102 includes a video encoder 210, a channel encoder 212, a punching unit 214, and a transmitter extraction unit 1700. Receiving device 104 includes a de-punching unit 220, a channel decoder 222, a video decoder 224, a frame estimation unit 226, a video encoder 228, and a receiver extraction unit 1702. The video encoder 210, channel encoder 212, punching unit 214, de-punching unit 220, channel decoder 222, video decoder 224, frame estimation unit 226, and video encoder 228 can operate in the same manner as described elsewhere in this disclosure.
[0247] However, in Figure 17 In the example, the transmitter decimation unit 1700 can apply a decimation mode to the encoded video data after the channel encoder 212 generates error correction data for the encoded video data. The decimation mode indicates the mode of the encoded video data that is not transmitted. For example, the transmitter decimation unit 1700 can cause the transmitting device 102 not to transmit encoded video data of a specific frame, a region of the frame, a block pattern within the frame, a specific view, etc. The receiver decimation unit 1702 can determine the decimation mode indication based on the frame reconstructed by the video decoder 224. The receiver decimation unit 1702 can send a decimation mode indication to the transmitting device 102. The transmitter decimation unit 1700 can apply the decimation mode indicated by the decimation mode indication.
[0248] Although the DVC-based scheme has been described Figure 17 However, the techniques disclosed herein related to sending a decimation mode indication from the receiving device 104 to the transmitting device 102 are not necessarily limited thereto. For example, in some examples, the image estimation unit 226 and the video encoder 228 may be omitted.
[0249] Figure 18 This is a conceptual diagram illustrating an example exchange of information including extraction pattern indications according to the technology of this disclosure. Figure 18 In the example, the transmitting device 102 can transmit the first set of frames (e.g., frames n-n1, frames nn) n1+1 The encoded video data of the frame (n) is sent to the receiving device 104. The transmitting device 102 can also send error correction data of the first set of frames.
[0250] Receiving device 104 can transmit a decimation mode indication indicating a decimation mode determined based on a first set of encoded frames, and transmitting device 102 can receive a decimation mode indication indicating a decimation mode determined based on a first set of encoded frames. Figure 18 In the example, the extraction mode extracts frames at a 1:2 ratio. In other words, encoded video data from one frame out of every two frames will be sent.
[0251] Therefore, transmitting device 102 can send the encoded video data of the second set of frames to receiving device 104. According to the decimation mode indicated by the received decimation mode instruction, transmitting device 102 skips the transmission of encoded video data from every other frame in the second set of frames. For example... Figure 18 As shown in the example, the index values of the frames in the second set of frames (e.g., n+2, n+4, n+n2) are increased by 2 instead of 1, as is the case for the first set of frames.
[0252] Subsequently, receiving device 104 can determine, based on the second set of frames, that a more appropriate decimation mode would be a 1:1 decimation mode (i.e., the decimation mode in which transmitting device 102 transmits the encoded video data for each frame). Therefore, in Figure 18 In the example, receiving device 104 can send a second decimation mode indication indicating a second decimation mode, and transmitting device 102 can receive the second decimation mode indication indicating a second decimation mode. Subsequently, transmitting device 102 can transmit encoded video data of a third set of frames. According to the second decimation mode, transmitting device 102 does not skip the transmission of encoded video data for any frames in the third set of frames. Therefore, as... Figure 18 As shown in the example, the index value (e.g., n+n2+1, n+n2+2, etc.) is increased by 1 instead of 2.
[0253] Figure 19 This is a flowchart illustrating an example operation of a transmitting device 102 according to the technology of this disclosure, wherein the transmitting device 102 receives a decimation mode indication. Figure 19 In the example, the video encoder 210 of the transmitting device 102 can encode a first set of frames of video data to generate first encoded video data (1900). The transmitting device 102 can then send the first encoded video data (1902) to the receiving device 104.
[0254] Furthermore, transmitting device 102 can receive from receiving device 104 a decimation mode indication (1904) that indicates a decimation mode determined based on a first set of frames. The decimation mode can be a mode in which encoded video data is not transmitted. For example, in some examples, the decimation mode indicates a mode that skips the transmission of encoded video data for an entire frame. In other words, transmitting device 102 may not transmit any encoded video data for a particular frame and may transmit some or all of the encoded video data for other frames. In some examples, the decimation mode indicates a mode that skips the transmission of encoded video data for a specific region within a frame. For example, the decimation mode may indicate that transmitting device 102 will skip the transmission of encoded video data associated with a specific block of a frame, such as... Figure 6 and Figure 9As shown in the example. In some examples where the video data is multi-view video data, the decimation mode can indicate a mode for skipping the transmission of encoded video data from a particular view. In such examples, a view can be associated with a sensor on the same user device (e.g., the same XR headset) or with a sensor on a different user device (e.g., different XR headsets worn by different users) or camera. In some examples, different decimation modes can exist for different regions rather than frames. For example, decimation may not be applied to the region of interest, and a decimation mode of limiting blocks or lower effective bits or higher frequency transform coefficients can be applied to areas of the frame outside the region of interest.
[0255] The video encoder 210 can encode a second set of frames of video data to generate second encoded video data (1906). Additionally, the transmitter decimation unit 1700 can apply a decimation mode to the second encoded video data to generate decimated video data (1908). The transmitting device 102 can send the decimated video data to the receiving device (1910).
[0256] In some examples, the transmitter decimation unit 1700 can determine the decimation mode. Therefore, in Figure 19 In the context of [the above context], transmitting device 102 can encode a third set of video data frames to generate third encoded video data, determine a second decimation mode indicating that the encoded video data has not been transmitted, and apply the second decimation mode to the third encoded video data to generate second decimated video data. Transmitting device 102 can send the second decimated video data to receiving device 104. Transmitting device 102 can also send a second decimation mode indication to the receiving device. The second decimation mode indication is used to indicate that the second decimation mode is applied to the third encoded video data.
[0257] The transmitter decimation unit 1700 can determine the decimation mode in various ways. For example, the transmitter decimation unit 1700 can test various decimation modes. When testing a decimation mode, the transmitter decimation unit 1700 can apply the decimation mode to the frame and reconstruct the frame from one or more previous original frames of the frame's error correction data and video data. The transmitter decimation unit 1700 can compare the reconstructed frame with the frame to determine the distortion level. The transmitter decimation unit 1700 can compare the distortion levels associated with different decimation modes to determine the decimation mode.
[0258] In some examples, Figure 19The operations are performed within the DVC context. Therefore, the channel encoder 212 of transmitting device 102 can generate first error correction data based on the first coded video data. Transmitting device 102 can send the first error correction data to receiving device 104. The channel encoder 212 can generate second error correction data based on the second coded video data. Transmitting device 102 can send the second error correction data to receiving device.
[0259] Figure 20 This is a flowchart illustrating an example operation of a receiving device 104 according to the technology of this disclosure, wherein the receiving device 104 transmits a decimation mode indication. Figure 20 In the example, receiving device 104 can receive first encoded video data from transmitting device 102 (2000). Video decoder 224 can perform a decoding process to reconstruct a first set of images based on the first error-corrected encoded video data (2002).
[0260] Additionally, the receiver extraction unit 1702 can determine an extraction mode (2004) based on a first set of frames that indicates a mode in which encoded video data is not transmitted. In some examples, the extraction mode indicates a mode that skips the transmission of encoded video data for the entire frame. In some examples, the extraction mode indicates a mode that skips the transmission of encoded video data for a specified area or block within the frame, such as in... Figure 6 and Figure 9 Examples are provided. In some examples, the video data is multi-view video data and the extraction mode indicates a mode that skips the transmission of encoded video data from the specified view's frame.
[0261] The receiver decimation unit 1702 can determine the decimation mode in one of a variety of ways. For example, the receiver decimation unit 1702 can apply one or more experimental decimation modes to error-corrected coded video data generated by the channel decoder 222 for a first set of frames to generate decimated coded video data. The receiver decimation unit 1702 can then cause the channel decoder 222 to apply an error correction process to modify the decimated coded video data based on the first error-corrected data to generate experimental error-corrected video data. The receiver decimation unit 1702 can then cause the video decoder 224 to apply a decoding process to reconstruct the first set of frames based on the experimental error-corrected video data. The receiver decimation unit 1702 can determine whether the decimation mode meets one or more criteria based on a comparison between the first set of frames reconstructed as based on the experimental error-corrected video data and the first set of frames reconstructed as based on the first error-corrected video data. For ease of explanation, the frame reconstructed based on the experimental error-corrected video data may be referred to as a "test frame," and the frame reconstructed based on the first error-corrected video data may be referred to as a "baseline frame." The receiver extraction unit 1702 can repeat the process with multiple trial extraction patterns until the receiver extraction unit 1702 identifies an extraction pattern that meets the criteria.
[0262] For example, receiver extraction unit 1702 can compare each test frame with a corresponding baseline frame to determine whether the test frame meets the standard. For example, if the sum of the differences between the test frame and the corresponding baseline frame is less than a certain amount, receiver extraction unit 1702 can determine that the test frame meets the standard. If at least a given number of test frames exceed a threshold, receiver extraction unit 1702 can select an extraction mode associated with the test frame.
[0263] In a more general example, receiver extraction unit 1702 can apply a function to the test screen and the corresponding baseline screen to generate a value. If the value is less than a threshold, receiver extraction unit 1702 can select an extraction mode associated with the test screen.
[0264] Furthermore, in some examples, receiver decimation unit 1702 can cancel the use of a decimation mode (e.g., revert to a mode in which all encoded video data is transmitted) or switch to a less aggressive decimation mode when certain conditions occur. For example, receiver decimation unit 1702 can cancel the use of a decimation mode in response to determining that a given number of baseline frames fail to meet a standard. For example, if the sum of the differences between the test frame and the corresponding baseline frame is greater than a certain amount, receiver decimation unit 1702 can determine that the test frame fails the standard. If at least a given number of test frames that fail the standard exceed a threshold, receiver decimation unit 1702 can cancel the use of the decimation mode or revert to a less aggressive decimation mode. In a more general example, receiver decimation unit 1702 can apply a function to the test frame and the corresponding baseline frame to generate a value. If this value is greater than a second threshold, receiver decimation unit 1702 can cancel the use of the decimation mode or revert to a less aggressive decimation mode.
[0265] The receiver decimation unit 1702 can send a decimation mode indication (2006) to the transmitting device 102, indicating the decimation mode determined. The receiver decimation unit 1702 can send the decimation mode indication using uplink control information (UCI) / media access control-control element (MAC-CE) messages, radio resource control (RRC) messages, sidelink control information (SCI) messages, or another type of message.
[0266] Receiving device 104 can receive extracted video data (2008) from transmitting device 102. The extracted video data may include second encoded video data for which an extraction mode has been applied. The second encoded video data is generated based on a second set of frames of the video data.
[0267] Video decoder 224 can perform a decoding process to reconstruct a second set of images based on second error-corrected coded video data (2010). Video decoder 224 can perform the same decoding process as described elsewhere in this disclosure.
[0268] In some examples, receiving device 104 can receive and use a decimation mode indication from transmitting device 102. Therefore, in Figure 20In the example, receiving device 104 can receive a second decimation mode indication indicating that a second mode of encoded video data has not been transmitted. Receiving device 104 can receive third error-corrected data and second decimated video data from transmitting device 102. The second decimated video data may include third encoded video data for which the second decimation mode has been applied. The third encoded video data can be generated based on a third set of frames of the video data. Channel decoder 222 can apply an error correction process to generate third error-corrected encoded video data based on the third encoded video data and the third error-corrected data. Video decoder 224 can apply a decoding process to reconstruct a third set of frames based on the third error-corrected encoded video data.
[0269] In some examples, Figure 20 The process can be executed in a DVC-based implementation. Therefore, receiving device 104 can receive first error-correcting data from transmitting device 102. Receiving device 104 applies an error-correcting process to modify first coded video data based on the first error-correcting data to generate first error-corrected coded video data. Receiving device 104 can perform a decoding process to reconstruct a first set of images based on the first error-corrected coded video data. Additionally, receiving device 104 can receive second error-correcting data from transmitting device 102. Receiving device 104 can apply an error-correcting process to generate second error-corrected coded video data based on the second coded video data, predictive coded video data generated by video encoder 228, and the second error-correcting data. Receiving device 104 can perform a decoding process to reconstruct a second set of images based on the second error-corrected coded video data.
[0270] During the video encoding process, a video encoder typically analyzes multiple encoding options and selects the optimal one. For example, a video encoder might analyze various ways to divide a maximum decoding unit (LCU) or macroblock into decoding units (CUs) and / or prediction units (PUs). In another example, a video encoder might analyze multiple intra-frame prediction modes to generate prediction blocks for a PU while performing intra-frame prediction. In yet another example, a video encoder might analyze multiple reference frames and motion vectors to generate prediction blocks for a PU while performing inter-frame prediction. This analysis and selection can be resource-intensive. For example, to be efficient, a video encoder might need to process multiple options in parallel, which increases the hardware complexity and power requirements of the video encoder. Analysis and selection can also involve multiple requests to read data and write data to memory, further increasing power requirements.
[0271] According to one or more techniques of this disclosure, a large portion of the process of analyzing and selecting encoding operations is transferred from a transmitting device (e.g., transmitting device 102) to a receiving device (e.g., receiving device 104). For example, the transmitting device may encode a first frame of video data to generate first encoded video data. The transmitting device may send the first encoded video data to the receiving device. The receiving device may receive the first encoded video data from the transmitting device and reconstruct the first frame based on the first encoded video data. Additionally, the receiving device may estimate a second frame of the video data based on the first frame. The second frame may be a frame that appears after the first frame in decoding order. The receiving device may generate encoding selection data for the estimated second frame. The encoding selection data indicates encoding selections for encoding the estimated second frame. The receiving device may send the encoding selection data for the second frame. The transmitting device may receive the encoding selection data for the second frame of the video data. The transmitting device may encode the second frame based on the encoding selection data to generate second encoded video data. The transmitting device may send the second encoded video data to the receiving device. The receiving device may receive the second encoded video data from the transmitting device. The receiving device may reconstruct the second frame based on the second encoded video data. In this way, because the transmitting device receives encoded selection data from the receiving device, the transmitting device does not need to perform resource-intensive analysis and selection processes while encoding the second screen, since the analysis and selection process for the second screen has already been performed at the receiving device. This reduces the resource requirements of the transmitting device.
[0272] Transmitting and receiving devices can communicate over low-range, low-power links on ultra-wideband (e.g., high-bandwidth) communication links. In some examples, the transmitting and receiving devices may communicate using time-division duplex (TDD), subband non-overlapping full-duplex (SBFD), or SFFD schemes. The low latency associated with this type of communication allows the transmitting device to receive coded selection data quickly enough to continue transmitting coded video data to meet a predetermined frame rate.
[0273] Figure 21 This is a block diagram illustrating example components of a transmitting device 102 according to the technology of this disclosure and a receiving device 104 that transmits encoded selection data to the transmitting device. Figure 21 In the example, the transmitting device 102 may include a video encoder 210, a channel encoder 212, and a punching unit 214. The receiving device 104 may include a de-punching unit 220, a channel decoder 222, a video decoder 224, a frame estimation unit 226, and a video encoder 228.
[0274] exist Figure 21In the examples provided, video encoder 210 can encode frames of video data. Unlike some of the examples given above, video encoder 210 can perform a complete video encoding process that may include intra-frame prediction and inter-frame prediction. In some examples, video encoder 210 can encode video data using video codecs such as H.264 / AVC, H.265 / HEVC, H.266 / VVC, Basic Video Decoding (EVC), AV1, etc. Channel encoder 212, puncturing unit 214, depuncturing unit 220, and channel decoder 222 can operate in the same manner as described elsewhere in this disclosure.
[0275] In addition, Figure 21 In the example, the video decoder 224 of the receiving device 104 can perform a video decoding process on the error-corrected coded video data generated by the channel decoder 222. The video decoder 224 can perform a complete video decoding process including intra-frame prediction and inter-frame prediction. The video decoder 224 can use the same video codec as the video encoder 210.
[0276] After the video decoder 224 reconstructs at least a portion of the video data frame, the frame estimation unit 226 can estimate the corresponding portion of the subsequent frame following the reconstructed frame. The frame estimation unit 226 can estimate the subsequent frame in the same manner as described elsewhere in this disclosure. Furthermore, the video encoder 228 can apply the video encoding process to the subsequent frame. In the example where both the video encoder 210 and the video decoder 224 use a video codec, the video encoder 228 can use the same codec.
[0277] However, according to one or more techniques of this disclosure, receiving device 104 may send encoding selection data 2100 to transmitting device 102. Encoding selection data 2100 indicates encoding selection for encoding the estimated subsequent frame. For example, encoding selection data may include motion parameters of blocks in the estimated subsequent frame. The block motion parameters may include motion vectors, reference frame indicators, merge candidate indices, affine motion parameters, and other data for determining predicted blocks in one or more reference frames. Therefore, in this example, video encoder 228 may encode the block using inter-frame prediction, and encoding selection data 2100 may include motion parameters instructing video encoder 228 how to encode the block using inter-frame prediction.
[0278] In some examples, the encoding selection data may include intra-prediction parameters for the blocks in the estimated subsequent frames. The intra-prediction parameters may include data indicating the intra-prediction mode (e.g., planar mode, DC mode, directional prediction mode, etc.) used by the video encoder 228 for intra-prediction of the blocks. In some examples, the encoding selection data may include other information, such as information describing how the video encoder 228 segments the estimated subsequent frames into blocks, whether residual prediction is used, whether and how intra-block copying (IBC) is used, and whether specific filters are used.
[0279] Video encoder 210 can use encoding selection data 2100 when encoding actual (non-estimated) subsequent frames. That is, instead of searching for different possibilities during the video encoding process, video encoder 210 can use the video encoding selection indicated by encoding selection data 2100. For example, encoding selection data 2100 can indicate that a specific block of a subsequent frame is encoded using a specific intra-prediction mode. Therefore, in this example, when encoding a subsequent frame, video encoder 210 can encode the specific block using the specific intra-prediction mode without analyzing different potential intra-prediction modes to select the specific intra-prediction mode. In another example, encoding selection data 2100 can indicate the motion vectors and reference frame for a specific block of a subsequent frame. Therefore, in this example, when encoding a subsequent frame, video encoder 210 can use the motion vectors to determine the predicted block in the reference frame without analyzing the potential reference frame and motion vectors. Video encoder 210 can then encode the specific block using the predicted block.
[0280] The transmitting device 102 can process the encoded video data of subsequent frames in the same way as other frames. Similarly, the receiving device 104 can process the encoded video data of subsequent frames in the same way as other encoded video data. Therefore, after the video decoder 224 reconstructs at least a portion of the subsequent frame, the frame estimation unit 226 can predict the corresponding portion of the frame following the subsequent frame, the video encoder 228 can encode the video data of the estimated subsequent frame, and transmit encoding selection data for the estimated subsequent frame; this cycle can be repeated. In this way, some of the burden of encoding video data can be transferred from the video encoder 210 of the transmitting device 102 to the video encoder 228 of the receiving device 104. This can reduce the resource requirements of the transmitting device 102.
[0281] about Figure 21The described process can be adapted for use with DVC technology. For example, transmitting device 102 can apply a decimation mode to the encoded video data (e.g., according to any of the examples provided elsewhere in this disclosure), such that transmitting device 102 transmits only some encoded video data, but still transmits error-corrected data of the decimated video data. The channel decoder 222 of receiving device 104 can receive error-corrected data for a specific frame from the transmitting device (e.g., via the de-punch unit 220). In this example, channel decoder 222 can apply an error-correction process to generate error-corrected encoded video data based on the encoded video data for the specific frame generated by the video encoder 228 of receiving device 104 and the error-corrected data for that specific frame. Video decoder 224 can decode the error-corrected encoded video data to reconstruct specific features.
[0282] In some cases, transmitting device 102 needs to transmit encoded video data according to a schedule. For example, transmitting device 102 may need to send encoded video data to receiving device 104 according to a predetermined frame rate per minute to support a specific application. Therefore, it is possible that transmitting device 102 may not receive the encoding selection data for the frame in time for encoding and transmission of the encoded video data. Therefore, in some examples, based on the determination that no encoding selection data for the frame has been received from receiving device 104 before the time limit expires, video encoder 210 may encode the frame without using the encoding selection data. Video encoder 210 may encode the frame using a finite video encoding process. Furthermore, in some examples, the encoded video data of the frame may include encoding selection data generated by video encoder 210. In some examples, the encoded video data of the frame may include data indicating that the encoded video data of the frame is not generated based on the encoding selection data generated by receiving device 104. The time limit may be subject to or defined based on the capabilities of transmitting device 102.
[0283] Encoding selection data and timing constraints can be defined on a per-image-fragment basis (e.g., slices, regions, etc.). Therefore, in this disclosure, discussions of encoding selection data, encoded video data, or other types of data for a frame can be applied only with respect to individual frames.
[0284] Figure 22 This is a communication diagram illustrating an example exchange of data, including encoded selection data, between a transmitting device 102 and a receiving device 104 according to the technology of this disclosure. Figure 22In the example, transmitting device 102 sends the encoded video data of frame n-1 to receiving device 104. Receiving device 104 can reconstruct frame n-1 based on the encoded video data of frame n-1. Additionally, receiving device 104 can estimate and encode frame n based on frame n-1. Receiving device 104 can send encoding selection data for frame n to transmitting device 102. Transmitting device 102 can encode frame n based on the encoding selection data for frame n and send the resulting encoded video data of frame n to receiving device 104. This process can be repeated multiple times. Therefore, in Figure 22 In the example, receiving device 104 may reconstruct frame n based on the encoded video data of frame n, estimate frame n+1 based on frame n and / or one or more other previously reconstructed frames, encode frame n+1, and send encoding selection data for frame n+1 to transmitting device 102.
[0285] Figure 23 This is a flowchart illustrating an example operation of a transmitting device 102 according to the technology of this disclosure, wherein the transmitting device 102 receives encoding selection data. Figure 23 In the example, video encoder 210 encodes a first frame of video data to generate first encoded video data (2300). In some examples, if transmitting device 102 has not yet received encoding selection data for the first frame, transmitting device 102 may perform a finite video encoding process on the first frame. Compared to a full video encoding process, a finite video encoding process can use relatively fewer computationally intensive decoding tools. For example, a finite video encoding process can use intra-frame prediction instead of inter-frame prediction.
[0286] Transmitting device 102 may send first coded video data (2302) to receiving device 104. In some examples, transmitting device 102 may apply a channelization coding process to the first coded video data to generate error correction data for the first coded video data. Transmitting device 102 may send both the first coded video data and the error correction data to receiving device 104.
[0287] Subsequently, transmitting device 102 can receive encoding selection data (2304) for a second frame of video data from receiving device 104. The encoding selection data can indicate encoding selections used to encode an estimate of the second frame. The second frame is decoded after the first frame. In some examples, the second frame may appear before or after the first frame in the output decoder. In some examples, the encoding selection data is entropy encoded. Therefore, in such examples, transmitting device 102 can entropy decode the encoding selection data. For example, transmitting device 102 can apply CABAC decoding, Golomb-Rice decoding, or another type of entropy decoding to the encoding selection data. In some examples, the encoding selection data is channel-coded. Therefore, transmitting device 102 can apply error correction operations to the encoding selection data based on error correction data.
[0288] The video encoder 210 of transmitting device 102 can encode the second frame based on encoding selection data to generate second encoded video data (2306). For example, the encoding selection data may include data indicating how a particular macroblock is segmented into CUs. In this example, the video encoder 210 may segment the macroblock into CUs in a manner indicated by the encoding selection data. In another example, the encoding selection data may indicate an intra-prediction mode for the block (e.g., CU or PU), and the video encoder 210 may use the indicated intra-prediction mode to encode the block. Thus, in this example, the encoding selection data received from receiving device 104 may include intra-prediction parameters for the block of the second frame, and the video encoder 210 may, as part of encoding the second frame, perform intra-prediction based on the intra-prediction parameters for the block of the second frame to generate a prediction block. The second encoded video data may include encoded video data based on the prediction block.
[0289] In another example, the encoded selection data received from receiving device 104 includes motion parameters of blocks in the second frame, and transmitting device 102 can, as part of encoding the second frame, perform motion compensation based on the motion parameters of the blocks in the second frame to generate predicted blocks. The second encoded video data includes encoded video data based on the predicted blocks.
[0290] The transmitting device 102 may send second encoded video data (2308) to the receiving device. In some examples, the second encoded video data does not include encoding selection data that instructs the transmitting device 102 to encode the second frame or the receiving device 104 to encode an estimate of the second frame. The second encoded video data may not necessarily include encoding selection data because the receiving device 104 generates and therefore already has encoding selection data.
[0291] In some examples, Figure 23 The operation can be used in conjunction with DVC technology. For example, transmitting device 102 can receive encoding selection data for a third frame of video data. The encoding selection data for the third frame can indicate encoding selections for encoding an estimate of the third frame. Video encoder 210 can encode the third frame based on the encoding selection data for the third frame to generate third-coded video data. Channel encoder 212 can apply a channel coding process to generate error correction data for the third-coded video data. Transmitting device 102 can send the error correction data of the third-coded video data to the receiving device without sending at least a portion of the third-coded video data.
[0292] Figure 24 This is a flowchart illustrating an example operation of a receiving device 104 according to the technology of this disclosure, wherein the receiving device 104 transmits encoding selection data. Figure 24 In the example, receiving device 104 can receive first encoded video data (2400) from sending device.
[0293] The video decoder 224 of the receiving device 104 can reconstruct a first frame of the video data based on the first encoded video data (2402). The frame estimation unit 226 of the receiving device 104 can estimate a second frame of the video data based on the first frame (2404). The second frame can be a frame that appears after the first frame in the decoding order.
[0294] The video encoder 228 of the receiving device 104 can generate encoding selection data (2406) for the estimated second frame. The encoding selection data indicates encoding selections for encoding the estimated second frame. For example, as part of encoding the estimated second frame, the video encoder 228 can perform motion compensation based on motion parameters of blocks in the second frame to generate prediction blocks. In this example, the encoding selection data may include motion parameters of blocks in the second frame from one or more processors. In some examples, as part of encoding the second frame, the video encoder 228 can perform intra-frame prediction based on intra-frame prediction parameters for blocks in the second frame to generate prediction blocks. In this example, the encoding selection data may include intra-frame prediction parameters for blocks in the second frame.
[0295] Receiving device 104 can send encoding selection data (2408) for a second frame to transmitting device 102. In some examples, receiving device 104 can apply entropy coding (e.g., CABAC coding, Golomb-Rice decoding, etc.) to the encoding selection data for the second frame before sending it. In some examples, receiving device 104 can perform a channel coding process on the encoding selection data to generate error correction data for the encoding selection data. Receiving device 104 can send the encoding selection data and the error correction data for the encoding selection data to transmitting device 102. In some examples, compared to other data transmissions in the data link (e.g., wireless sidelink channel 112, wireless uplink / downlink channel, etc.) between receiving device 104 and transmitting device 102, the communication interface 134 of receiving device 104 ( Figure 1 The coded selection data can be modulated at a lower modulation order. This increases the likelihood that the transmitting device 102 will correctly receive the coded selection data.
[0296] Subsequently, receiving device 104 can receive second encoded video data (2410) from transmitting device. Video decoder 224 can reconstruct a second scene (2412) based on the second encoded video data. In some examples, the second encoded video data does not include encoding selection data. Video decoder 224 can apply a decoding process that includes using encoding selection data to reconstruct the second scene based on the second encoded video data.
[0297] Figure 24 The process can be used in conjunction with DVC technology. For example, the frame estimation unit 226 can estimate a third frame of video data based on one or more of the first or second frames. The video encoder 228 can encode the estimated third frame to generate third encoded video data. The receiving device 104 can send third encoding selection data to the transmitting device 102. The third encoding selection data can indicate encoding selections for encoding the estimated third frame. Subsequently, the receiving device 104 can receive error correction data for the third frame. The channel decoder 222 can apply an error correction process to generate error-corrected encoded video data for the third frame based on the error correction data and the third encoded video data. The video decoder 224 can apply the error-corrected encoded video data based on the third frame to reconstruct the decoding process of the third frame. In some examples where the transmitting device 102 and the receiving device 104 use DVC technology, the error-corrected video data for the third frame does not include the third encoding selection data. However, the video decoder 224 can use the third encoding selection data generated by the video encoder 228 of the receiving device 104 to apply a decoding process to reconstruct the third frame based on the error-corrected encoded video data.
[0298] Figure 25 This is a conceptual diagram illustrating an example hierarchy of encoded video data according to the technology of this disclosure. More specifically, Figure 25 This shows the hierarchy of encoded video data generated using the H.264 / AVC video decoding standard. For example... Figure 25 As shown in the example, the Network Abstraction Layer (NAL) is the highest level of the layer. At the NAL, data is organized into NAL units. In some examples, NAL units are assigned to different packets or decoding blocks for transmission. NAL units in the NAL can include Sequence Parameter Sets (SPS) and Picture Parameter Sets (PPS) containing high-level syntax. NAL units in the NAL can also include Video Decoding Layer (VCL) NAL units. VCL NAL units can include slice NAL units containing slice-level data. A slice can be a series of macroblocks within a frame. Slices can include Instantaneous Decoder Refresh (IDR) slices and regular slices. Decoding an IDR slice does not depend on any other slice. Regular slices may have dependencies on other slices.
[0299] Each slice NAL unit can include a slice header and slice data. The slice header of the slice NAL unit includes information for decoding the slice data of the slice NAL unit. The slice data of the slice NAL unit consists of a series of macroblocks (MBs). Skip indicators can be scattered between the MBs. Each MB contains encoded video data of a specific block of the slice. Furthermore, as... Figure 25 As shown, a Block Markup (MB) can include a type indicator, prediction information, decoded block mode, quantization parameters (QP), and coded residual data. If the MB is coded using intra-frame prediction, the prediction data can indicate one or more intra-frame modes used to encode the MB. If the MB is coded using inter-frame prediction, the prediction data can indicate one or more reference frames and one or more motion vectors. The coded residual data of the MB can include the coded residual data of the luma blocks within the MB, the coded residual data of the Cb blocks within the MB, and the coded residual data of the Cr blocks within the MB. Generally, the coded residual data constitutes the maximum capacity portion of the encoded video data.
[0300] According to the technology disclosed herein, everything in the hierarchy at the macroblock layer except for the coded residual data can be coded selection data. Therefore, in some examples, receiving device 104 can send type data, prediction data, decoded block mode, and QP to transmitting device 102 for each MB of the estimated frame. Furthermore, in some examples, transmitting device 102 can send only the coded residual data of the MB to receiving device 104, without sending the MB's type data, prediction data, decoded block mode, or QP. In some examples, the coded selection data sent by receiving device 104 may include slice header data, SPS data, and PPS data. Transmitting device 102 and receiving device 104 can exchange information, or information indicating encoder and decoder capabilities can be pre-configured.
[0301] In some examples, transmitting device 102 may transmit data other than encoded residual data of some frames, some MBs, or some slices. Transmitting device 102 may transmit information (e.g., bits) at the network abstraction layer (e.g., as a frame-level control field), which may indicate whether a DVC-based approach should be used or whether conventional compression should be used for a particular frame.
[0302] Figure 26 This is a block diagram illustrating alternative example components of a transmitting device 102 according to one or more technologies of this disclosure. Figure 26 In the example, transmitting device 102 performs digital encoding and analog encoding on the video data. Transmitting device 102 transmits digitally encoded video data and analog-encoded video data to receiving device 104 via channel 230.
[0303] exist Figure 26 The example includes a video encoder 2600, a residual generation unit 2602, an analog encoder 2604, a reliability sorting unit 2606, an interleaving unit 2608, a channel encoder 2610, and a punching unit 2612. The video encoder 2600 can acquire video data and... Figure 2 The video encoder 210 operates in roughly the same way. For example... Figure 26 As shown in example A, video encoder 2600 can receive values of encoding parameters (e.g., encoding selection parameters) sent by receiving device 104. In some examples, video encoder 2600 can send values of encoding parameters and / or encoding selection data to receiving device 104. In this way, the video encoder 2600 of transmitting device 102, the video encoder of receiving device 104, and the video decoder of receiving device 104 can operate based on the same values of encoding parameters.
[0304] The video encoder 2600 can also output predicted data to the residual generation unit 2602. Furthermore, the video encoder 2600 can apply a higher level of quantization than the video encoder 210. The residual generation unit 2602 can generate residual data based on the predicted data and the video data.
[0305] The analog encoder 2604 can perform analog encoding operations on residual data. Examples of details of the analog encoding operations can be found in the following patents: U.S. Patent No. 11,553,184, filed December 29, 2020, entitled "Hybrid Digital-Analog Modulation for Transmission of Video Data"; U.S. Patent No. 11,431,962, filed December 29, 2020, entitled "Analog Modulated Video Transmission with Variable Symbol Rate"; and U.S. Patent No. 11,457,224, filed December 29, 2020, entitled "Interlaced Coefficients in Hybrid Digital-Analog Modulation for Transmission of Video Data".
[0306] For example, in some examples, the analog encoder 2604 can generate coefficients based on residual data. For example, the analog encoder 2604 can binarynize the residual data to generate coefficients. The analog encoder 2604 can quantize the coefficients. In other examples of generating coefficients based on video data, the analog encoder 2604 can perform more, fewer, or different steps. For example, in some examples, the analog encoder 2604 does not perform a quantization step. In other examples, the analog encoder 2604 does not perform a step of binarynizing the residual data.
[0307] Furthermore, the analog encoder 2604 can generate a coefficient vector. Each coefficient vector contains n coefficients. The analog encoder 2604 can generate the coefficient vector in one of several ways. For example, in one example, the analog encoder 2604 can generate the coefficient vector as a set of n consecutive coefficients based on the coefficient decoding order. Various coefficient decoding orders can be used, such as raster scan order, zigzag scan order, reverse raster scan order, vertical scan order, etc. In some examples, the coefficient vector may include one or more negative coefficients and one or more positive coefficients (i.e., signed coefficients). In some examples, the coefficient vector only includes non-negative coefficients (i.e., unsigned coefficients).
[0308] For each of the coefficient vectors, the analog encoder 2604 can determine the amplitude value of the coefficient vector based on a mapping pattern. For each corresponding allowed coefficient vector among a plurality of allowed coefficient vectors, the mapping pattern maps the corresponding allowed coefficient vector to a corresponding amplitude value among a plurality of amplitude values. The corresponding amplitude value is adjacent to at least one other amplitude value among a plurality of amplitude values in n-dimensional space, wherein the at least one other amplitude value is adjacent to the corresponding amplitude value in a monotonic digital line of amplitude values.
[0309] In some examples, to determine the amplitude value of the coefficient vector, the analog encoder 2604 can determine the position in n-dimensional space. The coordinates of the position in n-dimensional space are based on the coefficients of the coefficient vector, and the mapping pattern maps different positions in n-dimensional space to different amplitude values among multiple amplitude values. The analog encoder 2604 can determine the amplitude value of the coefficient vector as the amplitude value corresponding to a specific position in n-dimensional space.
[0310] Analog encoder 2604 can modulate analog signals based on the amplitude values of a coefficient vector. For example, analog encoder 2604 can determine an analog symbol based on a pair of amplitude values. An analog symbol can correspond to the phase shift and power of a point in the IQ plane with coordinates indicated by the amplitude value pair. Analog encoder 2604 can modulate the analog signal during the symbol sampling time based on the determined phase shift and power. The modem (e.g., communication interface 118) of transmitting device 102 can be configured to output analog signals.
[0311] In addition, Figure 26 In example A, the reliability sequencing unit 2606 can obtain encoded video data generated by the video encoder 2600. The reliability sequencing unit 2606 can obtain reliability-side information from the video encoder 2600. In some examples, the reliability sequencing unit 2606 can receive values of channel and compression state feedback (CCSF) parameters. The values of the CCSF parameters can provide information about the condition of channel 230 (e.g., signal-to-noise ratio, latency, network bandwidth congestion, etc.). In some examples, the values of the CCSF parameters provide information related to predicted reliability and quality. For example, the values of the CCSF parameters may include a decimation mode indicator. In some examples, the values of the CCSF parameters can enable the channel encoder 2610 to determine the decimation mode.
[0312] Interleaving unit 2608 can perform an interleaving process that ensures reliable and unreliable bits are evenly distributed across code blocks. For example, encoded video data can be divided into code blocks. Channel encoder 2610 can generate a separate error correction dataset for each code block. Before channel encoder 2610 generates error correction data, interleaving unit 2608 can interleave encoded video data between code blocks according to a predefined interleaving pattern. For example, encoded video data representing different neighboring pixels can be interleaved into different code blocks. After the channel decoding process is applied, the deinterleaving process performed at receiving device 104 is the reverse of the interleaving process. Therefore, if one of the code blocks is corrupted during transmission, pixels decoded from the corrupted code block can be spatially scattered across the frame among pixels decoded from the uncorrupted code block.
[0313] The channel encoder 2610 of the transmitting device 102 can perform a channel coding process on the video data obtained from the interleaving unit 2608. The channel encoder 2610 can, according to the information provided by the channel encoder 212 (…), Figure 2 The channel coding process can be performed using any example provided by the channel encoder 2610. The puncturing unit 2612 can perform bit puncturing operations on the error-correcting data generated by the channel encoder 2610. The puncturing unit 2612 can be configured according to the information provided by the puncturing unit 214 (…). Figure 2 Any example provided performs bit punching on the error correction data. Transmitting device 102 may transmit encoded video data and error correction data (e.g., bit punched error correction data) to receiving device 104 via channel 230.
[0314] Figure 27 This is a block diagram illustrating example alternative components of a receiving device 104 according to one or more technologies of this disclosure. Figure 27 The version of the receiving device 104 shown can be used with Figure 26 The version of the transmitting device 102 shown is compatible. Figure 27 In the example, the receiving device 104 includes an analog decoder 2700, a de-punching unit 2702, a channel decoder 2704, a deinterleaving unit 2706, a video decoder 2708, a reconstruction unit 2710, a frame estimation unit 2712, a video encoder 2714, a reliability unit 2716, and a feedback unit 2718.
[0315] The analog decoder 2700 can acquire analog encoded video data. The analog decoder 2700 can perform analog decoding operations to reconstruct residual data. Example details of the analog decoding operations can be found in U.S. Patent Nos. 11,553,184, 11,431,962, and 11,457,224.
[0316] For example, in some examples, the analog decoder 2700 can determine the amplitude values of multiple coefficient vectors based on an analog signal. For instance, the analog decoder 2700 can determine the phase shift and power at the symbol sampling time of the analog signal. The analog decoder 2700 can then determine a point in the IQ plane indicated by the determined phase shift and power. The analog decoder 2700 can then determine the amplitude value pairs as coordinates of the points in the IQ plane.
[0317] For each of the coefficient vectors, the analog decoder 2700 can determine the coefficients in the coefficient vector based on the amplitude value and mapping pattern of the coefficient vector. For each of a plurality of allowed coefficient vectors, the mapping pattern maps the allowed coefficient vector to a corresponding amplitude value among a plurality of amplitude values. The corresponding amplitude value is adjacent to at least one other amplitude value among a plurality of amplitude values in n-dimensional space, wherein the at least one other amplitude value is adjacent to the corresponding amplitude value in a monotonic digital line of amplitude values. Each of the coefficient vectors can include n coefficients. The value n can be greater than or equal to 2. In some examples, the analog decoder 2700 can determine the coefficients in the coefficient vector as coordinates of positions in n-dimensional space corresponding to amplitude values. The mapping pattern maps different positions in n-dimensional space to different amplitude values among a plurality of amplitude values. In some examples, the coefficient vector includes one or more negative coefficients and one or more positive coefficients. In other examples, the coefficient vector may include only non-negative coefficients.
[0318] In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain a sign value, where the sign value indicates the positive or negative sign of the coefficients in the coefficient vector. In such examples, the analog decoder 2700 may determine the absolute value of the coefficients in the coefficient vector based on the magnitude value and mapping pattern of the coefficient vector. The analog decoder 2700 may reconstruct the coefficients in the coefficient vector at least partially by applying the sign value to the absolute value of the coefficients in the coefficient vector. In some examples, as part of determining the coefficients, the analog decoder 2700 may obtain data representing shift values. In such examples, the shift value indicates the most negative coefficient in the coefficient vector. Additionally, in such examples, the analog decoder 2700 may determine the intermediate values of the coefficients in the coefficient vector based on the magnitude value and mapping pattern of the coefficient vector. The analog decoder 2700 may reconstruct the coefficients in the coefficient vector at least partially by adding shift values to each of the intermediate values of the coefficients in the coefficient vector.
[0319] Furthermore, the analog decoder 2700 can generate residual data based on the coefficients in the coefficient vector. For example, in one example, the analog decoder 2700 can dequantize the coefficients of the coefficient vector. In this example, the analog decoder 2700 can perform a debinarization process to convert the coefficients into digital sample values. For example, the analog decoder 2700 can apply an inverse DCT to the coefficients to convert the coefficients into digital sample values. In this way, the analog decoder 2700 can generate digital residual sample values.
[0320] The de-puncturing unit 2702 can obtain encoded video data and bit-punctured error correction data. The de-puncturing unit 2702 can apply a de-puncturing process to the bit-punctured error correction data to reconstruct the error correction data. The de-puncturing unit 2702 can be configured according to other parts of this disclosure. Figure 2 Any example provided by the de-drilling unit 220 for applying the de-drilling process.
[0321] Channel decoder 2704 can perform a channel decoding process, wherein the channel decoding process modifies encoded video data (e.g., encoded video data received via channel 230 or encoded video data generated by video encoder 2714 and modified by reliability unit 714 in some examples) based on error correction data. Channel decoder 2704 can be configured according to other parts of this disclosure. Figure 2 The channel decoder 222 provides any example to perform the channel decoding process.
[0322] The deinterleaving unit 2706 can perform deinterleaving operations on the error-corrected coded video data generated by the channel decoder 2704. For example, the deinterleaving process can reverse the interleaving process performed by the interleaving unit 2608 of the transmitting device 102. For example, the deinterleaving process can be performed according to the interleaving mode used by the interleaving unit 2608.
[0323] Video decoder 2708 can acquire encoded video data (e.g., deinterleaved encoded video data generated by deinterleaving unit 2706). Video decoder 2708 can perform a video decoding process on the encoded video data to reconstruct the video data frame. The video decoding process performed by video decoder 2708 can be the same as the video decoding process described in any of the examples provided elsewhere in this disclosure with respect to video decoder 224. Reconstruction unit 2710 of receiving device 104 can add residual data generated by analog decoder 2700 to the corresponding samples of the reconstructed video data generated by video decoder 2708, thereby completely reconstructing the video data frame.
[0324] In addition, Figure 27In the example, the frame estimation unit 2712 may estimate one or more frames based on previously reconstructed frames. As previously mentioned in this disclosure, the discussion of frames may be applied relative to segments of frames (such as slices). The frame estimation unit 2712 may estimate frames based on any of the examples provided elsewhere in this disclosure regarding frame estimation unit 226. The video encoder 2714 may perform a video encoding process on the estimated frames. As part of performing the video encoding process, the video encoder 2714 may determine encoding selection data, such as encoding selection data 2100, as previously discussed. The receiving device 104 may send the encoding selection data to the transmitting device 102. In some examples, the video encoder 2714 may send to the transmitting device 102 data such as information about... Figure 4 and Figure 5 The encoding parameters discussed enable the video encoder 2600 of transmitting device 102 to perform a limited encoding process in the same manner as the video encoder 2714 of receiving device 104. In some examples, the video encoder 2714 may determine a multi-view encoding cue and send it to transmitting device 102.
[0325] Reliability unit 2716 can be used in conjunction with reliability unit 1002 ( Figure 10 It operates in a roughly the same manner. Feedback unit 2718 can send predicted quality feedback (e.g., CCSF parameters) to transmitting device 102 based on the output of reliability unit 1002.
[0326] In some examples of this disclosure, the video encoder 2600 of the transmitting device 102 generates predicted data for a set of frames, and the residual generation unit 2602 can generate residual data based on the first predicted data and the first set of frames. The video encoder 2600 can apply a transform to the predicted data to generate a transform block, quantize the transform coefficients of the transform block, and apply entropy coding to the syntax elements representing the quantized transform coefficients to generate entropy-coded syntax elements. The first encoded video data may include entropy-coded syntax elements. The channel encoder 2610 can perform a channel coding process, wherein the channel coding process generates error correction data for the encoded video data, including entropy-coded syntax elements. The analog encoder 2604 can perform analog modulation on the residual data to generate analog-modulated residual data. The communication interface of the transmitting device 102 can transmit the analog-modulated residual data, the error correction data, and the encoded video data.
[0327] The communication interface of receiving device 104 can receive analog modulation residual data and error correction data and extracted video data from transmitting device. Extracted video data may include encoded video data for which decimation mode has been applied. Encoded video data is generated based on a set of video data frames. Channel decoder 2704 can apply an error correction process to generate error-corrected encoded video data based on the encoded video data and error correction data. Error-corrected encoded video data includes entropy-coded syntax elements representing quantization transform coefficients. Video decoder 2708 can perform a decoding process to reconstruct a second set of frames based on the error-corrected encoded video data. As part of the decoding process to reconstruct the frame set, video decoder 2708 can apply entropy decoding to the syntax elements to obtain quantization transform coefficients, inverse quantize the quantization transform coefficients to generate inverse quantization transform coefficients, and apply an inverse transform to the inverse quantization transform coefficients to generate prediction data. Analog decoder 2700 can demodulate analog modulation residual data to obtain residual data. Reconstruction unit 2710 can reconstruct the frame set based on the prediction data and residual data.
[0328] The following is a non-limiting list of the terms of one or more technologies under this disclosure.
[0329] Clause 1A. A method for decoding video data, the method comprising: obtaining error correction data from a transmitting device at a receiving device, wherein the error correction data provides error correction information and is generated based on coded video data of one or more blocks of a frame of the video data; generating prediction data of the frame at the receiving device using one or more decoding tools not used to generate the coded video data of one or more blocks, wherein the prediction data of the frame includes predictions of blocks of the frame based at least in part on one or more previously reconstructed frames of the video data; generating coded video data at the receiving device based on the prediction data of the frame; generating error-corrected coded video data at the receiving device using the error correction data to perform an error correction operation on the coded video data; and performing a reconstruction operation at the receiving device to reconstruct blocks of the frame based on the error-corrected coded video data, wherein the reconstruction operation is controlled by values of one or more parameters.
[0330] Clause 2A. The method described in Clause 1A further includes receiving the value of the parameter from the transmitting device at the receiving device.
[0331] Clause 3A. The method described in Clause 1A further includes determining the value of the parameter at the receiving device, without receiving the value of the parameter from the transmitting device.
[0332] Clause 4A. The method according to any one of Clauses 1A-3A, wherein: the parameters include one or more quantization parameters, generating coded video data includes using the quantization parameters to quantize transform coefficients generated from scene-based prediction data, and performing the reconstruction operation includes using the quantization parameters to inverse quantize the transform coefficients of the error-correcting coded video data.
[0333] Clause 5A. The method according to Clause 4A, wherein the method further comprises: calculating a quantization parameter based on the entropy ratio of the quantized transform coefficients to the unquantized transform coefficients.
[0334] Clause 6A. The method according to any one of Clauses 1A-5A, wherein: the parameters include a transform size parameter, generating coded video data includes applying a forward transform with a transform size indicated by the transform size parameter to the sampled domain data of the picture, and performing a reconstruction operation includes applying an inverse transform with a transform size indicated by the transform size parameter to the transform coefficients of the error-correcting coded video data.
[0335] Clause 7A. The method according to any one of Clauses 1A-6A, wherein: the parameters include parameters indicating the number of transform coefficients; generating coded video data includes including a set of transform coefficients in the coded video data, wherein the set of transform coefficients includes the indicated number of transform coefficients; and performing the reconstruction operation includes parsing the set of transform coefficients from the error-corrected coded video data, wherein the set of transform coefficients includes the indicated number of transform coefficients.
[0336] Clause 8A. The method according to Clause 7A, wherein: obtaining coded video data and error correction data includes receiving coded video data and error correction data from a transmitting device via a communication channel at a receiving device, and the method further includes applying an optimization process, wherein the optimization process determines the number of transform coefficients based on the signal-to-noise ratio of the data transmitted on the communication channel.
[0337] Clause 9A. The method according to any one of Clauses 1A-8A, wherein: the parameters include a bit width parameter for a plurality of index values; for each corresponding index value among the plurality of index values: performing a reconstruction operation includes parsing a first set of bits from error-corrected coded video data, wherein the first set of bits indicates transform coefficients having the corresponding index value, and the number of bits in the first set of bits is equal to the bit width indicated by the bit width parameter of the corresponding index value; generating coded video data includes including a second set of bits in the coded video data, wherein the second set of bits indicates transform coefficients having the corresponding index value, and the number of bits in the second set of bits is equal to the bit width indicated by the bit width parameter of the corresponding index value; and parsing a third set of bits from the error-corrected coded video data, wherein the third set of bits indicates transform coefficients having the corresponding index value, and the number of bits in the third set of bits is equal to the bit width indicated by the bit width parameter of the corresponding index value.
[0338] Clause 10A. The method according to any one of Clauses 1A-9A, wherein the method further comprises performing a bit de-punching operation on the error-corrected data before generating the error-corrected coded video data.
[0339] Clause 11A. The method according to any one of Clauses 1A-10A, wherein the parameters include one or more of the following: color space, transform size, quantization parameter, number of transform coefficients in the first coded video data, or number of bits for each transform coefficient in the first coded video data.
[0340] Clause 12A. The method according to any one of Clauses 1A-11A, wherein: the extraction mode defines the mode of anchor transform blocks and non-anchor transform blocks in the frame, the method further comprising receiving system bits of anchor transform blocks at a receiving device but not receiving system bits of non-anchor transform blocks, the system bits of anchor transform blocks representing transform coefficients in the anchor transform blocks, and the system bits of non-anchor transform blocks representing reduced bit-depth versions of the original transform coefficients in the non-anchor transform blocks; the error correction data includes error correction data of anchor transform blocks and error correction data of non-anchor transform blocks, wherein the error correction data of non-anchor transform blocks is based on the original transform coefficients in the non-anchor transform blocks, and generating error-corrected coded video data includes: performing error correction on the system bits of anchor transform blocks using the error correction data of anchor transform blocks; and performing error correction on portions of coded video data corresponding to non-anchor transform blocks using the error correction data of non-anchor transform blocks.
[0341] Clause 13A. The method described in Clause 12A further includes: determining a decimation mode at the receiving device; and sending the decimation mode from the receiving device to the transmitting device.
[0342] Clause 14A. The method according to any one of Clauses 1A-13A, wherein: the extraction mode defines the mode of anchor transform blocks and non-anchor transform blocks in the frame, the transform coefficients in the non-anchor transform blocks having a reduced bit depth relative to the anchor transform blocks; the receiving device receives the system bits of the anchor transform blocks, the system bits of the non-anchor transform blocks, and a correlation matrix, the system bits of the anchor transform blocks representing the transform coefficients in the anchor transform blocks, and the system bits of the non-anchor transform blocks representing a reduced bit depth version of the original transform coefficients in the non-anchor transform blocks; the reconstruction operation includes, for each non-anchor transform coefficient in the non-anchor transform blocks: calculating an interpolated value of the non-anchor transform coefficient at the receiving device based on the correlation matrix and the corresponding anchor transform coefficient; and calculating a reconstructed value of the non-anchor transform coefficient at the receiving device based on the interpolated value of the non-anchor transform coefficient in the error-corrected coded video data and the value of the non-anchor transform coefficient.
[0343] Clause 15A. A method for encoding video data, the method comprising: obtaining video data from a video source at a transmitting device; generating encoded video data of a first frame and encoded video data of a second frame of the video data based on a parameter set at the transmitting device; performing channel coding on the encoded video data of the first frame and the encoded video data of the second frame at the transmitting device to generate error-corrected data of the first frame and error-corrected data of the second frame; and transmitting the encoded video data of the first frame, the error-corrected data of the first frame, and the error-corrected data of the second frame at the transmitting device.
[0344] Clause 16A. The method described in Clause 15A further includes sending the value of the parameter from the transmitting device to the receiving device.
[0345] Clause 17A. The method according to Clause 16A, wherein: the parameters include one or more quantization parameters, and generating coded video data includes using the quantization parameters to quantize the transform coefficients of the first frame and the transform coefficients of the second frame.
[0346] Clause 18A. The method according to any one of Clauses 16A-17A, wherein: the parameters include a transform size parameter, and generating coded video data includes applying a forward transform to residual data blocks of a first frame and residual data blocks of a second frame, wherein the forward transform has a transform size indicated by the transform size parameter.
[0347] Clause 19A. The method according to any one of Clauses 16A-18A, wherein: the parameters include parameters indicating the number of transform coefficients, and generating coded video data of the first frame and coded video data of the second frame includes a set of transform coefficients in the coded video data of the first frame and the coded video data of the second frame, the set of transform coefficients including the indicated number of transform coefficients.
[0348] Clause 20A. The method according to any one of Clauses 16A to 10A, wherein the parameters include one or more of the following: color space, transform size, quantization parameter, number of transform coefficients in the encoded video data, or number of bits per transform coefficient in the encoded video data.
[0349] Clause 21A. A method for encoding video data, the method comprising: obtaining video data from a video source at a transmitting device; generating transform blocks based on the video data at the transmitting device; determining, at the transmitting device, which transform blocks in the transform blocks are anchor transform blocks; calculating a correlation matrix of a set of transform blocks at the transmitting device; generating a bit-reduced non-anchor transform matrix at the transmitting device; and transmitting the anchor transform blocks, non-anchor transform blocks, and correlation matrix to a receiving device at the transmitting device.
[0350] Clause 22A. The method according to Clause 21A further includes receiving an indication of the extraction mode from the receiving device at the transmitting device.
[0351] Clause 23A. An apparatus comprising: a memory configured to store video data; a communication interface; and one or more processes implemented in a circuit and coupled to the memory, the one or more processors being configured to perform the method according to any one of Clauses 1A-22A.
[0352] Clause 24A. An apparatus comprising components for performing the method according to any one of Clauses 1A-22A.
[0353] Clause 25A. A computer-readable data storage medium having instructions stored thereon, which, when executed, cause a device to perform the method according to any one of Clauses 1A-22A.
[0354] Clause 1B. An apparatus for processing video data, the apparatus comprising: a memory configured to store video data; and a communication interface configured to obtain error correction data from a transmitting device, wherein the error correction data provides error correction information about a frame of the video data; one or more processes implemented in circuitry and coupled to the memory, the one or more processors configured to: generate prediction data for a frame, wherein the prediction data for the frame includes predictions of blocks of the frame based at least in part on one or more previously reconstructed frames of the video data; generate coded video data based on the prediction data for the frame, wherein the coded video data includes transform blocks, wherein the transform blocks include transform coefficients; scale the bits of the transform coefficients of the transform blocks based on reliability values of bit positions; use the error correction data to generate error-corrected coded video data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstruct a frame based on the error-corrected coded video data.
[0355] Clause 2B. The device as described in Clause 1B, wherein one or more processors are further configured to generate a reliability value at the receiving device.
[0356] Clause 3B. The apparatus as described in Clause 2B, wherein one or more processors are configured to generate a reliability value based on statistics regarding the occurrence of errors in bit positions.
[0357] Clause 4B. The device pursuant to any one of Clauses 2B or 3B, wherein one or more processors are configured to generate a reliability value based on the reliability characteristics of individual regions of a frame of video data.
[0358] Clause 5B. An apparatus pursuant to any one of Clauses 2B-4B, wherein one or more processors are configured to generate reliability values based on a noise model.
[0359] Clause 6B. The device according to any one of Clauses 1B-5B, wherein the communication interface is further configured to send a reliability value to the transmitting device.
[0360] Clause 7B. The device according to any one of Clauses 1B-5B, wherein the communication interface is further configured to receive reliability values from the transmitting device.
[0361] Clause 8B. An apparatus for processing video data, the method comprising: a memory configured to store video data; and one or more processes implemented in circuitry and coupled to the memory, the one or more processors configured to: acquire the video data; acquire predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated frame generated by a receiving device; adjust one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; perform a video coding process to generate coded video data based on one or more frames of the acquired video data, wherein the video coding process is controlled by the video coding parameters; perform channel coding processing on the coded video data to generate channel-coded data, wherein the channel coding processing is controlled by the channel coding parameters; and a communication interface configured to transmit the channel-coded data to the receiving device.
[0362] Clause 9B. The apparatus according to Clause 8B, wherein: the video coding parameters include quantization parameters, one or more processors are configured to: as part of adjusting the video coding parameters, adjust the quantization parameters, and one or more processors are configured to: as part of performing the video coding process, use the quantization parameters to quantize the transform coefficients of one or more frames of transform blocks.
[0363] Clause 10B. An apparatus according to any one of Clauses 8B-9B, wherein: the channel coding parameters include a low-density parity-check (LDPC) graph, one or more processors are configured to: adjust the LDPC graph as part of adjusting the channel coding parameters, and one or more processors are configured to: use the LDPC graph to generate codewords included in the channel coding data as part of performing the channel coding process.
[0364] Clause 11B. The apparatus according to any one of Clauses 8B-10B, wherein: the channel-coded data includes error-correcting data, one or more processors are further configured to adjust one or more bit-puncturing parameters based on prediction quality feedback, and one or more processors are configured to perform a bit-puncturing process on the error-correcting data, wherein the bit-puncturing process is controlled by one or more bit-puncturing parameters.
[0365] Clause 11B. A method of processing video data, the method comprising: obtaining error correction data from a transmitting device at a receiving device, wherein the error correction data provides error correction information about a frame of video data; generating prediction data of the frame at the receiving device, wherein the prediction data of the frame includes predictions of blocks of the frame based at least in part on one or more previously reconstructed frames of the video data; generating coded video data at the receiving device based on the prediction data of the frame, wherein the coded video data includes transform blocks, wherein the transform blocks include transform coefficients; scaling bits of the transform coefficients of the transform blocks at the receiving device based on reliability values of bit positions; generating error-corrected coded video data at the receiving device using the error correction data to perform an error correction operation on the scaled bits of the transform coefficients of the transform blocks; and reconstructing a frame at the receiving device based on the error-corrected coded video data.
[0366] Clause 12B. The method described in Clause 11B further includes generating a reliability value at the receiving device.
[0367] Clause 13B. The method according to Clause 12B, wherein generating a reliability value includes generating a reliability value at the receiving device based on statistics regarding the occurrence of errors in bit positions.
[0368] Clause 14B. The method according to any one of Clauses 12B or 13B, wherein generating the reliability value includes generating the reliability value at the receiving device based on the reliability characteristics of individual regions of the video data frame.
[0369] Clause 15B. The method according to any one of Clauses 12B-14B, wherein generating the reliability value includes generating the reliability value at the receiving device based on a noise model.
[0370] Clause 16B. The method according to any one of Clauses 11B-15B further includes transmitting a reliability value from the receiving device to the transmitting device.
[0371] Clause 17B. The method according to any one of Clauses 11B-15B further includes receiving a reliability value from the transmitting device at the receiving device.
[0372] Clause 18B. A method for processing video data, the method comprising: acquiring video data; acquiring predictive quality feedback, wherein the predictive quality feedback is based on the reliability of an estimated frame generated by a receiving device; adjusting one or more of video coding parameters or channel coding parameters based on the predictive quality feedback; performing a video coding process to generate coded video data based on one or more frames of the acquired video data, wherein the video coding process is controlled by the video coding parameters; performing channel coding processing on the coded video data to generate channel-coded data, wherein the channel coding processing is controlled by the channel coding parameters; and transmitting the channel-coded data to the receiving device.
[0373] Clause 19B. The method according to Clause 18B, wherein: video coding parameters include quantization parameters, adjusting video coding parameters includes adjusting quantization parameters, and performing a video coding process includes using quantization parameters to quantize the transform coefficients of one or more transform blocks of a frame.
[0374] Clause 20B. The method according to any one of Clauses 18B-19B, wherein: the channel coding parameters include a low-density parity-check (LDPC) graph, adjusting the channel coding parameters includes adjusting the LDPC graph, and performing the channel coding process includes using the LDPC graph to generate codewords included in the channel-coded data.
[0375] Clause 21B. The method according to any one of Clauses 18B-20B, wherein: the channel-coded data includes error-correcting data, and the method further comprises: adjusting one or more bit puncturing parameters based on prediction quality feedback, and performing a bit puncturing process on the error-correcting data, wherein the bit puncturing process is controlled by one or more bit puncturing parameters.
[0376] Clause 22B. An apparatus comprising components for performing the method according to any one of Clauses 11B-21B.
[0377] Clause 23B. A computer-readable data storage medium having instructions stored thereon, which, when executed, cause a device to perform the method according to any one of Clauses 11B-21B.
[0378] Clause 1C. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: acquire a first set of multi-view frames of the video data, wherein the first set of multi-view frames includes a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; transmit first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multi-view frames; receive a multi-view encoding prompt from the receiving device; acquire a second set of multi-view frames of the video data, wherein the second set of multi-view frames includes a third frame and a fourth frame, the third frame being from the first viewpoint and the fourth frame being from the second viewpoint; perform a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompt received from the receiving device to generate second encoded video data, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame; and transmit the second encoded video data to the receiving device.
[0379] Clause 2C. The apparatus according to Clause 1, wherein one or more processors are further configured to: receive an updated multiview encoding prompt from the receiving device after sending second encoded video data to the receiving device; obtain a third set of multiview frames of the video data, wherein the third set of multiview frames includes a fifth frame and a sixth frame, the fifth frame being from a first viewpoint and the sixth frame being from a second viewpoint; encode the third set of multiview frames based on the updated multiview encoding prompt received from the receiving device to generate third encoded video data; and send the third encoded video data to the receiving device.
[0380] Clause 3C. The device according to any one of Clauses 1C-2C, wherein the multi-view encoding prompt includes one or more of the following: relative shift between blocks of the first and second screens, brightness correction between the first and second screens, inter-block shift between anchor blocks and reconstructed blocks, or motion data for reference shift.
[0381] Clause 4C. A device according to any one of Clauses 1C-3C, wherein: the device is an extended reality (XR) headset, and one or more processors are configured to: receive from a receiving device virtual element data generated based on a first set of multi-view images and a second set of multi-view images; and output the virtual element data for display in an XR scene.
[0382] Clause 5C. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: obtain first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multi-view frames of the video data, the first set of multi-view frames including a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; determine a multi-view encoding cue based on the first encoded video data; transmit the multi-view encoding cue to the transmitting device; and obtain second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multi-view frames including a third frame and a fourth frame, the second encoded video data being encoded using a multi-view encoding process, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame based on the multi-view encoding cue.
[0383] Clause 6C. The apparatus as described in Clause 5C, wherein one or more processors are further configured to decode second coded video data.
[0384] Clause 7C. The device according to any one of Clauses 5C-6C, wherein the multiview encoding prompt is a first multiview encoding prompt, and one or more processors are further configured to: determine a second multiview encoding prompt based on second encoded video data; send the second multiview encoding prompt to a transmitting device; and obtain third encoded video data from the transmitting device, wherein the third encoded video data is based on a third set of multiview frames including a fifth frame and a sixth frame, the third encoded video data being encoded using a multiview encoding process, wherein the multiview encoding process is based on the second multiview encoding prompt to reduce interview redundancy between the fifth frame and the sixth frame.
[0385] Clause 8C. The device according to any one of Clauses 5C-7C, wherein: the multi-view encoding prompt includes a depth map indicating the depth of an object represented in a first screen and a second screen, and one or more processors are configured to determine the depth map based on the first screen and the second screen as part of determining the multi-view encoding prompt.
[0386] Clause 9C. The device as described in Clauses 5C to 8C, wherein the multi-view encoded prompt includes one or more illuminance compensation factors, and one or more processors are configured to determine the illuminance compensation factors based on a first screen and a second screen as part of determining the multi-view encoded prompt.
[0387] Clause 10C. An apparatus according to any one of Clauses 5C-9C, wherein: the transmitting device is an extended reality (XR) headset, and one or more processors are further configured to: process a second set of images to generate virtual element data; and transmit the virtual element data to the XR headset.
[0388] Clause 11C. A method for processing video data, the method comprising: obtaining a first set of multi-view frames of the video data, wherein the first set of multi-view frames includes a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multi-view frames; receiving a multi-view encoding prompt from the receiving device; obtaining a second set of multi-view frames of the video data, wherein the second set of multi-view frames includes a third frame and a fourth frame, the third frame being from the first viewpoint and the fourth frame being from the second viewpoint; performing a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompt received from the receiving device to generate second encoded video data, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame; and transmitting the second encoded video data to the receiving device.
[0389] Clause 12C. The method according to Clause 11C further comprises: receiving an update multiview encoding prompt from the receiving device after sending the second encoded video data to the receiving device; obtaining a third set of multiview frames of the video data, wherein the third set of multiview frames includes a fifth frame and a sixth frame, the fifth frame being from a first viewpoint and the sixth frame being from a second viewpoint; encoding the third set of multiview frames based on the update multiview encoding prompt received from the receiving device to generate third encoded video data; and sending the third encoded video data to the receiving device.
[0390] Clause 13C. The method according to any one of Clauses 11C-12C, wherein the multi-view encoding prompt includes one or more of the following: relative shift between blocks of the first and second screens, brightness correction between the first and second screens, inter-block shift between anchor blocks and reconstructed blocks, or motion data for reference shift.
[0391] Clause 14C. The method according to any one of Clauses 11C-13C, wherein the method further comprises: receiving from a receiving device virtual element data generated based on a first set and a second set of multi-view images; and outputting the virtual element data for display in an extended reality (XR) scene.
[0392] Clause 15C. A method for processing video data, the method comprising: obtaining first encoded video data from a transmitting device, wherein the first encoded video data is based on a first set of multi-view frames of the video data, the first set of multi-view frames including a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; determining a multi-view encoding cue based on the first encoded video data; sending the multi-view encoding cue to the transmitting device; and obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multi-view frames including a third frame and a fourth frame, the second encoded video data being encoded using a multi-view encoding process, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame based on the multi-view encoding cue.
[0393] Clause 16C. The method described in Clause 15C further includes decoding the second coded video data.
[0394] Clause 17C. The method according to any one of Clauses 15C-16C, wherein the multiview coding prompt is a first multiview coding prompt, and the method further includes: determining a second multiview coding prompt based on second coded video data; sending the second multiview coding prompt to a transmitting device; obtaining third coded video data from the transmitting device, wherein the third coded video data is based on a third set of multiview frames including a fifth frame and a sixth frame, the third coded video data being encoded using a multiview coding process, wherein the multiview coding process reduces interview redundancy between the fifth frame and the sixth frame based on the second multiview coding prompt.
[0395] Clause 18C. The method according to any one of Clauses 15C-17C, wherein: the multi-view encoding prompt includes a depth map indicating the depth of objects represented in a first screen and a second screen, and determining the multi-view encoding prompt includes determining the depth map based on the first screen and the second screen.
[0396] Clause 19C. The method according to Clauses 15C-18C, wherein the multi-view encoding prompt includes one or more illuminance compensation factors, and determining the multi-view encoding prompt includes determining the illuminance compensation factors based on a first screen and a second screen.
[0397] Clause 20C. The method according to any one of Clauses 15C-19C, wherein: the transmitting device is an extended reality (XR) headset, and the method further comprises: processing a second set of images to generate virtual element data; and transmitting the virtual element data to the XR headset.
[0398] Clause 21C. An apparatus comprising: means for acquiring a first set of multi-view frames of video data, wherein the first set of multi-view frames includes a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; means for transmitting first encoded video data to a receiving device, wherein the first encoded video data is based on the first set of multi-view frames; means for receiving a multi-view encoding prompt from the receiving device; means for acquiring a second set of multi-view frames of video data, wherein the second set of multi-view frames includes a third frame and a fourth frame, the third frame being from the first viewpoint and the fourth frame being from the second viewpoint; means for performing a multi-view encoding process on the second set of multi-view frames based on the multi-view encoding prompt received from the receiving device to generate second encoded video data, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame; and means for transmitting the second encoded video data to the receiving device.
[0399] Clause 22C. An apparatus comprising: means for obtaining first encoded video data from a transmitting device, wherein the first encoded video data is a first set of multi-view frames based on the video data, the first set of multi-view frames including a first frame and a second frame, the first frame being from a first viewpoint and the second frame being from a second viewpoint; means for determining multi-view encoding cues based on the first encoded video data; means for sending the multi-view encoding cues to the transmitting device; and means for obtaining second encoded video data from the transmitting device, wherein the second encoded video data is based on a second set of multi-view frames including a third frame and a fourth frame, the second encoded video data being encoded using a multi-view encoding process, wherein the multi-view encoding process reduces inter-view redundancy between the third frame and the fourth frame based on the multi-view encoding cues.
[0400] Clause 1D. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: encode a first set of frames of the video data to generate first encoded video data; transmit the first encoded video data to a receiving device; receive from the receiving device a decimation mode indication indicating a decimation mode determined based on the first set of frames, wherein the decimation mode is a mode in which the encoded video data is not transmitted; encode a second set of frames of the video data to generate second encoded video data; apply the decimation mode to the second encoded video data to generate decimated video data; and transmit the decimated video data to the receiving device.
[0401] Clause 2D. The apparatus according to Clause 1D, wherein one or more processors are configured to: generate first error correction data based on first coded video data; transmit the first error correction data to a receiving device; generate second error correction data based on second coded video data; and transmit the second error correction data to a receiving device.
[0402] Clause 3D. The device according to any one of Clauses 1D-2D, wherein the extraction mode indicates a mode for skipping the transmission of encoded video data of the full frame.
[0403] Clause 4D. The device according to any one of Clauses 1D-3D, wherein the extraction mode indicates a mode for transmitting encoded video data that skips a specific region within the frame.
[0404] Clause 5D. The device according to any one of Clauses 1D-4D, wherein the video data is multi-view video data and the extraction mode indicates a mode for skipping the transmission of encoded video data from a particular view.
[0405] Clause 6D. The device according to any one of Clauses 1D-5D, wherein: the decimation mode indication is a first decimation mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, the decimated video data is first decimated video data, and one or more processors are further configured to: encode a third set of frames of video data to generate third encoded video data; determine a second decimation mode indicating a second mode in which encoded video data is not transmitted; apply the second decimation mode to the third encoded video data to generate second decimated video data; transmit the second decimated video data to a receiving device; and transmit a second decimation mode indication to the receiving device, the second decimation mode indication indicating that the second decimation mode is applied to the third encoded video data.
[0406] Clause 7D. An apparatus according to any one of Clauses 1D-6D, wherein: one or more processors are configured as part of encoding a first set of frames: generating first prediction data for the first set of frames; generating residual data based on the first prediction data and the first set of frames; applying a transform to the first prediction data to generate a transform block; quantizing the transform coefficients of the transform block; applying entropy coding to syntax elements representing the quantized transform coefficients to generate a first entropy-coded syntax element, wherein the first encoded video data includes the first entropy-coded syntax element; the one or more processors are further configured to perform analog modulation on the residual data to generate first analog-modulated residual data, and the apparatus further includes a communication interface configured to transmit the first analog-modulated residual data and the first encoded video data.
[0407] Clause 8D. A device according to any one of Clauses 1D-7D, wherein: the device is an extended reality (XR) headset and includes a display system, and one or more processors are further configured to: receive virtual element data from a receiving device; and the display system is configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0408] Clause 9D. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: receive first coded video data from a transmitting device; perform a decoding process to reconstruct a first set of frames based on the first coded video data; determine, based on the first set of frames, an extraction mode indicating a mode in which the coded video data has not been transmitted; send to the transmitting device an extraction mode indication indicating the determined extraction mode; receive extracted video data from the transmitting device, wherein the extracted video data includes second coded video data to which the extraction mode has been applied, wherein the second coded video data is generated based on a second set of frames of the video data; and perform a decoding process to reconstruct a second set of frames based on the second coded video data.
[0409] Clause 10D. The apparatus according to Clause 9D, wherein one or more processors are further configured to: receive first error-correcting data from a transmitting device; apply an error-correcting process to modify first coded video data based on the first error-correcting data to generate first error-corrected coded video data; wherein the one or more processors are configured to perform a decoding process to reconstruct a first set of images based on the first error-corrected coded video data; wherein the one or more processors are further configured to: receive second error-correcting data from a transmitting device; apply an error-correcting process to generate second error-corrected coded video data based on the second coded video data and the second error-correcting data; and wherein the one or more processors are configured to perform a decoding process to reconstruct a second set of images based on the second error-corrected coded video data.
[0410] Clause 11D. The device according to any one of Clauses 9D-10D, wherein the extraction mode indicates a mode for skipping the transmission of encoded video data that skips the entire frame.
[0411] Clause 12D. An apparatus according to any one of Clauses 9D-11D, wherein one or more processors are configured as part of determining an extraction mode: applying the extraction mode to first coded video data to generate decimated coded video data; applying an error correction process to modify the decimated coded video data based on the first error correction data to generate trial error-corrected video data; applying a decoding process to reconstruct a first set of frames based on the trial error-corrected video data; and determining whether the extraction mode satisfies the standard based on a comparison between the first set of frames reconstructed as based on the trial error-corrected video data and the first set of frames reconstructed as based on the first video data.
[0412] Clause 13D. The device according to any one of Clauses 9D-12D, wherein the extraction mode indicates a mode for skipping the transmission of encoded video data in a specified area within the frame.
[0413] Clause 14D. The device according to any one of Clauses 9D-13D, wherein the video data is multi-view video data and the extraction mode indicates a mode for skipping the transmission of encoded video data from the specified view.
[0414] Clause 15D. The apparatus according to any one of Clauses 9D-14D, wherein: the decimation mode indication is a first decimation mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, the decimated video data is first decimated video data, and one or more processors are further configured to: receive a second decimation mode indication indicating a second mode in which encoded video data is not transmitted; receive second decimated video data from a transmitting device, wherein the second decimated video data includes third encoded video data to which the second decimation mode has been applied, wherein the third encoded video data is generated based on a third set of frames of video data; and apply a decoding process to reconstruct the third set of frames based on the third encoded video data.
[0415] Clause 16D. The apparatus according to any one of Clauses 9D to 15D, wherein: the apparatus further includes a communication interface configured to receive analog modulation residual data, the second coded video data including entropy-coded syntax elements representing quantization transform coefficients; one or more processors are configured as part of an application decoding process to reconstruct a second set of images: applying entropy decoding to the syntax elements to obtain quantization transform coefficients; dequantizing the quantization transform coefficients to generate inverse quantization transform coefficients; applying the inverse transform to the inverse quantization transform coefficients to generate prediction data; demodulating the analog modulation residual data to obtain residual data; and reconstructing a second set of images based on the prediction data and the residual data.
[0416] Clause 17D. A device according to any one of Clauses 9D-16D, wherein: one or more processors are further configured to process a second set of frames to generate virtual element data, and the transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0417] Clause 18D. A method comprising: encoding a first set of frames of video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving from the receiving device a decimation mode indication indicating a decimation mode determined based on the first set of frames, wherein the decimation mode is a mode in which the encoded video data is not transmitted; encoding a second set of frames of video data to generate second encoded video data; applying the decimation mode to the second encoded video data to generate decimated video data; and transmitting the decimated video data to the receiving device.
[0418] Clause 19D. The method according to Clause 18D further includes: generating first error correction data based on first coded video data; sending the first error correction data to a receiving device; generating second error correction data based on second coded video data; and sending the second error correction data to the receiving device.
[0419] Clause 20D. The method according to any one of Clauses 18D-19D, wherein the extraction mode indicates a mode for transmitting encoded video data that skips the entire frame.
[0420] Clause 21D. The method according to any one of Clauses 18D-20D, wherein the extraction mode indicates a mode for transmitting coded video data that skips a specific region within the frame.
[0421] Clause 22D. The method according to any one of Clauses 18D-21D, wherein the video data is multi-view video data, and the extraction mode indicates a mode for skipping the transmission of encoded video data from a particular view.
[0422] Clause 23D. The method according to any one of Clauses 18D to 22D, wherein: the decimation mode indication is a first decimation mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, the decimated video data is first decimated video data, and the method further comprises: encoding a third set of frames of video data to generate third encoded video data; determining a second decimation mode indicating a second mode in which encoded video data is not transmitted; applying the second decimation mode to the third encoded video data to generate second decimated video data; transmitting the second decimated video data to a receiving device; and transmitting the second decimation mode to the receiving device, the second decimation mode indication indicating that the second decimation mode is applied to the third encoded video data.
[0423] Clause 24D. The method according to any one of Clauses 18D-23D, wherein: encoding a first set of frames comprises: generating first prediction data for the first set of frames; generating residual data based on the first prediction data and the first set of frames; applying a transform to the first prediction data to generate a transform block; quantizing the transform coefficients of the transform block; applying entropy coding to a syntax element representing the quantized transform coefficients to generate a first entropy-coded syntax element, wherein the first encoded video data includes the first entropy-coded syntax element; the method further comprises: performing analog modulation on the residual data to generate first analog-modulated residual data, and transmitting the first analog-modulated residual data and the first encoded video data.
[0424] Clause 25D. The method according to any one of Clauses 18D-24D, wherein the method further comprises: receiving virtual element data from a receiving device; and displaying one or more virtual elements in an extended reality (XR) scene based on the virtual element data.
[0425] Clause 26D. A method comprising: receiving first coded video data from a transmitting device; applying a decoding process to reconstruct a first set of frames based on the first coded video data; determining, based on the first set of frames, an extraction mode indicating that the coded video data has not been transmitted; sending to the transmitting device an extraction mode indication indicating the determined extraction mode; receiving extracted video data from the transmitting device, wherein the extracted video data includes second coded video data to which the extraction mode has been applied, wherein the second coded video data is generated based on a second set of frames of the video data; and performing a decoding process to reconstruct the second set of frames based on the second coded video data.
[0426] Clause 27D. The method according to Clause 26D, wherein the method further comprises: receiving first error correction data from a transmitting device; applying an error correction process to modify first coded video data based on the first error correction data to generate first error-corrected coded video data; wherein performing a decoding process to reconstruct a first set of images includes performing a decoding process to reconstruct a first set of images based on the first error-corrected coded video data; wherein the method further comprises: receiving second error correction data from a transmitting device; applying an error correction process to generate second error-corrected coded video data based on the second coded video data and the second error correction data; and wherein performing a decoding process to reconstruct a second set of images includes performing a decoding process to reconstruct a second set of images based on the second error-corrected coded video data.
[0427] Clause 28D. The method according to any one of Clauses 26D-27D, wherein the extraction mode indicates a mode for transmitting encoded video data that skips the entire frame.
[0428] Clause 29D. The method according to any one of Clauses 26D-28D, wherein determining the extraction pattern comprises: applying the extraction pattern to first coded video data to generate decimated coded video data; applying an error correction process to modify the decimated coded video data based on the first error correction data to generate experimental error-corrected video data; applying a decoding process to reconstruct a first set of frames based on the experimental error-corrected video data; and determining whether the extraction pattern satisfies the standard based on a comparison between the first set of frames reconstructed as based on the experimental error-corrected video data and the first set of frames reconstructed as based on the first video data.
[0429] Clause 30D. The method according to any one of Clauses 26D-29D, wherein the extraction mode indicates a mode for transmitting coded video data that skips a specified area within the frame.
[0430] Clause 31D. The method according to any one of Clauses 26D-30D, wherein the video data is multi-view video data, and the extraction mode indicates a mode for skipping the transmission of encoded video data from the specified view.
[0431] Clause 32D. The method according to any one of Clauses 26D-31D, wherein: the decimation mode indication is a first decimation mode indication, the mode in which encoded video data is not transmitted is a first mode in which encoded video data is not transmitted, the decimated video data is first decimated video data, and the method further comprises: receiving a second decimation mode indication indicating a second mode in which encoded video data is not transmitted; receiving second decimated video data from a transmitting device, wherein the second decimated video data includes third encoded video data to which the second decimation mode has been applied, wherein the third encoded video data is generated based on a third set of frames of video data; and applying a decoding process to reconstruct the third set of frames based on the third encoded video data.
[0432] Clause 33D. The method according to any one of Clauses 26D to 32D, wherein: the method further comprises receiving analog modulation residual data, the second coded video data including entropy-coded syntax elements representing quantization transform coefficients; applying a decoding process to reconstruct a second set of images comprising: applying entropy decoding to the syntax elements to obtain quantization transform coefficients; inverse quantizing the quantization transform coefficients to generate inverse quantization transform coefficients; applying the inverse transform to the inverse quantization transform coefficients to generate prediction data; demodulating the analog modulation residual data to obtain residual data; and reconstructing a second set of images based on the prediction data and the residual data.
[0433] Clause 34D. The method according to any one of Clauses 26D to 33D, wherein: the method further comprises processing a second set of images to generate virtual element data, and the transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0434] Clause 35D. An apparatus comprising: means for encoding a first set of frames of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device an indication of a decimation mode determined based on the first set of frames, wherein the decimation mode is a mode in which the encoded video data is not transmitted; means for encoding a second set of frames of video data to generate second encoded video data; means for applying the decimation mode to the second encoded video data to generate decimated video data; and means for transmitting the decimated video data to the receiving device.
[0435] Clause 36D. An apparatus comprising: means for receiving first encoded video data from a transmitting device; means for performing a decoding process to reconstruct a first set of frames based on the first encoded video data; means for determining an extraction mode based on the first set of frames that indicates a mode in which the encoded video data has not been transmitted; means for sending to the transmitting device an extraction mode indication indicating the determined extraction mode; means for receiving extracted video data from the transmitting device, wherein the extracted video data includes second encoded video data to which the extraction mode has been applied, wherein the second encoded video data is generated based on a second set of frames of the video data; and means for performing a decoding process to reconstruct a second set of frames based on the second encoded video data.
[0436] Clause 1E. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: encode a first frame of the video data to generate first encoded video data; transmit the first encoded video data to a receiving device; receive from the receiving device encoding selection data for a second frame of the video data, wherein: the encoding selection data for the second frame indicates encoding selections for encoding an estimate of the second frame, and the second frame follows the first frame in decoding order; encode the second frame based on the encoding selection data for the second frame to generate second encoded video data; and transmit the second encoded video data to the receiving device.
[0437] Clause 2E. The apparatus according to Clause 1E, wherein: the encoding selection data received from the receiving device includes motion parameters of blocks of a second frame, one or more processors are configured as part of encoding the second frame to perform motion compensation based on the motion parameters of blocks of the second frame to generate prediction blocks, and the second encoded video data includes encoded video data based on prediction blocks.
[0438] Clause 3E. An apparatus pursuant to any one of Clauses 1E-2E, wherein: the encoding selection data received from the receiving apparatus includes intra-prediction parameters of blocks of the second frame, one or more processors are configured as part of encoding the second frame to perform intra-prediction based on the intra-prediction parameters of blocks of the second frame to generate prediction blocks, and the second encoded video data includes encoded video data based on the prediction blocks.
[0439] Clause 4E. The device pursuant to any one of Clauses 1E-3E, wherein the second encoded video data does not include encoding selection data.
[0440] Clause 5E. The device pursuant to any one of Clauses 1E-4E, wherein one or more processors are configured to perform entropy decoding on encoded selection data for a second screen.
[0441] Clause 6E. The device pursuant to any one of Clauses 1E-5E, wherein one or more processors are further configured to: generate first error correction data based on first coded video data; and transmit the first coded video data and the first error correction data to the receiving device.
[0442] Clause 7E. The device pursuant to any one of Clauses 1E-6E, wherein one or more processors are further configured to: encode a third frame without using the encoding selection data for the third frame, based on the determination that no encoding selection data for the third frame has been received from the receiving device before the time limit expires.
[0443] Clause 8E. The apparatus according to any one of Clauses 1E-7E, wherein one or more processors are further configured to: receive encoding selection data for a third frame of video data, wherein the encoding selection data for the third frame indicates encoding selection for encoding an estimate of the third frame; encode the third frame based on the encoding selection data for the third frame to generate third coded video data; apply a channel coding process for generating error correction data for the third coded video data; and transmit the error correction data of the third coded video data to a receiving device without transmitting at least a portion of the third coded video data.
[0444] Clause 9E. A device pursuant to any one of Clauses 1E-8E, wherein: the device is an extended reality (XR) headset and includes a display system, and one or more processors are further configured to: receive virtual element data from a receiving device; and the display system is configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0445] Clause 10E. An apparatus comprising: a memory configured to store video data; and one or more processors implemented in circuitry and coupled to the memory, the one or more processors being configured to: receive first coded video data from a transmitting device; reconstruct a first frame of the video data based on the first coded video data; estimate a second frame of the video data based on the first frame, wherein the second frame is a frame that appears after the first frame in decoding order; generate encoding selection data for the estimated second frame, wherein the encoding selection data indicates encoding selections for encoding the estimated second frame; transmit the encoding selection data for the second frame to the transmitting device; receive second coded video data from the transmitting device; and reconstruct the second frame based on the second coded video data.
[0446] Clause 11E. The apparatus according to Clause 10E, wherein: one or more processors are configured as part of encoding the estimated second frame, to perform motion compensation based on motion parameters of blocks of the second frame to generate predicted blocks, and encoding selection data includes motion parameters of blocks of the second frame, and second encoded video data includes encoded video data based on predicted blocks.
[0447] Clause 12E. An apparatus according to any one of Clauses 10E-11E, wherein: one or more processors are configured as part of encoding a second frame, to perform intra-frame prediction based on intra-frame prediction parameters for blocks of the second frame to generate prediction blocks, encoding selection data including intra-frame prediction parameters for blocks of the second frame, and second encoded video data including encoded video data based on prediction blocks.
[0448] Clause 13E. The device pursuant to any one of Clauses 10E-12E, wherein: the second coded video data does not include coded selection data; and one or more processors are configured, as part of an application decoding process, to use the coded selection data to reconstruct a second image based on the second coded video data.
[0449] Clause 14E. The device according to any one of Clauses 10E-13E, wherein one or more processors are configured to entropy encode the encoding selection data for the second frame before transmitting the encoding selection data for the second frame.
[0450] Clause 15E. The apparatus according to any one of Clauses 10E-14E, wherein: one or more processors are further configured to: estimate a third frame of video data based on one or more of a first frame or a second frame; perform an encoding process to encode the estimated third frame to generate third encoded video data, wherein third encoding selection data indicates encoding selection for encoding the estimated third frame; send the third encoding selection data to a transmitting device; receive error correction data of the third frame from the transmitting device; apply the error correction process to generate error-corrected encoded video data of the third frame based on the error correction data of the third frame and the third encoded video data; and apply the error-corrected encoded video data of the third frame to reconstruct a decoding process of the third frame.
[0451] Clause 16E. The apparatus as described in Clause 15E, wherein: the error-correcting coded video data of the third frame does not include third coding selection data, and one or more processors are configured, as part of an application decoding process, to use the third coding selection data to reconstruct the third frame based on the error-correcting coded video data of the third frame.
[0452] Clause 17E. The apparatus according to any one of Clauses 10E-16E, wherein one or more processors are configured to: apply a channel coding process to coding selection data for a second frame to generate error correction data for the coding selection data for the second frame; and transmit the error correction data for the coding selection data for the second frame to a transmitting device.
[0453] Clause 18E. The device according to any one of Clauses 10E-17E, wherein the device includes a communication interface configured to modulate coded selection data at a lower modulation order compared to other data transmissions in the data link between the device and the transmitting device.
[0454] Clause 19E. A device pursuant to any one of Clauses 10E-18E, wherein: one or more processors are further configured to process a second set of images to generate virtual element data, and the transmitting device is an XR headset configured to display one or more virtual elements in an extended reality (XR) scene based on the virtual element data.
[0455] Clause 20E. A method for processing video data, the method comprising: encoding a first frame of the video data to generate first encoded video data; transmitting the first encoded video data to a receiving device; receiving from the receiving device encoding selection data for a second frame of the video data, wherein: the encoding selection data for the second frame indicates encoding selections for encoding an estimate of the second frame, and the second frame follows the first frame in a decoding order; encoding the second frame based on the encoding selection data for the second frame to generate second encoded video data; and transmitting the second encoded video data to the receiving device.
[0456] Clause 21E. The method according to Clause 20E, wherein: the encoding selection data received from the receiving device includes motion parameters of blocks of the second frame, encoding the second frame includes performing motion compensation based on the motion parameters of the blocks of the second frame to generate predicted blocks, and the second encoded video data includes encoded video data based on the predicted blocks.
[0457] Clause 22E. The method according to any one of Clauses 20E-21E, wherein: encoding the selected data received from the receiving device includes intra-frame prediction parameters for blocks of the second frame, encoding the second frame includes performing intra-frame prediction based on the intra-frame prediction parameters for blocks of the second frame to generate prediction blocks, and the second coded video data includes coded video data based on prediction blocks.
[0458] Clause 23E. The method according to any one of Clauses 20E to 22E, wherein the second encoded video data does not include encoding selection data.
[0459] Clause 24E. The method according to any one of Clauses 20E-23E further includes entropy decoding of the encoded selection data used for the second screen.
[0460] Clause 25E. The method according to any one of Clauses 20E to 24E further includes: generating first error correction data based on first coded video data; and transmitting the first coded video data and the first error correction data to a receiving device.
[0461] Clause 26E. The method according to any one of Clauses 20E-25E further includes: encoding the third screen without using the encoding selection data for the third screen, based on the determination that no encoding selection data for the third screen has been received from the receiving device before the time limit expires.
[0462] Clause 27E. The method according to any one of Clauses 20E-26E further includes: receiving coding selection data for a third frame of video data, wherein the coding selection data for the third frame indicates coding selection for encoding an estimate of the third frame; encoding the third frame based on the coding selection data for the third frame to generate third coded video data; applying a channel coding process to generate error correction data for the third coded video data; and transmitting the error correction data of the third coded video data to a receiving device without transmitting at least a portion of the third coded video data.
[0463] Clause 28E. The method according to any one of Clauses 20E-27E, wherein: the device is an extended reality (XR) headset and includes a display system, and the method further includes: receiving virtual element data from a receiving device; and displaying one or more virtual elements in an XR scene on the display system based on the virtual element data.
[0464] Clause 29E. A method for processing video data, the method comprising: receiving first coded video data from a transmitting device; reconstructing a first frame of the video data based on the first coded video data; estimating a second frame of the video data based on the first frame, wherein the second frame is a frame that appears after the first frame in decoding order; generating encoding selection data for the estimated second frame, wherein the encoding selection data indicates encoding selections for encoding the estimated second frame; transmitting the encoding selection data for the second frame to the transmitting device; receiving second coded video data from the transmitting device; and reconstructing the second frame based on the second coded video data.
[0465] Clause 30E. The method according to Clause 29E, wherein: encoding the estimated second frame includes performing motion compensation based on motion parameters of blocks in the second frame to generate predicted blocks, and encoding selection data includes motion parameters of blocks in the second frame, and second encoded video data includes encoded video data based on predicted blocks.
[0466] Clause 31E. The method according to any one of Clauses 29E-30E, wherein: the second frame is subjected to intra-frame prediction to generate a prediction block by performing intra-frame prediction based on intra-frame prediction parameters for the blocks of the second frame, the encoded selection data includes the intra-frame prediction parameters for the blocks of the second frame, and the second encoded video data includes encoded video data based on the prediction blocks.
[0467] Clause 32E. The method according to any one of Clauses 29E to 31E, wherein: the second coded video data does not include coded selection data; and the application of the decoding process includes using the coded selection data to reconstruct a second image based on the second coded video data.
[0468] Clause 33E. The method according to any one of Clauses 29E-32E, wherein entropy encoding of the encoding selection data for the second screen occurs before the encoding selection data for the second screen is transmitted.
[0469] Clause 34E. The method according to any one of Clauses 29E-33E further includes: estimating a third frame of video data based on one or more of a first frame or a second frame; performing an encoding process to encode the estimated third frame to generate third encoded video data, wherein third encoding selection data indicates encoding selection for encoding the estimated third frame; sending the third encoding selection data to a transmitting device; receiving error correction data of the third frame from the transmitting device; applying an error correction process to generate error-corrected encoded video data of the third frame based on the error correction data of the third frame and the third encoded video data; and a decoding process to reconstruct the third frame using the error-corrected encoded video data of the third frame.
[0470] Clause 35E. The method according to any one of Clause 34E, wherein: the error-corrected video data of the third frame does not include third coding selection data, and the application of the decoding process includes using the third coding selection data to reconstruct the third frame based on the error-corrected video data of the third frame.
[0471] Clause 36E. The method according to any one of Clauses 29E-35E further includes: applying a channel coding process to coding selection data for a second screen to generate error correction data for the coding selection data for the second screen; and transmitting the error correction data for the coding selection data for the second screen to a transmitting device.
[0472] Clause 37E. The method according to any one of Clauses 29E-36E further includes: modulating the coded selection data at a lower modulation order compared to other data transmissions in the data link between the device and the transmitting device.
[0473] Clause 38E. The method according to any one of Clauses 29E-37E, wherein: it further comprises processing a second set of images to generate virtual element data, and the transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
[0474] Clause 39E. An apparatus comprising: means for encoding a first frame of video data to generate first encoded video data; means for transmitting the first encoded video data to a receiving device; means for receiving from the receiving device encoding selection data for a second frame of the video data, wherein: the encoding selection data for the second frame indicates encoding selections for encoding an estimate of the second frame, and the second frame is in decoding order after the first frame; means for encoding the second frame based on the encoding selection data for the second frame to generate second encoded video data; and means for transmitting the second encoded video data to the receiving device.
[0475] Clause 40E. An apparatus comprising: means for receiving first coded video data from a transmitting device; means for reconstructing a first frame of the video data based on the first coded video data; means for estimating a second frame of the video data based on the first frame, wherein the second frame is a frame that appears after the first frame in a decoding order; means for generating encoding selection data for the second frame, wherein the encoding selection data for the second frame indicates encoding selections for encoding the second frame; means for transmitting the encoding selection data for the second frame to the transmitting device; means for receiving second coded video data from the transmitting device; and means for reconstructing the second frame based on the second coded video data.
[0476] It will be appreciated that, depending on the example, certain actions or events of any of the techniques described herein can be performed in a different order, and can be added, combined, or omitted entirely (e.g., the practice does not require all the described actions or events). Furthermore, in some examples, actions or events can be performed simultaneously rather than sequentially, for example, through multithreading, interrupt handling, or multiple processors.
[0477] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code, and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium, or a communication medium that includes any medium facilitating, for example, the transmission of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium accessible by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0478] By way of example and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer specifically to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. The combinations described above should also be included within the scope of computer-readable media.
[0479] Instructions can be executed by one or more processors (e.g., programmable processors), such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.
[0480] The techniques disclosed herein can be implemented in a wide vari...
Claims
1. An apparatus comprising: The memory is configured to store video data; as well as One or more processors are implemented in a circuit and coupled to the memory, the one or more processors being configured to: The first frame of the video data is encoded to generate first encoded video data; Send the first encoded video data to the receiving device; The receiving device receives encoding selection data for a second frame of the video data, wherein: The encoding selection data for the second frame indicates the encoding selection used to encode the estimate of the second frame, and The second screen appears after the first screen in the decoding order. The second frame is encoded based on the encoding selection data used for the second frame to generate second encoded video data; and The second encoded video data is sent to the receiving device.
2. The device according to claim 1, wherein: The encoded selection data received for the second frame includes motion parameters of the blocks in the second frame. The one or more processors are configured as part of encoding the second frame, to perform motion compensation based on the motion parameters of blocks in the second frame to generate predicted blocks, and The second encoded video data includes encoded video data based on the prediction block.
3. The device according to claim 1, wherein: The encoding selection data used for the second frame includes intra-frame prediction parameters for the blocks of the second frame. The one or more processors are configured as part of encoding the second frame, to perform intra-frame prediction based on intra-frame prediction parameters of blocks in the second frame to generate prediction blocks, and The second encoded video data includes encoded video data based on the prediction block.
4. The device according to claim 1, wherein, The second encoded video data does not include the encoding selection data.
5. The device according to claim 1, wherein, The one or more processors are configured to perform entropy decoding on the encoded selection data used for the second screen.
6. The device according to claim 1, wherein, The one or more processors are further configured to: First error correction data is generated based on the first encoded video data; and The first encoded video data and the first error correction data are sent to the receiving device.
7. The device according to claim 1, wherein, The one or more processors are further configured to: Based on the determination that no encoding selection data for the third screen was received from the receiving device before the time limit expired, the third screen was encoded without using the encoding selection data for the third screen.
8. The device according to claim 1, wherein, The one or more processors are further configured to: Receive encoding selection data for a third frame of the video data, wherein the encoding selection data for the third frame indicates encoding selection for encoding an estimate of the third frame; The third frame is encoded based on the encoding selection data used for the third frame to generate third encoded video data; The application generates the channel coding process for the error correction data of the third encoded video data; and Error correction data of the third encoded video data is sent to the receiving device without sending at least a portion of the third encoded video data.
9. The device according to claim 1, wherein: The device is an extended reality (XR) headset and includes a display system, and The one or more processors are further configured to: Receive virtual element data from the receiving device; and The display system is configured to display one or more virtual elements in an XR scene based on the virtual element data.
10. An apparatus comprising: The memory is configured to store video data; as well as One or more processors are implemented in a circuit and coupled to the memory, the one or more processors being configured to: Receive the first encoded video data from the transmitting device; The first frame of the video data is reconstructed based on the first encoded video data; The second frame of the video data is estimated based on the first frame, and the second frame is a frame that appears after the first frame in the decoding order; Generate encoding selection data for the second screen, wherein the encoding selection data for the second screen indicates encoding selection for encoding the second screen; Send encoding selection data for the second screen to the sending device; Receive second encoded video data from the transmitting device; and The second frame is reconstructed based on the second encoded video data.
11. The device according to claim 10, wherein: The one or more processors are configured as part of encoding the second frame, to perform motion compensation based on the motion parameters of blocks in the second frame to generate predicted blocks, and The encoding selection data used for the second frame includes motion parameters of the blocks in the second frame. The second encoded video data includes encoded video data based on the prediction block.
12. The device according to claim 10, wherein: The one or more processors are configured as part of encoding the second frame to perform intra-frame prediction based on intra-frame prediction parameters for blocks of the second frame to generate prediction blocks. The encoding selection data for the second frame includes intra-frame prediction parameters for the blocks of the second frame, and The second encoded video data includes encoded video data based on the prediction block.
13. The device according to claim 10, wherein: The second encoded video data does not include encoding selection data for the second frame; and The one or more processors are configured to reconstruct the second frame based on the second encoded video data using encoding selection data for the second frame.
14. The device according to claim 10, wherein, The one or more processors are configured to entropy encode the encoding selection data for the second screen before sending the encoding selection data for the second screen.
15. The apparatus according to claim 10, wherein: The one or more processors are further configured to: The third frame of the video data is estimated based on one or more of the first frame or the second frame; An encoding process is performed to encode the estimated third frame to generate third encoded video data, wherein third encoding selection data indicates encoding selection for encoding the estimated third frame; Send the third encoding selection data to the transmitting device; Receive error correction data of the third screen from the transmitting device; The error correction process is applied to generate error-corrected coded video data for the third frame based on the error correction data and the third coded video data; and The decoding process of the third scene is reconstructed using error-corrected coded video data.
16. The device according to claim 15, wherein: The error-correction encoded video data of the third frame does not include the third encoding selection data, and The one or more processors are configured, as part of the decoding process, to use the third encoding selection data to reconstruct the third frame based on the error-corrected encoded video data of the third frame.
17. The device according to claim 10, wherein, The one or more processors are configured to: The channel coding process is applied to the coding selection data used for the second frame to generate error correction data for the coding selection data used for the second frame; and Error correction data for encoding selection data of the second screen is sent to the transmitting device.
18. The device according to claim 10, wherein, The device includes a communication interface configured to modulate the encoded selection data for the second frame at a lower modulation order compared to other data transmissions in the data link between the device and the transmitting device.
19. The apparatus according to claim 10, wherein: The one or more processors are further configured to process a second set of frames to generate virtual element data, and The transmitting device is an extended reality (XR) headset configured to display one or more virtual elements in an XR scene based on the virtual element data.
20. A method for processing video data, the method comprising: The first frame of the video data is encoded to generate first encoded video data; Send the first encoded video data to the receiving device; The receiving device receives encoding selection data for a second frame of the video data, wherein: The encoding selection data for the second frame indicates the encoding selection used to encode the estimate of the second frame, and The second screen appears after the first screen in the decoding order. The second frame is encoded based on the encoding selection data used for the second frame to generate second encoded video data; and The second encoded video data is sent to the receiving device.
21. The method of claim 20, wherein: The encoding selection data used for the second frame includes motion parameters of the blocks in the second frame. Encoding the second frame includes performing motion compensation based on the motion parameters of blocks in the second frame to generate predicted blocks, and The second encoded video data includes encoded video data based on the prediction block.
22. The method of claim 20, wherein: The encoding selection data used for the second frame includes intra-frame prediction parameters for the blocks of the second frame. Encoding the second frame includes performing intra-frame prediction based on the intra-frame prediction parameters of the blocks in the second frame to generate prediction blocks, and The second encoded video data includes encoded video data based on the prediction block.
23. The method of claim 20, further comprising: First error correction data is generated based on the first encoded video data; as well as The first encoded video data and the first error correction data are sent to the receiving device.
24. The method of claim 20, further comprising: Receive encoding selection data for a third frame of the video data, wherein the encoding selection data for the third frame indicates encoding selection for encoding an estimate of the third frame; The third frame is encoded based on the encoding selection data used for the third frame to generate third encoded video data; The channel coding process for generating error correction data for the third encoded video data; and Error correction data of the third encoded video data is sent to the receiving device without sending at least a portion of the third encoded video data.
25. A method for processing video data, the method comprising: Receive the first encoded video data from the transmitting device; The first frame of the video data is reconstructed based on the first encoded video data; The second frame of the video data is estimated based on the first frame, and the second frame is a frame that appears after the first frame in the decoding order; Generate encoding selection data for the second screen, wherein the encoding selection data for the second screen indicates encoding selection for encoding the second screen; Send encoding selection data for the second screen to the sending device; Receive second encoded video data from the transmitting device; and The second frame is reconstructed based on the second encoded video data.
26. The method of claim 25, wherein: Encoding the second frame includes performing motion compensation based on the motion parameters of blocks in the second frame to generate predicted blocks, and The encoding selection data used for the second frame includes motion parameters of the blocks in the second frame. The second encoded video data includes encoded video data based on the prediction block.
27. The method of claim 25, wherein: Encoding the second frame includes performing intra-frame prediction based on the intra-frame prediction parameters of the blocks in the second frame to generate prediction blocks. The encoding selection data used for the second frame includes intra-frame prediction parameters for the blocks of the second frame, and The second encoded video data includes encoded video data based on the prediction block.
28. The method according to claim 25, wherein: The second encoded video data does not include encoding selection data for the second frame; and The second frame is reconstructed based on the second encoded video data using the encoding selection data for the second frame.
29. The method of claim 25, further comprising: The third frame of the video data is estimated based on one or more of the first frame or the second frame; An encoding process is performed to encode the estimated third frame to generate third encoded video data, wherein third encoding selection data indicates encoding selection for encoding the estimated third frame; Send the third encoding selection data to the transmitting device; Receive error correction data of the third screen from the transmitting device; The error correction process is applied to generate error-corrected coded video data for the third frame based on the error correction data and the third coded video data; and The decoding process of the third scene is reconstructed using error-corrected coded video data.
30. The method of claim 25, further comprising: The channel coding process is applied to the coding selection data used for the second frame to generate error correction data for the coding selection data used for the second frame; as well as Error correction data for encoding selection data of the second screen is sent to the transmitting device.
Citation Information
Patent Citations
Analog modulated video transmission with variable symbol rate
US11431962B2
Interlaced coefficients in hybrid digital-analog modulation for transmission of video data
US11457224B2
Hybrid digital-analog modulation for transmission of video data
US11553184B2