Adaptive loop filtering for color format support
By introducing the ALF chroma flag and slice header signaling into the video bitstream, the problem of ineffective utilization of chroma components in existing video coding and decoding standards is solved, flexible ALF filtering of multiple color components is achieved, and the video output quality is improved.
Patent Information
- Application Number
- CN202180028173.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-14
- Filing Date
- 2021-04-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-04-15
AI Technical Summary
Existing video codec standards lack flexibility when processing video data with similar characteristics in different color channels, especially for video data in RGB and 4:4:4 formats. As a result, the chrominance components cannot be effectively filtered using the adaptive loop filter (ALF), affecting video output performance.
The ALF chroma flag is introduced into the video bitstream, and the chroma ALF filter is controlled through slice header data signaling, allowing flexible ALF filtering processing of multiple color components, including luminance and chroma components.
The operational flexibility and output performance of video codec equipment are improved, especially for video formats with similar characteristics of chrominance and luminance, and the applicability and efficiency of ALF filtering are enhanced.
Smart Images

Figure CN115398920B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] This application relates to video coding. For example, aspects of the application relate to systems, apparatuses, methods, and computer readable media for improving loop filters, such as adaptive loop filters (ALFs) (referred to as “systems and techniques”). In some examples, the systems and techniques can code (e.g., encode and / or decode) video data in different color formats (e.g., a 4:4:4 color format, a 4:2:0 color format, and / or other color formats). BACKGROUND
[0002] Many devices and systems allow video data to be processed and output for consumption. Digital video data includes a large amount of data to meet the needs of consumers and video providers. For example, consumers of video data desire the highest quality video with high fidelity, resolution, frame rate, and the like. As a result, the large amount of video data needed to meet these demands places a burden on communication networks and devices that process and store the video data.
[0003] Various video coding techniques can be used to compress video data. Video coding is performed in accordance with one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Motion Picture Expert Group (MPEG) coding, VP9, Alliance for Open Media (AOMedia) Video 1 (AV1), and the like. Video coding often employs prediction methods (e.g., inter-prediction, intra-prediction, and the like) that take advantage of redundancy present in video pictures or sequences. An important goal of video coding techniques is to compress video data into a form that uses a lower bit rate, while avoiding or minimizing degradations to video quality. As ever-evolving video services become available, coding techniques with better coding efficiency are needed. SUMMARY
[0004] Systems and techniques for coding (e.g., encoding and / or decoding) pictures and / or video content are described. In certain video coding standards (e.g., the Versatile Video Coding (VVC) standard), an adaptive loop filter (ALF) filters luma components with an adaptive filter bank and a classifier such that filters pre-stored in memory or signaled (e.g., via an adaptation parameter set (APS)) are identified and used for ALF filtering. Chroma components of a bitstream are filtered with a single filter (e.g., a 5x5 filter) whose coefficients are signaled once per APS. However, the operation of such video coding standards lacks flexibility and can not be efficient for videos with similar characteristics for different color channels (e.g., video data having a red, green, blue (RGB) format (where each pixel has a red component, a green component, and a blue component), video data having luma and chroma components per pixel, such as video data in a 4:4:4 format, or other video data). In such formats where color components have similar characteristics, color component data that is not subject to ALF filtering in the aforementioned video coding standards can benefit from ALF filtering. However, syntax elements of some video coding standards signal an ALF chroma identifier in an APS (e.g., using one or more APS syntax elements). Using APS syntax elements to control ALF filtering on chroma components lacks flexibility.
[0005] Examples described herein add a flag associated with ALF filtering for multiple color components (e.g., an added ALF chroma flag in addition to an existing ALF luma flag) to APS signaling. In some cases, an ALF chroma identifier for indicating chroma ALF filtering is included in slice header data (e.g., using a syntax structure described below) rather than being included in ALF data (e.g., signaled using an alf data syntax structure as described herein). Using slice header data for chroma ALF filter signaling can improve operation of a video coding device (e.g., an encoding device, a decoding device, or a combined encoding-decoding device) to improve flexibility of ALF filtering. Using slice header data for chroma ALF filter signaling can also improve video output performance (e.g., for chroma data in video formats having similar characteristics for chroma and luma). In some examples, ALF flag signaling can be used to apply ALF filtering and identify an ALF map for ALF filtering for a particular color component of video data (e.g., a block of video data carrying certain color component specific data) (e.g., in a slice of a picture of the video data). For example, a decoder receiving a bitstream including video data having a 4:4:4 format can identify a presence of an ALF flag (e.g., an ALF chroma filter signal flag) in the bitstream. The ALF flag can indicate that chroma ALF filtering is available for a slice of a picture of the video data. Additional information in the bitstream can indicate an ALF map (e.g., slice_alf_chroma_map_signalled, slice_alf_chroma2_map_signalled, etc.) that provides information used in ALF filtering for at least a portion of the slice.
[0006] According to one illustrative example, an apparatus for decoding video data is provided. The apparatus includes a memory and at least one processor (e.g., configured in a circuit) coupled to the memory. The at least one processor is configured to obtain a video bitstream, the video bitstream including adaptive loop filter (ALF) data, determine a value of an ALF chroma filter signal flag from the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in the video bitstream, and process at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0007] According to another illustrative example, a method of decoding video data is provided. The method includes obtaining a video bitstream including adaptive loop filter (ALF) data, determining, from the ALF data, a value of an ALF chroma filter signal flag that indicates whether chroma ALF filter data is signaled in the video bitstream, and processing at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0008] In another example, a non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to obtain a video bitstream including adaptive loop filter (ALF) data, determine, from the ALF data, a value of an ALF chroma filter signal flag that indicates whether chroma ALF filter data is signaled in the video bitstream, and process at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0009] In another example, an apparatus for decoding video data is provided. The apparatus includes means for obtaining a video bitstream including adaptive loop filter (ALF) data, means for determining, from the ALF data, a value of an ALF chroma filter signal flag that indicates whether chroma ALF filter data is signaled in the video bitstream, and means for processing at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0010] Some aspects of the method, apparatus, and computer readable medium described above include obtaining a slice header of a slice of video data from a video bitstream, obtaining a slice header of a slice of video data from a video bitstream, determining, from the slice header, a value of an ALF chroma identifier that indicates whether an ALF can be applied to one or more chroma components of the slice, and processing at least a portion of the slice of the video data based on the ALF chroma identifier from the slice header.
[0011] Some aspects of the method, apparatus, and computer readable medium described above include determining, from the slice header, a value of a chroma format identifier, the value of the chroma format identifier and the value of the ALF chroma identifier indicating which chroma component of the one or more chroma components the ALF is capable of being applied to.
[0012] In some aspects, a value of an ALF chroma filter signal flag indicates that the chroma ALF filter data is signaled in the video bitstream. In some aspects, the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least a portion of a slice of video data.
[0013] Some aspects of the method, apparatus, and computer-readable medium described above include obtaining, based on the value of the ALF chroma filter signal flag, chroma ALF filter data to be used for processing at least a portion of a slice of video data; and applying the chroma ALF filter data to at least a portion of a slice of the video data.
[0014] Some aspects of the method, apparatus, and computer-readable medium described above include inferring a value of the ALF chroma filter signal flag to be zero when the value of the ALF chroma filter signal flag is not present in the ALF data.
[0015] Some aspects of the method, apparatus, and computer-readable medium described above include obtaining, based on the value of the ALF chroma filter signal flag, luma ALF filter data to be used for one or more chroma components of at least one block of a video bitstream; and applying the luma ALF filter data to the one or more chroma components of the at least one block of the video bitstream.
[0016] Some aspects of the method, apparatus, and computer-readable medium described above include obtaining a slice header of a slice of video data from a video bitstream; determining a value of a chroma format identifier from the slice header; and processing one or more chroma components of at least one block of the video bitstream using luma ALF filter data based on the value of the chroma format identifier from the slice header.
[0017] Some aspects of the method, apparatus, and computer-readable medium described above include processing a value of an ALF chroma filter signal flag from ALF data to determine that chroma ALF filter data is signaled in a video bitstream.
[0018] Some aspects of the method, apparatus, and computer-readable medium described above include determining an ALF application parameter set (APS) identifier for a first color component of at least a portion of a slice; and determining an ALF map for the first color component of at least a portion of the slice.
[0019] Some aspects of the method, apparatus, and computer-readable medium described above include enabling ALF filtering for at least two non-luma components of at least a portion of a slice based on a component of the at least a portion of the slice that includes a shared characteristic.
[0020] In some aspects, the at least two non-luma components of the at least portion of the slice include a red component, a green component, and a blue component of the at least portion of the slice.
[0021] In some aspects, the at least two non-luma components of the at least portion of the slice include a chroma component of the at least portion of the slice.
[0022] In some aspects, the at least portion of the slice includes video data in a 4:4:4 format.
[0023] Some aspects of the method, apparatus, and computer-readable medium described above include enabling ALF filtering for at least two non-luma components of an at least portion of a slice based on the at least portion of the slice including non-4:2:0 format video data.
[0024] Some aspects of the method, apparatus, and computer-readable medium described above include determining a chroma type array variable for an at least portion of a slice, determining an ALF chroma application parameter set (APS) identifier for a first component of the at least portion of the slice based on the chroma type array variable for the at least portion of the slice, and determining a signaled ALF map for the first component of the at least portion of the slice.
[0025] Some aspects of the method, apparatus, and computer-readable medium described above include determining a second signaled ALF map for a second component of the at least portion of the slice based on the chroma type array variable.
[0026] Some aspects of the method, apparatus, and computer-readable medium described above include performing ALF filtering for the first component and the second component of the at least portion of the slice using the signaled ALF map and the second signaled ALF map.
[0027] Some aspects of the method, apparatus, and computer-readable medium described above include determining a third signaled ALF map for a third component of the at least portion of the slice based on the chroma type array variable.
[0028] In some aspects, the first component is a luma component, where the second component is a first chroma component, and where the third component is a second chroma component.
[0029] In some aspects, the first component is a red component, where the second component is a green component, and where the third component is a blue component.
[0030] Some aspects of the method, apparatus, and computer-readable medium described above include performing ALF processing for each component of the at least portion of the slice based on the chroma type array variable.
[0031] According to another illustrative example, an apparatus for encoding video data is provided. The apparatus includes a memory and at least one processor (e.g., configured in a circuit) coupled to the memory. The at least one processor is configured to: generate adaptive loop filter (ALF) data; determine a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generate the video bitstream including the ALF data.
[0032] According to another illustrative example, a method of encoding video data is provided. The method includes: generating adaptive loop filter (ALF) data; determining a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generating the video bitstream including the ALF data.
[0033] In another example, a non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, cause the one or more processors to: generate adaptive loop filter (ALF) data; determine a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generate the video bitstream including the ALF data.
[0034] In another example, an apparatus for encoding video data is provided. The apparatus includes: means for generating adaptive loop filter (ALF) data; means for determining a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and means for generating the video bitstream including the ALF data.
[0035] Some aspects of the above-described method, apparatus, and computer readable medium for encoding video data include determining a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether ALF can be applied to one or more chroma components of a slice of the video data; and including the value of the ALF chroma identifier in a slice header of the video bitstream.
[0036] Some aspects of the above-described methods, apparatuses, and computer readable medium for encoding video data include determining a value for a chroma format identifier, the value for the chroma format identifier and a value for an ALF chroma identifier indicating which chroma component of one or more chroma components an ALF can be applied to; and including the value for the chroma format identifier in a slice header of the video bitstream.
[0037] In some aspects, a value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream. In some aspects, the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least a portion of a slice.
[0038] Some aspects of the above-described methods, apparatuses, and computer readable medium for encoding video data include determining a value for a chroma format identifier, wherein the value for the chroma format identifier indicates one or more chroma components of at least one block of the video bitstream to be processed using luma ALF filter data; and including the value for the chroma format identifier in a slice header of the video bitstream.
[0039] In some aspects, the apparatus includes a mobile device (e.g., a mobile telephone or so-called “smart phone,” a tablet or other type of mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a television, a vehicle (or a computing device of a vehicle), or other device. In some aspects, the apparatus includes at least one camera for capturing one or more pictures or video frames. For example, the apparatus can include a camera (e.g., an RGB camera) or multiple cameras for capturing one or more pictures and / or one or more videos including video frames. In some aspects, the apparatus includes a display for displaying one or more pictures, videos, notifications, or other displayable data. In some aspects, the apparatus includes a transmitter configured to transmit one or more video frames and / or syntax data to at least one device over a transmission medium. In some aspects, the processor includes a neural processing unit (NPU), a central processing unit (CPU), a graphics processing unit (GPU), or other processing device or component.
[0040] This Summary is intended to identify some, but not all, of the claimed subject matter. The subject matter should be understood from the entire description of the patent, with appropriate
[0041] The foregoing will be more fully understood from the following description, claims, and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0042] An illustrative embodiment of the present application is described in detail below with reference to the following drawings:
[0043] Figure 1A is a block diagram illustrating an example of an encoding device and a decoding device according to some examples;
[0044] Figure 1B is a diagram illustrating an example implementation of a picture divided into tiles and slices according to the techniques of this disclosure;
[0045] Figure 1C is a diagram illustrating an example implementation of a filter unit for performing color component based ALF according to the techniques of this disclosure;
[0046] Figure 2A is a conceptual diagram illustrating an example of adaptive loop filter (ALF) filter support including a 5x5 diamond according to some examples;
[0047] Figure 2B is a conceptual diagram illustrating another example of ALF filter support including a 7x7 diamond according to some examples;
[0048] Figure 3A is a conceptual diagram illustrating aspects of subsampled positions for vertical gradients according to some examples;
[0049] Figure 3B is a conceptual diagram illustrating aspects of subsampled positions for horizontal gradients according to some examples;
[0050] Figure 3C is a conceptual diagram illustrating aspects of subsampled positions for diagonal gradients according to some examples;
[0051] Figure 3D is a conceptual diagram illustrating aspects of subsampled positions for diagonal gradients according to some examples;
[0052] Figure 4 is a flow diagram illustrating aspects of ALF support for different color formats according to some examples;
[0053] Figure 5 is a flow diagram illustrating an example of a process for decoding video data according to some examples;
[0054] Figure 6 is a flow diagram illustrating an example of a process for decoding video data according to some examples;
[0055] Figure 7 is a block diagram illustrating an example video encoding device according to some examples; and
[0056] Figure 8 FIG. 1 is a block diagram illustrating an example video decoding device, in accordance with some examples. DETAILED DESCRIPTION
[0057] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some of them can be applied in combination, as will be apparent to one of ordinary skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of embodiments of the application. It will be apparent, however, that various embodiments can be practiced without
[0058] The ensuing description provides exemplary embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the ensuing description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing an exemplary embodiment. It will be apparent to those skilled in the art that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.
[0059] Video coding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques can include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques to reduce or remove redundancies inherent in video sequences. A video encoder can partition each picture of an original video sequence into rectangular regions called video blocks or coding units. Blocks can include coding tree blocks (CTBs), prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise indicated, a reference to a “block” generally can refer to such a video block (e.g., a CTB, coding block, prediction block, transform block, or other suitable block or sub-block, as will be understood by one of skill in the art). Moreover, each of these blocks can also be interchangeably referred to herein as a “unit” (e.g., a coding tree unit (CTU), coding unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit can indicate a coding logical unit encoded in a bitstream, while a block can indicate a portion of a video frame buffer for which processing is directed. In some standards, a coding tree block (CTB) constitutes a CTU and is structured to carry individual color components of video data. For example, a CTU can include a first CTB for a luma component of the CTU, a second CTB for a chroma-blue (Cb) component of the CTU, and a third CTB for a chroma-red (Cr) component of the CTU.
[0060] A video block can be encoded using a particular prediction mode. For inter-prediction modes, a video encoder can search for a block that is similar to the block being encoded in a frame (or picture) located at another temporal position, referred to as a reference frame or reference picture. The video encoder can limit the search to a certain spatial displacement from the block to be encoded. The best match can be located using a two-dimensional (2D) motion vector that includes a horizontal displacement component and a vertical displacement component. For intra-prediction modes, the video encoder can use a spatial prediction technique to form a predicted block based on data from neighboring blocks previously encoded within the same picture.
[0061] A video encoder can determine a prediction error. For example, the prediction error can be determined as a difference between the pixels values in the block being encoded and the predicted block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error using a transform coding (e.g., in the form of a discrete cosine transform (DCT), a discrete sine transform (DST), or other suitable transform) to produce transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and the motion vectors can be represented using syntax elements and, along with control information, form a coded representation of the video sequence. In some examples, the video encoder can entropy code the syntax elements, further reducing the number of bits needed to represent them.
[0062] A video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., a predicted block) for decoding a current frame. For example, the video decoder can add the predicted block and the compressed prediction error. The video decoder can determine the compressed prediction error by weighting transform basis functions using the quantized coefficients. The difference between the reconstructed frame and the original frame is referred to as a reconstruction error.
[0063] In some cases, one or more adaptive loop filters (ALFs) can be applied individually to color components in video data to improve the quality of the output video. For example, an ALF can be applied to a picture or a block of a picture after the picture or block has been reconstructed using inter prediction or intra prediction. In some cases, ALF filtering can be used to correct or repair artifacts introduced during reconstruction of the picture or block.
[0064] In some video formats, certain color components include additional data compared to other color components. For example, video data having a 4:2:0 format includes luma data having a higher resolution than associated chroma data. In such video formats, ALF filtering can only be applied to the higher resolution luma data. However, in other formats, such as red green blue (RGB) data and 4:4:4 format data (where all color components have the same sampling rate), the different color components have similar characteristics. Some video coding standards (e.g., the EVC standard) are structured such that ALF filtering is applied to the luma component, and non-luma components (e.g., chroma components) do not receive ALF filtering. In such cases, output performance can be improved by applying ALF to more than one color component (e.g., in addition to being able to apply ALF filtering to the luma component, ALF filtering is also applied to one or more chroma components). Aspects described herein provide additional flexibility and efficient signaling for formats where ALF filtering of multiple color components results in an improvement in the output image.
[0065] For example, aspects described herein can include applying ALF filtering to multiple color components of a video bitstream to improve performance. In one example, an RGB format video bitstream can include a picture that is divided into slices that include multiple CTBs. Each slice can include separate CTBs for a red color component, a green color component, and a blue color component. In another example, a video bitstream that includes a luminance and chrominance components (e.g., in a luminance (Y)-chrominance blue (Cb)-chrominance red (Cr) format, referred to as a YCbCr format) can include a picture that is divided into slices. Each slice can include a CTB for a luminance component and CTBs for two chrominance components (e.g., a CTB for a Cb component and a CTB for a Cr component). While some video coding standards emphasize ALF filtering of a single color component (e.g., a luminance CTB), examples described herein provide slice header based signaling to flexibly allow ALF filtering of additional color components of video data (e.g., ALF filtering of either or both of the chrominance CTBs in addition to the luminance CTB).
[0066] In some examples, an ALF chroma filter signal flag (e.g., in the alf data syntax structure) is added to the ALF data in the video bitstream (e.g., in the parameter set, in the header data such as a slice header, etc.). In one illustrative example, the ALF chroma filter signal flag can include an alf chroma filter signal flag syntax element in the alf data syntax structure. The ALF chroma filter signal flag, operating together with the ALF luma filter signal flag, can indicate that chroma filter data is signaled, or is not signaled. In some cases, the ALF chroma filter signal flag can be used with a slice ALF chroma identifier (e.g., signaled as a slice alf chroma idc syntax element) that is signaled in the slice header, rather than in the ALF data (e.g., rather than in the alf data syntax structure) to indicate ALF filtering of additional color components (e.g., non-luma components, such as chroma components). For example, the ALF chroma filter signal flag (e.g., an alf chroma filter signal flag syntax element in the alf data syntax structure) can indicate that ALF is available for one or more chroma components, and the slice ALF chroma identifier (e.g., a slice alf chroma idc in the slice header) can have values (e.g., values from 0 to 3) that each indicate a different ALF chroma option. In one example, a first value can indicate that ALF filtering is to be applied to a first chroma component, a second value can indicate that ALF filtering is to be applied to a second chroma component, a third value can indicate that ALF filtering is to be applied to both the first chroma component and the second chroma component, and a fourth value can indicate that ALF filtering is not to be applied to either the first chroma component or the second chroma component.
[0067] The techniques described herein can be applied to any existing video codec (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codec), MPEG5 Efficient Video Coding (EVC) (e.g., implemented in ETM5.0), Versatile Video Coding (VVC), Joint Exploration Model (JEM), VP9, AV1), and / or can be an effective coding tool for any video coding standard that is being developed and / or future video coding standards.
[0068] Figure 1A1 is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device (also referred to as a client device). The source device and / or the receiving device may include an electronic device, such as a mobile or landline telephone handset (e.g., a smartphone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, a server device in a server system (e.g., a video streaming server system or other suitable server system) including one or more server devices, a head-mounted display (HMD), a head-up display (HUD), smart glasses (e.g., virtual reality (VR) glasses, augmented reality (AR) glasses, or other smart glasses), or any other suitable electronic device.
[0069] The components of system 100 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0070] Although system 100 is shown as including certain components, one of ordinary skill in the art will appreciate that system 100 may include more than Figure 1A For example, in some cases, system 100 may further include one or more memory devices other than storage 108 and storage 118 (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) in communication with and / or electrically connected to the one or more memory devices, one or more wireless interfaces for performing wireless communications (e.g., including one or more transceivers and a baseband processor for each wireless interface), one or more wired interfaces for performing communications over one or more hardwired connections (e.g., a serial interface such as a Universal Serial Bus (USB) input, a lighting connector, and / or other wired interfaces), and / or Figure 1A Other components not shown.
[0071] The coding and decoding techniques described herein can be applied to video coding and decoding in various multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcasting or transmission, encoding of digital video for storage on a data storage medium, decoding of digital video for storage on a data storage medium, or other applications. In some examples, system 100 can support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.
[0072] The encoding device 104 (or encoder) can be used to encode video data using a video codec standard or protocol to generate an encoded video bitstream. Examples of video codec standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including its Scalable Video Codec (SVC) and Multi-View Video Codec (MVC) extensions, and High Efficiency Video Codec (HEVC) or ITU-T H.265. Various extensions to HEVC exist to handle multi-layer video codecs, including range and screen content codec extensions, 3D Video Codec (3D-HEVC) and Multi-View Extension (MV-HEVC), and Scalable Extension (SHVC). HEVC and its extensions are developed by the Joint Collaboration Team on Video Coding and Decoding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG), and the Joint Collaboration Team on 3D Video Codec Extensions (JCT-3V).
[0073] MPEG and ITU-T VCEG also formed the Joint Exploratory Video Team (JVET) to explore and develop new video codec tools for the next generation of video codec standards, called Versatile Video Codec (VVC). The reference software is called the VVC Test Model (VTM). The goal of VVC is to provide significant improvements in compression performance over the existing HEVC standard, helping to deploy higher quality video services and emerging applications (e.g., 360° immersive multimedia, high dynamic range (HDR) video, etc.). VP9 and Alliance for Open Media (AOMedia) Video 1 (AV1) are other video codec standards to which the techniques described in this article can be applied.
[0074] Many of the embodiments described herein can be performed using a video codec such as MPEG5 EVC, VVC, HEVC, AVC, and / or extensions thereof. However, the techniques and systems described herein can also be applicable to other video coding standards, such as MPEG4 or other MPEG standards, Joint Photographic Experts Group (JPEG) (or other still picture coding standards), VP9, AV1, extensions thereof, or other suitable coding standards that are already available or not yet available or developed. Thus, although the techniques and systems described herein can be described with reference to a particular video coding standard, one of ordinary skill in the art will appreciate that the description should not be interpreted as applying only to that particular standard.
[0075] Reference Figure 1A The video source 102 can provide the video data to the encoding device 104. The video source 102 can be part of the source device or can be part of a device other than the source device. The video source 102 can include a video capture device (e.g., a video camera, a still camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider that provides video data, a video feed interface that receives video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.
[0076] The video data from the video source 102 can include one or more input pictures. A picture can also be referred to as a “frame.” A picture or frame is a still image, which in some cases is part of a video. In some examples, the data from the video source 102 can be a still image that is not part of a video. In some video coding specifications, a video sequence can include a series of pictures. A picture can include three sample arrays, denoted as S L , S Cb , and S Cr . S L is a two-dimensional array of luma samples, S Cb is a two-dimensional array of Cb chroma samples, and S Cr is a two-dimensional array of Cr chroma samples. Chroma samples can also be referred to herein as “chrominance” samples. In other cases, a picture can be monochrome and can include only an array of luma samples. A pixel can refer to a point in a picture that includes luma and chroma samples. For example, a given pixel can contain a luma sample from the S L array, a Cb chroma sample value from the S Cb array, and a Cr chroma sample value from the S Cr array.
[0077] The encoder engine 106 (or encoder) of the encoding device 104 encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more coded video sequences. A coded video sequence (CVS) includes a series of access units (AUs) starting from an AU that has a random access point picture in the base layer and has certain properties until and excluding the next AU that has a random access point picture in the base layer and has certain properties. For example, certain properties of the random access point picture that starts the CVS can include a random access skipped leading (RASL) picture flag (e.g., NoRaslOutputFlag) equal to 1. Otherwise, a random access point picture (RASL flag equal to 0) does not start a CVS. An access unit (AU) includes one or more coded pictures and control information corresponding to the coded pictures that share the same output time. Coded slices of a picture are encapsulated at the bitstream level into data units called network abstraction layer (NAL) units. For example, in some video standards, a video bitstream can include one or more CVSs that include NAL units. Each of the NAL units has a NAL unit header. In one example, for H.264 / AVC, the header is one byte (except for multi-layer extensions), and for HEVC, the header is two bytes. Syntax elements in the NAL unit header take up specified bits and are thus visible to all types of systems and transport layers, such as transport streams, real-time transport (RTP) protocols, file formats, etc.
[0078] There are two types of NAL units in some video standards, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units include coded picture data that forms the encoded video bitstream. For example, the bit sequence that forms the encoded video bitstream is present in VCL NAL units. A VCL NAL unit can include one slice or slice segment (described below) of coded picture data, and a non-VCL NAL unit includes control information related to one or more coded pictures. In some cases, a NAL unit can be referred to as a packet. An HEVC AU includes VCL NAL units containing coded picture data and non-VCL NAL units corresponding to the coded picture data, if any. Among other information, a non-VCL NAL unit can contain parameter sets that have high-level information related to the encoded video bitstream. For example, parameter sets can include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). In some cases, each slice or other portion of a bitstream can reference a single active PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.
[0079] A NAL unit can include a bit sequence forming an encoded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as an encoded representation of a picture in a video. In some cases, the encoder engine 106 can generate an encoded representation of a picture by partitioning each picture into a plurality of slices. A slice is independent of other slices such that information in the slice is encoded without relying on data from other slices within the same picture. A slice includes one or more slice segments, including an independent slice segment and one or more dependent slice segments that depend on a previous slice segment, if any.
[0080] In some examples, the encoder engine 106 can partition each picture into subpictures, slices, and tiles, such as described in the VVC standard. Figure 1B is a diagram from the VVC standard showing an example of slice and tile partitioning of a picture 121. As shown, the picture 121 is partitioned into one or more tile rows and one or more tile columns. A tile can be defined as a series of CTUs covering a rectangular region of a picture. In some cases, the CTUs in a tile are scanned in a raster scan order in the tile. A slice can include an integer number of complete tiles or an integer number of consecutive complete rows of CTUs within a tile of a picture. For example, each vertical slice boundary can also be a vertical tile boundary. As described below, each CTU can include a plurality of coding tree blocks (CTBs). A subpicture can include one or more slices that collectively cover a rectangular region of a picture. For example, each subpicture boundary can also be a slice boundary, and each vertical subpicture boundary can also be a vertical tile boundary. In some cases, all CTUs in a subpicture belong to the same tile. In some cases, all CTUs in a tile belong to the same subpicture.
[0081] In some video standards, a slice is partitioned into CTBs of luma samples and chroma samples. A CTU of luma samples and one or more CTBs of chroma samples, along with syntax of the samples, is referred to as a coding tree unit (CTU). As described herein, video data can be structured into CTBs of different color components. ALF filtering of different color components can be applied to CTBs of each color component. A CTU can also be referred to as a “treeblock” or a “largest coding unit” (LCU). In some standards, a CTU is the basic processing unit for encoding. A CTU can be divided into multiple coding units (CUs) of different sizes. A CU contains arrays of luma and chroma samples called coding blocks (CBs).
[0082] A luma and chroma CB can be further divided into prediction blocks (PBs). A PB is a block of samples of a luma component or a chroma component that is inter- predicted or intra-block copy predicted using the same motion parameters. A luma PB and one or more chroma PBs, together with associated syntax, form a prediction unit (PU). For inter prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and used for inter prediction of the luma PB and the one or more chroma PBs. The motion parameters can also be referred to as motion information. A CB can also be partitioned into one or more transform blocks (TBs). A TB represents a square block of color component samples on which a residual prediction error signal is coded applying a residual transform, e.g., the same two-dimensional transform in some cases. A transform unit (TU) represents a TB of luma and chroma samples and corresponding syntax elements. Transform coding will be described in more detail below.
[0083] The size of a CU corresponds to the size of the coding mode and can be square in shape. For example, the size of the CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other appropriate size up to the size of the corresponding CTU. The phrase "NxN" is used herein to refer to pixel dimensions of a video block in both the vertical and horizontal dimensions (e.g., 8 pixels by 8 pixels). The pixels in a block can be arranged in rows and columns. In some embodiments, a block can not have the same number of pixels in the horizontal direction as in the vertical direction as a block. Syntax data associated with a CU can describe, for example, partitioning of the CU into one or more PUs. The partitioning mode can differ between whether the CU is intra or inter prediction mode encoded. PUs can be partitioned into non-square shapes. Syntax data associated with a CU can also describe partitioning of the CU into one or more TUs, for example, according to a CTU. The shape of a TU can be square or non-square.
[0084] According to some video coding standards, transform units (TUs) can be used to perform a transform. The TUs can differ for different CUs. The size of a TU can be determined based on the size of the PUs within a given CU. The TUs can be as large as the PUs, or can be smaller than the PUs. In some examples, residual samples corresponding to a CU can be subdivided into smaller units using a quad-tree structure referred to as a residual quad-tree (RQT). Leaf nodes of the RQT can correspond to TUs. Pixel difference values associated with a TU can be transformed to produce transform coefficients. The transform coefficients can be quantized by the encoder engine 106.
[0085] Once the pictures of the video data are partitioned into CUs, the encoder engine 106 predicts each PU using a prediction mode. The prediction units or blocks are subtracted from the original video data to obtain residuals (as described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. The prediction mode can include intra prediction (or intra-picture prediction) or inter prediction (or inter-picture prediction). Intra prediction exploits the correlation between spatially neighboring samples in a picture. For example, using intra prediction, each PU is predicted from neighboring image data in the same picture, e.g., using DC prediction to find an average value for the PU, planar prediction to fit a planar surface to the PU, directional prediction to extrapolate from neighboring data, or any other suitable type of prediction. Inter prediction uses the temporal correlation between pictures to derive a motion compensated prediction of a block of image samples. For example, using inter prediction, each PU is predicted from image data in one or more reference pictures (that precede or follow the current picture in output order) using motion compensated prediction. The decision whether to use inter- or intra-picture prediction to code a picture region can be made, for example, at the CU level.
[0086] The encoder engine 106 and the decoder engine 116 (described in more detail below) can be configured to operate according to a given video coding standard, such as EVC. According to some video coding standards, a video coder, such as the encoder engine 106 and / or the decoder engine 116, partitions a picture into a plurality of coding tree units (CTUs) (where a CTB of luma samples and one or more CTBs of chroma samples, along with syntax for the samples, are referred to as a CTU). The video coder can partition a CTU according to a tree structure, such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concepts of multiple split types, such as the separation between CUs, PUs, and TUs of some standards. The QTBT structure includes two levels, including a first level of partitioning according to quadtree partitioning and a second level of partitioning according to binary tree partitioning. A root node of the QTBT structure corresponds to a CTU. Leaf nodes of the binary trees correspond to coding units (CUs).
[0087] In the MTT partitioning structure, a block can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. Ternary tree partitioning is a partitioning that splits one block into three sub-blocks. In some examples, a ternary tree partitioning divides a block into three sub-blocks without dividing the original block through a center. The partitioning types (e.g., quadtree, binary tree, and ternary tree) in the MTT can be symmetric or asymmetric.
[0088] In some examples, a video coder can use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video coder can use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luma component and another QTBT or MTT structure for the two chroma components (or two QTBT and / or MTT structures for the respective chroma components).
[0089] A video coder can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures in accordance with certain standards. For purposes of illustration, the description herein can refer to QTBT partitioning. However, it should be understood that the techniques of this disclosure can also apply to video coders configured to use quadtree partitioning or other types of partitioning.
[0090] In some examples, one or more slices of a picture are assigned a slice type. Slice types include intra coded slices (I slices), inter coded P slices, and inter coded B slices. An I slice (intra coded frame, independently decodable) is a slice of a picture that is coded only with intra prediction, and thus is independently decodable, as the I slice only requires intra data to predict any prediction units or prediction blocks of the slice. A P slice (uni-directional predicted frame) is a slice of a picture that can be coded with intra prediction and uni-directional inter prediction. Each prediction unit or prediction block within a P slice is coded with either intra prediction or inter prediction. When inter prediction is applied, the prediction unit or prediction block is predicted from only one reference picture, and thus the reference samples come from only one reference region of a frame. A B slice (bi-directional predicted frame) is a slice of a picture that can be coded with intra prediction and inter prediction (e.g., bi-directional prediction or uni-directional prediction). Prediction units or prediction blocks of a B slice can be bi-directionally predicted from two reference pictures, where each picture contributes one reference region, and the sample sets of the two reference regions are weighted (e.g., with equal or different weights) to produce the prediction signal for the bi-directionally predicted block. As noted above, slices of a picture are independently coded. In some cases, a picture can be coded as only one slice.
[0091] As noted above, intra-picture prediction exploits the correlation between spatially neighboring samples within a picture. There are multiple intra-prediction modes (also referred to as “intra modes”). In some examples, intra-prediction of a luma block includes 35 modes, including a planar mode, a DC mode, and 33 angular modes (e.g., diagonal intra-prediction modes and angular modes adjacent to the diagonal intra-prediction modes). The 35 modes of intra-prediction are indexed as shown in Table 1 below. In other examples, more intra modes can be defined, including prediction angles that can not be represented by the 33 angular modes. In other examples, the prediction angles associated with the angular modes can be different than the prediction angles used in some standards.
[0092] Intra prediction mode Associated name 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34
[0093] Table 1: Specification of Intra prediction modes and associated names
[0094] Inter-picture prediction uses temporal correlation between pictures in order to derive a motion-compensated prediction of a block of picture samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Ax, Ay), with Ax specifying the horizontal displacement and Ay specifying the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector (Ax, Ay) can be of integer sample precision (also referred to as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of the reference frame. In certain cases, the motion vector (Ax, Ay) can have fractional sample precision (also referred to as fractional pixel precision or non-integer precision) in order to more accurately capture the motion of underlying objects without being limited to the integer pixel grid of the reference frame. The precision of a motion vector can be indicated by the quantization level of the motion vector. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., ¼ pixel, ½ pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture in order to derive the prediction signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values at fractional positions. The previously decoded reference picture is indicated by a reference index (refIdx) to a reference picture list. The motion vector and the reference index can be referred to as motion parameters. Two types of inter-picture prediction can be performed, including uni-directional prediction and bi-directional prediction.
[0095] For inter prediction using bi-directional prediction, two sets of motion parameters (Ax0, Ay0, refldx0 and Ax1, Ay1, refldx1) are used to generate two motion-compensated predictions (either from the same reference picture or possibly from different reference pictures). For example, for bi-directional prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are combined to arrive at the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference pictures available for bi-directional prediction are stored in two separate lists, denoted as List 0 and List 1. The motion parameters can be derived at the encoder using a motion estimation process.
[0096] For inter prediction using uni -prediction, one set of motion parameters (Ax0, Ay0, refIdx0) is used to generate a motion-compensated prediction from a reference picture. For example, for uni -prediction, each prediction block uses at most one motion-compensated prediction signal, and P prediction units are generated.
[0097] A PU can include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when a PU is encoded using intra prediction, the PU can include data describing an intra prediction mode for the PU. As another example, when a PU is encoded using inter prediction, the PU can include data defining a motion vector for the PU. The data defining the motion vector for the PU can describe, for example, a horizontal component of the motion vector (Ax), a vertical component of the motion vector (Ay), a resolution of the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), a reference picture to which the motion vector points, a reference index, a reference picture list (e.g., list 0, list 1, or list C) for the motion vector, or any combination thereof.
[0098] Following the performance of prediction using intra and / or inter prediction, the encoding device 104 can perform a transform and quantization. For example, following prediction, the encoder engine 106 can calculate residual values corresponding to a PU. The residual values can include pixel difference values between a current block of pixels (PU) being encoded and a prediction block (e.g., a predicted version of the current block) used to predict the current block. For example, following the generation of the prediction block (e.g., using inter prediction or intra prediction), the encoder engine 106 can produce a residual block by subtracting the prediction block produced by the prediction unit from the current block. The residual block includes a set of pixel difference values quantifying the differences between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.
[0099] Any residual data that can remain following the performance of prediction is transformed using a block transform, which can be based on a discrete cosine transform (DCT), a discrete sine transform (DST), an integer transform, a wavelet transform, other suitable transform function, or any combination thereof. In some cases, one or more block transforms (e.g., cores of size 32x32, 16x16, 8x8, 4x4, or other suitable size) can be applied to the residual data in each CU. In some examples, TUs can be used for the transform and quantization processes implemented by the encoder engine 106. A given CU having one or more PUs can also include one or more TUs. As described in further detail below, residual values can be transformed into transform coefficients using a block transform, and TUs can be used to quantize and scan the residual values to produce serialized transform coefficients for entropy coding.
[0100] In some embodiments, after the CU's PUs are used for intra- or inter- prediction encoding, the encoder engine 106 can calculate residual data for the TUs of the CU. The PUs can include pixel data in the spatial domain (or pixel domain). As previously noted, the residual data can correspond to pixel differences between pixels of the unencoded picture and prediction values corresponding to the PUs. The encoder engine 106 can form one or more TUs containing the residual data for the CU (which includes the PUs) and can transform the TUs to produce transform coefficients for the CU. The TUs can include coefficients in the transform domain after application of a block transform.
[0101] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, during quantization, a coefficient having an n-bit value can be rounded down to an m-bit value, where n is greater than m.
[0102] Once quantization is performed, the encoded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), partitioning information, and any other suitable data, such as other syntax data. The different elements of the encoded video bitstream can be entropy encoded by the encoder engine 106. In some examples, the encoder engine 106 can scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy encoded. In some examples, the encoder engine 106 can perform an adaptive scan. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 can entropy encode the vector. For example, the encoder engine 106 can use context adaptive variable length coding, context adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval partitioning entropy coding, or another suitable entropy encoding technique.
[0103] The output 110 of the encoding device 104 can transmit the NAL units that make up the encoded video bitstream data to a decoding device 112 of a receiving device over a communication link 120. The input 114 of the decoding device 112 can receive the NAL units. The communication link 120 can include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network can include any wireless interface or combination of wireless interfaces and can include any suitable wireless network (e.g., the Internet or other wide area networks, local area networks, personal area networks, campus area networks, wireless networks, wireless personal area networks, or any other such network based on any combination of IEEE 802.11 principles and architectures, Bluetooth® or other suitable short or long range wireless technologies). The wired network can include any wired interface or combination of wired interfaces and can include any suitable wired network (e.g., based on copper wire, fiber optic cable, etc.). The communication link 120 can be any combination of wireless and wired links. For example, the communication link 120 can be a network adapted to carry video data, such as a cellular network, the Internet, or a Wi-Fi network. TM , radio frequency (RF), ultra-wideband (UWB), WiFi-Direct, cellular, long term evolution (LTE), WiMax TMWired networks can include any wired interfaces (e.g., fiber, Ethernet, powerline Ethernet, coaxial cable Ethernet, digital signal line (DSL), etc.). Wired and / or wireless networks can be implemented using various devices such as base stations, routers, access points, bridges, gateways, switches, etc. Modulated coded video bitstream data can be transmitted to receiving devices according to a communication standard such as a wireless communication protocol.
[0104] In some examples, the encoding device 104 can store the encoded video bitstream data in a storage 108. The output 110 can retrieve the encoded video bitstream data from the encoder engine 106 or from the storage 108. The storage 108 can include any of a variety of distributed or locally accessed data storage media. For example, the storage 108 can include a hard drive, a storage disk, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. The storage 108 can also include a decoded picture buffer (DPB) for storing reference pictures used in inter prediction. In another example, the storage 108 can correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. In such cases, a receiving device including the decoding device 112 can access stored video data from the storage device via streaming or download. The file server can be any type of server capable of storing encoded video data and transmitting that encoded video data to a receiving device. Example file servers include web servers (e.g., for a website), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on a file server. The transmission of encoded video data from the storage 108 can be a streaming transmission, a download transmission, or a combination thereof.
[0105] Input 114 of the decoding device 112 receives encoded video bitstream data and can provide the video bitstream data to a decoder engine 116 or to a storage 118 for later use by the decoder engine 116. For example, the storage 118 can include a DPB for storing reference pictures used in inter prediction. The receiving device including the decoding device 112 can receive encoded video data to be decoded via the storage 108. The encoded video data can be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the receiving device. The communication medium used to transmit the encoded video data can include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from a source device to a receiving device.
[0106] The decoder engine 116 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more encoded video sequences that make up the encoded video data. The decoder engine 116 can rescale and perform inverse transforms on the encoded video bitstream data. The residual data is passed to a prediction stage of the decoder engine 116. The decoder engine 116 predicts blocks of pixels (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (residual data).
[0107] The video decoding device 112 can output the decoded video to a video destination device 119, which can include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, the video destination device 119 can be part of the receiving device that includes the decoding device 112. In some aspects, the video destination device 119 can be part of a separate device from the receiving device.
[0108] In some embodiments, the video encoding device 104 and / or the video decoding device 112 can be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 can also include other hardware or software that is necessary for implementing the above-described coding techniques, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof. The video encoding device 104 and the video decoding device 112 can be integrated as part of a combined encoder / decoder (CODEC) in the respective device.
[0109] Figure 1AThe example system shown is one illustrative example that can be used herein. Techniques for processing video data using the techniques described herein can be performed by any digital video encoding and / or decoding device. Although generally the techniques of this disclosure are performed by a video encoding device or a video decoding device, the techniques can also be performed by a combined video encoder-decoder, typically referred to as a "CODEC." Additionally, the techniques of this disclosure can also be performed by a video preprocessor. Source device and receive device are merely examples of such CODEC devices, where the source device generates coded video data for transmission to the receive device. In some examples, the source device and receive device can operate in a substantially symmetrical manner such that each device includes video encoding and decoding components. Hence, the example system can support one-way or two-way video transmission between video devices, e.g., for video streaming, video playback, video broadcasting, or video telephony.
[0110] As previously noted, some video coding standards define a bitstream that includes a set of NAL units, including VCL NAL units and non-VCL NAL units. The VCL NAL units include coded picture data that forms a coded video bitstream. For example, the bit sequence that forms a coded video bitstream is present in VCL NAL units. The non-VCL NAL units can contain parameter sets with high-level information related to the coded video bitstream, among other information. For example, the parameter sets can include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). Examples of the goals of parameter sets include bitrate efficiency, error resilience, and providing a system layer interface. Each slice references a single active PPS, SPS, and VPS to access information that the decoding device 112 can use to decode the slice. A parameter set identifier (ID) can be coded for each parameter set, including a VPS ID, an SPS ID, and a PPS ID. The SPS includes the SPS ID and the VPS ID. The PPS includes the PPS ID and the SPS ID. Each slice header includes the PPS ID. Using the IDs, the active parameter sets for a given slice can be identified.
[0111] A PPS includes information that applies to all slices in a given picture. As such, all slices in one picture refer to the same PPS. Slices in different pictures can also refer to the same PPS. An SPS includes information that applies to all pictures in the same coded video sequence (CVS) or bitstream. As previously mentioned, a coded video sequence is a series of access units (AUs) that start with a random access point picture (e.g., an instantaneous decoding reference (IDR) picture or a broken link access (BLA) picture, or other appropriate random access point picture) in the base layer and have certain properties (as described above) up to and not including the next AU in the base layer that has a random access point picture and has certain properties (or the end of the bitstream). The information in an SPS can not change from picture to picture in a coded video sequence. The same SPS can be used for pictures in a coded video sequence. A VPS includes information that applies to all layers in a coded video sequence or bitstream. A VPS includes a syntax structure with syntax elements that apply to the entire coded video sequence. In some embodiments, a VPS, SPS, or PPS can be sent in-band with the encoded bitstream. In some embodiments, a VPS, SPS, or PPS can be sent out-of-band in a separate transmission from the NAL units containing the coded video data.
[0112] Various chroma formats can be used for video. A chroma format syntax element can be used to specify chroma samples. For example, syntax element chroma format idc specifies chroma samples relative to luma samples, such as in the VVC and / or EVC standards. In some cases, the value of chroma format idc shall be in the range of 0 to 2, inclusive.
[0113] Depending on the value of chroma format idc, the values of variables SubWidthC and SubHeightC are assigned as specified in VVC clause 6.2, and variable ChromaArrayType is assigned. For example, the values of variables SubWidthC and SubHeightC can be assigned as follows:
[0114]
[0115] Table 2 SubWidthC and SubHeightC values derived from chroma format idc and separate colour plane flag
[0116] In some examples, variable ChromaArrayType is assigned as follows:
[0117] If chroma format idc is equal to 0, ChromaArrayType is set equal to 0.
[0118] Otherwise, ChromaArrayType is set equal to chroma format idc.
[0119] The variables SubWidthC and SubHeightC are specified in Table 1 below, depending on the chroma format sampling structure specified by chroma format idc. ISO / IEC can specify other values for chroma format idc, SubWidthC and SubHeightC in the future.
[0120] chroma_format_idc Chroma format SubWidthC SubHeightC 0 Monochrome 1 1 1 4:2:0 2 2 2 4:2:2 2 1 3 4:4:4 1 1
[0121] Table 1 SubWidthC and SubHeightC values derived from chroma format idc
[0122] In monochrome sampling, there is only one sample array, which is nominally considered to be the luma array. In 4:2:0 sampling, each of the two chroma arrays (e.g., for Cb and Cr) has half the height and half the width of the luma array. In 4:2:2 sampling, each of the two chroma arrays has the same height as the luma array and half the width. In 4:4:4 sampling, each of the two chroma arrays has the same height and width as the luma array.
[0123] In the field of video coding, filters can be applied to improve the quality of decoded or reconstructed video signals. In some cases, a filter can be applied as a post filter, where the filtered frame is not used for prediction of future frames. In some cases, a filter can be applied as an in-loop filter, where the filtered frame is used for prediction of one or more future frames. For example, an in-loop filter can filter a picture after performing reconstruction on the picture (e.g., after adding a residual to a prediction) and before the picture is output and / or before the picture is stored in a picture buffer (e.g., a decoded picture buffer). For example, a filter can be designed by minimizing an error between an original signal and a decoded filtered signal. Examples of filters include a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter.
[0124] Figure 1C An example implementation of a filter unit 122 that can be used to process video pictures or blocks using ALF filtering in accordance with examples herein is shown. In some cases, the filter unit 122 can be implemented as Figure 7 the filter unit 63 of FIG. 1 and / orFigure 8 Filter units 63 and 91 may, for example, in some cases perform the techniques of this disclosure in conjunction with other components of video encoding device 104 or video decoding device 112. In some examples, filter units 63 and 91 may be post-processing units that may perform the techniques of this disclosure, for example, outside of video encoding device 104 and video decoding device 112 (e.g., after the decoded video is output from video decoding device 112).
[0125] exist Figure 1C In the example of FIG, the filter unit 122 includes a deblocking filter 124, a sample adaptive offset (SAO) filter 126, and an adaptive loop filter (ALF) / geometric transform-based adaptive loop filter (GALF) filter 128. The SAO filter 126 can be configured, for example, to determine offset values for samples of a block. The deblocking filter 124 can be used to compensate for the use of block structure units during the encoding and decoding process. The ALF filter 128 can be used to minimize the error (e.g., mean square error) between the original samples and the decoded samples by using an adaptive filter, which can be a Wiener-based adaptive filter or other suitable adaptive filter. The ALF filter 128 can be configured to determine parameters for filtering the current block based on signaled parameters for filtering the color components of the current block, for example. As further described herein, the signaled parameters can be based on the sampling format, which can improve device performance by providing ALF filtering for multiple color components in some video signals (e.g., RGB or 4:4:4 format luma-chroma-chroma format video data).
[0126] In some examples, after the deblocking filter 124, the ALF 128 can be applied (e.g., with block-based filter adaptation). In some cases, for the luma component of the block, one of 25 filters can be selected through a classification process for each block (e.g., for each 4×4 block or other sized block) based on local statistical estimates such as gradient and directionality. To benefit from the symmetric properties of the filter, the ALF utilized can employ a filter coefficient transformation process. More details on the ALF design are provided below. Figure 7 and Figure 8 The loop filter is further generally described.
[0127] The filter unit 122 may include Figure 1C Fewer filters may be shown and / or additional filters and / or other components may be included. Figure 1CThe particular filters shown can be implemented in different orders. Other loop filters (either in the coding loop or after the coding loop) can also be used to smooth pixel transitions or otherwise improve video quality. When in the coding loop, decoded video blocks in a given frame or picture can be stored in a decoded picture buffer (DPB). The DPB stores reference pictures for subsequent motion compensation (e.g., for inter-prediction). The DPB can be part of or separate from additional memory that stores decoded video for later presentation on a display device (such as, Figure 1A the display of video destination device 119.
[0128] As noted above, different sample formats for color components can result in different performance outcomes when applying filtering. Some video coding standards emphasize ALF filtering for luma components (in which case ALF filtering does not apply to non-luma components, such as chroma components), such as due to the prevalence of video data having a 4:2:0 format. However, other formats (such as 4:4:4 format data RGB data, etc.) can benefit from improved picture quality when ALF filtering is applied to more than one color component (e.g., two or three color components in an RGB format or luma-chroma format, such as YUV and YCbCr formats).
[0129] Various filter shapes can be used. Two diamond filter shapes are used in some implementations (as shown in Figure 2A and Figure 2B In certain examples, Figure 2A the 5x5 diamond in Figure 2B may be used to filter luma samples, and Figure 2A the 5x5 diamond in may be used for chroma samples.
[0130] In certain cases, block classification can be performed. For example, for the luma component of a pixel (luma samples), each 4x4 block can be classified as one of 25 classes. The classification index C is derived based on its directionality D and the amount of activity
[0131]
[0132] In some examples, to compute D and the 1-D Laplacian operator is first used to compute gradients for horizontal, vertical, and two diagonal directions:
[0133]
[0134]
[0135]
[0136]
[0137] where indices i and j refer to the coordinates of the top-left sample within the 4x4 block, and R(i,j) indicates the reconstructed sample at coordinate (i,j).
[0138] To reduce the complexity of block classification, a sub-sampled 1-D Laplacian computation is applied. As Figure 3A , Figure 3B , Figure 3C and show the sub-sampled Laplacian computation Figure 3D The same sub-sampling positions are used for gradient computation of all directions, including vertical gradient Figure 3A , horizontal gradient Figure 3B , and diagonal gradient Figure 3C and Figure 3D .
[0139] The maximum and minimum values of the gradients of the horizontal and vertical directions D are set as:
[0140]
[0141] The maximum and minimum values of the gradients of the two diagonal directions D are set as:
[0142]
[0143] To derive the value of the directionality D, these values are compared with each other and with two thresholds t1 and t2:
[0144] Step 1. If and are both true, then D is set to 0.
[0145] Step 2. If continue from Step 3; otherwise, continue from Step 4.
[0146] Step 3. If then D is set to 2; otherwise, D is set to 1.
[0147] Step 4. If then D is set to 4; otherwise, D is set to 3.
[0148] The activity value A is computed as follows:
[0149]
[0150] where the A value is further quantized to the range of 0 to 4 (inclusive), and the quantized value is denoted as
[0151] For chroma components in a picture, no classification method is applied (e.g., a single set of ALF coefficients is applied for each chroma component.
[0152] In some cases, geometric transformations of filter coefficients can be applied. For example, in some examples, before filtering each 4x4 luma block, depending on the gradient value computed for the block, a geometric transformation such as a rotation or a diagonal and vertical flip is applied to the filter coefficients f(k, l). In some cases, this is equivalent to applying these transformations to the samples in the filter support region. By aligning their directionality, the transformations can make different blocks for which ALF is applied more similar.
[0153] Three geometric transformations can be introduced, including diagonal, vertical flip, and rotation:
[0154] Diagonal: f D (k, l) = f(l, k),
[0155] Vertical flip: f V (k, l) = f(k, K - l - 1),
[0156] Rotation: f R (k, l) = f(K - l - 1, k),
[0157] where K is the size of the filter, and 0 < k, l < K - 1 are the coefficient coordinates, with position (0, 0) at the top-left corner and position (K - 1, K - 1) at the bottom-right corner. Depending on the gradient value computed for the block, a transformation is applied to the filter coefficients f(k, l). The following table summarizes the relationship between the transformations and the four gradients for the four directions.
[0158] Gradient value Transform <![CDATA[g d2 <g d1 And g h <g v ]]> No transform g d2 g d1 and g v g h ]]> Diagonal g d1 g d2 g h g v ]]> Vertical flip g d1 g d2 and g v g h ]]> Rotation
[0159] Table 4 Mapping of the gradient computed for a block to a transformation
[0160] In some aspects, an adaptation parameter set (APS) can be used to signal filter parameters or filter data (e.g., filter coefficients and / or other parameters) in a bitstream, such as in an APS NAL unit. An APS can have an associated type, such as an ALF type or a luma mapping and chroma scaling (LMCS) type (e.g., as defined in the VVC standard or other video coding standards). An APS can include a set of luma filter parameters, one or more sets of chroma filter parameters, or a combination thereof. An APS can be used in various video coding standards, such as VVC, EVC, and the like. In some cases, signaling of an APS can be limited. For example, a group of tiles (e.g., a group of one or more tiles, such as those shown in FIG. 3) can only signal an index of an APS for the group of tiles (e.g., in a group of tiles header). Figure 1B In some aspects, an adaptation parameter set (APS) can be used to signal filter parameters or filter data (e.g., filter coefficients and / or other parameters) in a bitstream, such as in an APS NAL unit. An APS can have an associated type, such as an ALF type or a luma mapping and chroma scaling (LMCS) type (e.g., as defined in the VVC standard or other video coding standards). An APS can include a set of luma filter parameters, one or more sets of chroma filter parameters, or a combination thereof. An APS can be used in various video coding standards, such as VVC, EVC, and the like. In some cases, signaling of an APS can be limited. For example, a group of tiles (e.g., a group of one or more tiles, such as those shown in FIG. 3) can only signal an index of an APS for the group of tiles (e.g., in a group of tiles header).
[0161] In some examples, each APS can be identified by a unique identifier (e.g., adaptation_parameter_set_id) that is used to reference the current APS information from other syntax elements. APSs can be shared across pictures and can be different for different parts of a picture (e.g., for different tile groups within a picture). When tile group alf enabled flag is equal to 1, the APS is referenced by the tile group header and the ALF parameters carried in the APS can be implemented out-of-band (e.g., provided by an external device outside of the video encoding device), providing benefits in some cases.
[0162] As described herein, aspects are described in which APS can be used to indicate (e.g., using one or more flags or other syntax) whether ALF filtering is available for certain color components (e.g., chroma components) in addition to existing signaling (e.g., flags) that indicate ALF filtering for other color components (e.g., luma component). Additional signaling can be implemented using slice headers to improve flexibility and available output quality associated with ALF filtering for additional color components (e.g., ALF filtering for chroma in addition to luma).
[0163] For example, described herein are examples in which filter applicability can be controlled at the block level (e.g., CTB level). In some examples, a first flag is signaled to indicate whether ALF is applied to the luma component of a block (e.g., luma CTB), and a second flag is signaled to indicate whether ALF is available for the chroma component of the block (e.g., one or more chroma CTBs, such as a Cb CTB and / or a Cr CTB). Such flags can be signaled, for example, in ALF data (e.g., an alf data syntax structure, as described below). In some aspects, for chroma CTB signaling, a flag can be signaled using an alf chroma ctb present flag syntax element to indicate whether ALF is available for application, as described in more detail below.
[0164] At the decoder side, when ALF is enabled for a block (e.g., CTB), each sample R(i,j) within the block (e.g., together with the CTB or within a coding block of the CTB, such as a CU) is filtered, resulting in a sample value R'(i,j) as shown below, where L denotes the filter length, f m,n denotes the filter coefficients, and f(k,l) denotes the decoded filter coefficients.
[0165]
[0166] In some cases, fixed filters can be used. For example, some ALF designs use a fixed filter set that is provided to the decoder as side information to initialize. In some cases, there are a total of 64 7×7 filters (e.g., each filter contains 13 coefficients). For each classification category, a mapping is applied to define which 16 fixed filters from the 64 filters can be used for the current category. The selection index (0-15) of each category is signaled as the fixed filter index. When using an adaptively derived filter, the difference between the fixed filter coefficients and the adaptive filter coefficients is signaled.
[0167] In some cases, temporal filters may be used. For example, to further benefit from the temporal correlation of the video data, the ALF design may exploit the reuse of ALF coefficients signaled earlier in APS NAL units. Each APS is identified by a unique adaptation_parameter_set_id that is used to reference the current APS information from other syntax elements (e.g., from the tile group header). In some cases, all signaled APSs with a unique set identifier value are stored in an APS buffer (e.g., with a size of up to 32 (entries)). To enable random access (RA) codec configurations, the encoder's choice of APS adaptation_parameter_set_id to use is constrained. For example, to maintain temporal scalability, only temporal filters from the same or lower temporal layer may be used.
[0168] An example of the adaptive loop filter data syntax is as follows:
[0169]
[0170]
[0171] As described above, ALF filtering in some video codec standards (e.g., MPEG5 EVC) includes filtering the luma component using an APS filter bank and a classifier. The adaptive filter bank and classifier may indicate filtering selected from pre-stored or signaled filters in the APS filter bank. In some such examples, the chroma component may be filtered using a single 5×5 filter, with coefficients signaled once per APS.
[0172] As described further above, such coding operations can become inefficient for coding of RGB format video, 4:4:4 chroma format video, or video having other formats in which color components (e.g., luma and chroma components) have similar characteristics. For example, in 4:4:4 chroma format, all three color components are present at full resolution. With this shared full resolution of all three color components, two non-luma components (e.g., chroma components) can benefit from more advanced filtering (e.g., applying luma-type ALF to all three components). As described herein, luma-type ALF filtering can refer to ALF filtering for higher resolution or more complex data (e.g., when compared to the smaller filter described above with each APS signaling coefficients once). When a chroma ALF flag is set to allow luma-type ALF filtering of chroma components, examples described herein can signal ALF filter data at slice header level, which can enable more flexible ALF filtering of chroma data.
[0173] Another issue present in ALF filter design is that some video coding standards utilize an alf chroma idc syntax element that is signaled in an APS. The alf chroma idc syntax element is used to control filtering of certain chroma component data (whether filtering is on or off for certain video data). In some such standards, alf chroma idc equal to 0 specifies that the chroma adaptive loop filter set is not signaled and not applied to Cb and Cr color components, alf chroma idc greater than 0 indicates that the chroma ALF set is signaled, alf chroma idc equal to 1 indicates that the chroma ALF set is applied to the Cb color component, alf chroma idc equal to 2 indicates that the chroma ALF set is applied to the Cr color component, and alf chroma idc equal to 3 indicates that the chroma ALF set is applied to both Cb and Cr color components. When such signaling is indicated in an APS without an option to change the settings for subsequent slices that share the APS signaling, there is no option to adjust the ALF filtering as needed to account for changes that occur at the slice level. Thus, such APS signaling decreases performance by preventing the use of appropriate chroma ALF filtering.
[0174] Systems, methods, and computer readable media are described for improving filtering (e.g., adaptive loop filter (ALF), deblocking, and / or other filtering) and enabling video data to be coded with different color formats (e.g., 4:4:4 color format, 4:2:0 color format, and / or other color formats). Examples described herein improve over the prior art (e.g., techniques based on video coding standards) by implementing additional flexibility in ALF filtering of additional color components (e.g., chroma components) to improve coding efficiency and / or output performance for some video coding devices and networks. Such improvements can be applied to any video coding standard, such as those described above that emphasize luma ALF filtering. As one possible implementation, the following changes to MPEG5 EVC are proposed, and examples are described below in the context of the EVC standard. It will be clear that similar implementations can be used with other standards having the above-described properties.
[0175] As part of the above-described improvements in flexibility, in some aspects, the signaling of chroma filters is controlled by a separate flag signaled in the APS. As shown below, the slice ALF chroma identifier (e.g., slice_alf_chroma_idc syntax element) is moved from the ALF data (e.g., alf_data syntax structure) to the slice header (e.g., slice_header() syntax structure) to allow more flexible signaling and more frequent changing of settings from the slice ALF chroma identifier below to match available chroma ALF filtering to the video data format (e.g., to match to a 4:4:4 format, etc.). Examples language is described below, where portions are modified to indicate aspects as described herein. Changes relative to the EVC standard are shown in “changes” below, and the full standard is shown in “full standard” below. <highlight>"and" <highlightend>" between symbols to indicate new language and strikeout to indicate old removed language (e.g., <highlight> <highlightend>"and" and <highlight> <highlightend> ”)。
[0176]
[0177] <highlight> <highlightend>
[0178]
[0179] <highlight> <highlightend>
[0180] As noted above for ALF data (alf data syntax structure) and slice header syntax structure, aspects described herein can use ALF data to signal both a flag for luma ALF and a flag for chroma ALF. As indicated by the above syntax, the alf chroma filter signal flag syntax element specifies whether chroma filter data is signaled in the bitstream (e.g., in an APS). In some cases, the filter data can include filter coefficients. The slice alf chroma idc syntax element (e.g., slice_header() syntax element) in the slice header data of the bitstream specifies whether ALF is applied to one or more chroma components (e.g., Cb and / or Cr color components) in the slice. In some examples, if the ALF chroma filter signal flag value (e.g., value of the alf chroma filter signal flag syntax element) does not include an explicit value in the bitstream, the value can be inferred to be zero. In such examples, the decoder can determine from the ALF data that the ALF chroma filter signal flag value is zero even when no explicit flag value is signaled in the bitstream. The associated slice ALF chroma identifier (e.g., slice alf chroma idc syntax element) can indicate when ALF is applied to chroma data, such as after the ALF chroma filter signal flag value (e.g., value of the alf chroma filter signal flag syntax element) is used to identify that chroma ALF is signaled in the bitstream. In described implementations, a non-zero value of slice alf chroma idc (e.g., alf chroma indication from slice header) indicates that ALF can be applied to chroma components of the video data. Particular non-zero values (e.g., 1, 2, 3, etc.) can be used to map which particular color components ALF can be applied to.
[0181] For example, the ChromaArrayType variable describes the format of the video signal (e.g., whether the signal has chroma at all, such as a monochrome format or a format with one or more chroma components). The slice_alf_chroma_idc syntax element specifies the application of ALF to the chroma components. slice_alf_chroma_idc=0 specifies that no ALF is applied to the chroma components. slice_alf_chroma_idc=0 can be signaled even from ChromaArrayType=2 (indicating the presence of chroma components in the video signal). However, if chromaArrayType==0 (indicating that the video signal is luma-only), then slice_alf_chroma_idc cannot be signaled to have a value greater than 0, which would indicate that chroma filtering is to be performed.
[0182] Figure 4 4 shows aspects of a process 400 for providing ALF support for different color formats for a decoding device (e.g., decoding device 112) according to some examples. Figure 4 As shown, operation 402 of process 400 involves a decoding device obtaining an encoded bitstream (also referred to as a bitstream or video bitstream). The encoded bitstream may include both ALF data (e.g., the alf_data() syntax structure above) and slice header data (e.g., the slice_header() syntax structure above) for the encoded video data of the encoded bitstream. As part of processing the encoded video data, in operation 404, the decoding device determines whether ALF is enabled for the encoded bitstream. If ALF is not enabled (e.g., not available for processing luma or chroma color components), then in operation 405, decoding continues without ALF filtering until a setting is changed to enable ALF for a portion of the encoded video data (e.g., for one or more slices of the video data).
[0183] If ALF is available, an ALF luma filter signal flag (e.g., the alf luma filter signal flag syntax element from the alf data( ) syntax structure above) and an ALF chroma filter signal flag (e.g., the alf chroma filter signal flag syntax element from the alf data( ) syntax structure above) can be included (e.g., added by the encoding device 104) in the ALF data as part of the encoded bitstream. As described herein, in certain cases, the values of these flags can be inferred. For example, in one implementation, the value of the ALF luma filter signal flag and / or the ALF chroma filter signal flag can be explicitly signaled to be 1, and when the value is not present in the bitstream, the value can be inferred to be 0. At operation 406, the decoding device can process the ALF data to determine whether the ALF chroma filter signal flag has a value of 0. If the value of the ALF chroma filter signal flag is not 0 (e.g., the value of the flag is 1 or some other value not equal to 1), then at operation 407, the decoding device can determine that the chroma filter data is signaled in the APS (e.g., rather than being indicated in the slice header data).
[0184] In Figure 4 In implementations where the ALF chroma filter flag is determined to be 0, the decoding device can determine whether ALF is available for application to any chroma components at operation 408 by determining whether a value of a slice ALF chroma indication (e.g., slice alf chroma idc or ChromaArrayType) is a non-zero value (e.g., a value of 1 or some other value not equal to 1). If the value of the slice ALF chroma indication (e.g., slice alf chroma idc or ChromaArrayType) is determined to not be a non-zero value (e.g., is determined to be a value of 0), then the decoding device can determine at operation 409 that ALF is not to be applied to chroma color components (e.g., associated with the slice data including the slice ALF chroma identifier) of a current slice of video data. Based on this determination, the decoding device cannot apply ALF to the chroma color components of the current slice. If the slice ALF chroma identifier is determined by the decoding device to be a non-zero value (e.g., a value of 1 or some other value), then the decoding device can determine at operation 410 that ALF is available (or can be applied) to one or more chroma color components (e.g., Cb component and / or Cr component) of a slice or a portion thereof (e.g., a block of the slice such as a CTU, CU, CTB, CB, etc.) of video data. The decoding device can then apply ALF to the one or more chroma components based on the particular value of the slice ALF chroma identifier, as discussed further below.
[0185] In some examples, the ALF parameter signaling and applicability is made dependent on a chroma format identifier (id) that indicates a chroma format of the video data, such as a 4:2:0 format, a 4:2:2 format, a 4:4:4 format, or other chroma formats. In some cases, a chroma type array variable (e.g., ChromaArrayType) can be used as the chroma format id to indicate the chroma format of the video data. In some examples, if the video is coded with ChromaArrayType equal to 3 (corresponding to a 4:4:4 format) or not equal to 4:2:0, ALF filtering is enabled for both non-luma components (e.g., Cb and Cr components) of the coded video data. In some examples, if the video data is coded using a chroma format other than 4:2:0 format (e.g., using a 4:4:4 format or a 4:2:2 format), each chroma component can reference a separate APS to access the best filter set. In some examples, if the video data is coded using a format other than 4:2:0 format, an ALF classifier is performed on the non-luma components to produce an index to a particular filter. In some examples, for each chroma component, a separate block-based applicability map (e.g., where block-level flags are signaled) is coded for video data having a format other than 4:2:0 format. In some examples, for video data having a 4:2:0 format and / or for video data having a 4:2:2 format, chroma filtering can be performed using a single filter without a classifier, and no map is signaled, in which case each block can be filtered conditionally.
[0186] An example of syntax structures and semantics that can be used as part of operation 410 to determine how to apply available chroma ALF is as follows (where changes relative to the EVC standard are shown as, <highlight>"and" <highlightend>" between symbols to indicate new language and strikeout to indicate old, removed language:
[0187]
[0188] <highlight>
[0189]
[0190] <highlightend>
[0191]
[0192] alf_ctb_flag[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] equal to 1 specifies that the adaptive loop filter is applied to the coding tree block of the luma component of the coding tree unit at luma location ( xCtb, yCtb ). alf_ctb_flag[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] equal to 0 specifies that the adaptive loop filter is not applied to the coding tree block of the luma of the coding tree unit at luma location ( xCtb, yCtb ).
[0193] When alf_ctb_flag[ xCtb » CtbLog2SizeY ][ yCtb » CtbLog2SizeY ] is not present, it is inferred to be equal to slice_alf_enabled_flag.
[0194] <highlight>
[0195] When slice_chroma_alf_enabled_flag is not present, it is inferred to be equal to 0.
[0196]
[0197] <highlightend>
[0198] For each coding tree unit with luma coding tree block position (rx, ry), where rx = 0... PicWidthlnCtbsY - 1 and ry = 0... PicHeightlnCtbsY - 1, the following applies:
[0199] - When alf_ctb_flag[rx][ry] is equal to 1, the specified coding tree block luma type filtering process is invoked with recPicture L , alfPicture L , the referenced APS identification slice_alf_luma_aps_id, and luma coding tree block position (xCtb, yCtb) set equal to (rx « CtbLog2SizeY, ry « CtbLog2SizeY) as inputs and the modified filtered picture alfPicture L as output.
[0200] - When ChromaArrayType is equal to 3, the coding tree block luma type filtering process for chroma samples of Cb and Cr color components is invoked as follows:
[0201] - When alf_ctb_chroma_flag[rx][ry] is equal to 1, the specified coding tree block luma type filtering process is invoked with recPictureCb, alfPictureCb, the referenced APS identification slice_alf_chroma_aps_id, and chroma coding tree block position (xCtb, yCtb) set equal to (rx « CtbLog2SizeY, ry « CtbLog2SizeY) as inputs and the modified filtered picture alfPictureCb as output.
[0202] - When alf_ctb_chroma2_flag[rx][ry] is equal to 1, the specified coding tree block luma type filtering process is invoked with recPictureCr, alfPictureCr, the referenced APS identification slice_alf_chroma2_aps_id, and chroma coding tree block position (xCtb, yCtb) set equal to (rx « CtbLog2SizeY, ry « CtbLog2SizeY) as inputs and the modified filtered picture alfPictureCr as output.
[0203] - Otherwise, when ChromaArrayType is in the range of 1 to 2, inclusive, and slice_alf_chroma_idc is greater than 0, the following applies:
[0204] - When sliceChromaAlfEnableFlag is equal to 1, the specified coding tree block chroma type filtering process is invoked with recPicture set equal to recPictureCb, alfPicture set equal to alfPictureCb, the referenced APS identified by slice_alf_chroma_aps_id, and chroma coding tree block position (xCtbC, yCtbC) set equal to ((rx « CtbLog2SizeY) / SubWidthC, (ry « CtbLog2SizeY) / SubHeightC) as inputs, and the output is the modified filtered picture alfPictureCb.
[0205] - When sliceChroma2AlfEnableFlag is equal to 1, the specified coding tree block chroma type filtering process is invoked with recPicture set equal to recPictureCr, alfPicture set equal to alfPictureCr, the referenced APS identified by slice_alf_chroma_aps_id, and chroma coding tree block position (xCtbC, yCtbC) set equal to ((rx « CtbLog2SizeY) / SubWidthC, (ry « CtbLog2SizeY) / SubHeightC) as inputs, and the output is the modified filtered picture alfPictureCr.
[0206] The above are examples of aspects implemented as modifications to EVC, but similar modifications can be made to aspects in other video coding standards. The above syntax structures (e.g., the if (sps_alf_flag) syntax structure) can be part of a slice header syntax structure (e.g., the slice_header() syntax structure) in an EVC encoded bitstream. Once a decoding device has determined that ALF is generally available (e.g., for luma and chroma components), and is signaled at the slice level (e.g., rather than in an APS as described as part of operation 407), the decoding device can process the above syntax structures to determine the particular chroma components for which ALF is applied.
[0207] For example, as shown in the example slice header syntax above, if slice alf enable flag is true (e.g., value 1), then when slice alf chroma idc (e.g., ALF chroma indication from slice header) is greater than 0, the decoding device can check the ChromaArrayType variable for values 1 or 2. If this statement is true (e.g., ChromaArrayType is 1 or 2), then an ALF application parameter set (APS) identifier is applied to the first color component (e.g., luma component, red component, green component, blue component, etc.) of the slice data. The slice alf map signalled syntax element indicates the ALF filter values to be used when processing the first color component with ALF. ChromaArrayType of 3 with slice alf enable flag provides an ALF APS identifier for a second color component (e.g., first chroma component such as Cb or Cr component, red component, green component, blue component, etc.) of the slice of video data. Similarly, ChromaArrayType of 3 with slice alf 2 enable flag provides an ALF APS identifier for a third color component (e.g., second chroma component such as Cb or Cr component). The identifier (e.g., slice alf chroma aps id for the second component, or slice alf chroma 2 aps id for the third color component, red component, green component, blue component, etc.) is used with the signaled map (e.g., slice alf chroma map signalled for the second color component, or slice alf chroma 2 map signalled for the third color component) to identify the ALF parameters to be used when filtering the corresponding data of the slice video data.
[0208] In some cases, there are multiple syntax elements that control the application of ALF to chroma components of video data. For example, the slice_alf_chroma_aps_id syntax element mentioned above is a group flag that specifies, by number, that ALF is applied to chroma. For example, the slice_alf_chroma_aps_id syntax element can indicate for a current slice that a first chroma component (e.g., Cb) is filtered, or a second chroma component (e.g., Cr) is filtered, or both the first and second chroma components (e.g., Cb and Cr) are filtered, or no chroma components are filtered. As an example, for the slice_alf_chroma_aps_id syntax element, different chroma components can share the same identifier (ID), such as when the chroma format is 4:2:0. In the case of 4:4:4 format, different color components can have different IDs (e.g., aps_id). The Alf_map syntax element (e.g., slice_alf_chroma_map_signalled) specifies whether the Alf filter is on or off for a given block (e.g., luma CTB). The slice_alf_chroma_map_signalled syntax element specifies that additional syntax elements (e.g., Alf_map syntax elements) signaling ALF on / off maps are signaled for a first chroma component (e.g., Cb). In some cases, the slice_alf_chroma_map_signalled syntax element is signaled for non-4:0:0, 4:2:0 video. The slice_alf_chroma2_map_signalled syntax element specifies that additional syntax elements signaling ALF on / off maps are signaled for a second chroma component (e.g., Cr). In some cases, the slice_alf_chroma2_map_signalled syntax element is signaled for non-4:0:0, 4:2:0 video.
[0209] Figure 5 A process 500 of decoding image and / or video data is shown in accordance with some examples. In some aspects, the process 500 can be implemented in, or by, a system or device having a memory and one or more processors configured to perform the operations of the process 500. In some aspects, the process 500 is implemented in instructions stored in a computer-readable storage medium. For example, when processed by one or more processors of a coding system or device (e.g., system 100), the instructions cause the system or device to perform the operations of the process 500. In other aspects, other implementations are possible in accordance with the details provided herein.
[0210] At block 505, the process 500 includes obtaining a video bitstream. The video bitstream includes adaptive loop filter (ALF) data. In one illustrative example, the ALF data can be signaled using the alf data syntax structure described herein.
[0211] At block 510, the process 500 includes determining a value of an ALF chroma filter signal flag from the ALF data. The value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream. The ALF chroma filter signal flag can also be referred to herein as an ALF flag. In one illustrative example, the ALF chroma filter signal flag can include the alf chroma filter signal flag syntax element in the alf data syntax structure described herein. In some cases, the ALF filter data can include ALF filter coefficients (e.g., f(k,l)) and / or other parameters. In some examples, the process 500 can include inferring the value of the ALF chroma filter signal flag to be zero when the value of the ALF chroma filter signal flag is not present in the ALF data. In some examples, the process 500 can include processing the value of the ALF chroma filter signal flag from the ALF data to determine that chroma ALF filter data is signaled in the video bitstream.
[0212] At block 515, the process 500 includes processing at least a portion of a slice of video data based on the value of the ALF chroma filter signal flag. The portion of the slice can include a block (e.g., CTB, CB, CTU, CB, etc.) of the slice, a plurality of blocks (e.g., two or more CTBs, CBs, CTUs, CBs, etc.) of the slice, or the entire slice. In some aspects, the at least a portion of the slice includes video data in a 4:4:4 format or video data in a non-4:2:0 format. Various examples of processing the at least a portion of the slice of video data are described herein.
[0213] In some examples, process 500 can include obtaining, from a video bitstream, a slice header (e.g., a slice_header( ) syntax structure described herein) for a slice of video data. Process 500 can include determining, from the slice header, a value for an ALF chroma identifier (also referred to herein as a slice ALF chroma identifier). The value for the ALF chroma identifier indicates whether ALF can be applied to one or more chroma components of the slice. In one illustrative example, the ALF chroma identifier can include a slice_alf_chroma_idc syntax element included in the slice_header( ) syntax structure described herein. Process 500 can also include processing at least a portion of the slice based on the ALF chroma identifier from the slice header. In some cases, process 500 can include determining, from the slice header (e.g., from the slice_header( ) syntax structure), a value for a chroma format identifier. For example, the value for the chroma format identifier and the value for the ALF chroma identifier indicate which of the one or more chroma components ALF is applicable to. In one illustrative example, the chroma format identifier can include the ChromaArrayType variable described herein. In some aspects, a value for an ALF chroma filter signal flag (e.g., an alf_chroma_filter_signal_flag syntax element in an alf_data syntax structure) indicates that chroma ALF filter data is signaled in the video bitstream (e.g., and thus ALF is available for the one or more chroma components). In some cases, the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least a portion of the slice. In some examples, based on the value for the ALF chroma filter signal flag, process 500 can include obtaining chroma ALF filter data to be used for processing at least a portion of the slice. Process 500 can also include applying the chroma ALF filter data to at least a portion of the slice of video data.
[0214] In some examples, based on the value for the ALF chroma filter signal flag, process 500 can include obtaining luma ALF filter data to be used for one or more chroma components of at least one block of the video bitstream. Process 500 can further include applying the luma ALF filter data to the one or more chroma components of the at least one block of the video bitstream.
[0215] In some examples, process 500 can include obtaining a slice header for a slice of video data from a video bitstream. As described above, process 500 can include determining a value of a chroma format identifier (e.g., a value of a ChromaArrayType variable from a slice header()) syntax structure) from the slice header. Based on the value of the chroma format identifier from the slice header, process 500 can include processing one or more chroma components of at least one block of the video bitstream using luma ALF filter data.
[0216] As described above, process 500 can include processing a value of an ALF chroma filter signal flag from ALF data to determine that chroma ALF filter data is signaled in the video bitstream. In some examples, process 500 can include determining an ALF application parameter set (APS) identifier for a first color component of at least a portion of a slice. Process 500 can also include determining an ALF map for the first color component of the at least a portion of the slice. In one illustrative example, the ALF map can include a slice_alf_chroma_map_signalled syntax element, a slice_alf_chroma2_map_signalled syntax element, and / or other syntax elements described above or that can be signaled using. In some examples, process 500 can include enabling ALF filtering of at least two non-luma components of at least a portion of a slice based on a component of the at least a portion of the slice that includes a common characteristic. In some cases, the at least two non-luma components of the at least a portion of the slice include a red component, a green component, and a blue component of the at least a portion of the slice. In some cases, the at least two non-luma components of the at least a portion of the slice include one or more chroma components (e.g., a Cb component and / or a Cr component) of the at least a portion of the slice.
[0217] In some examples, process 500 can include enabling ALF filtering of at least two non-luma components of at least a portion of a slice based on the at least a portion of the slice including video data of a non-4:2:0 format.
[0218] As described above, in some cases, the at least portion of the slice includes video data in a 4:4:4 format. In some examples, process 500 can include determining a chroma type array variable for the at least portion of the slice. Process 500 can include determining, based on the chroma type array variable for the at least portion of the slice, an ALF chroma application parameter set (APS) identifier for a first component of the at least portion of the slice. Process 500 can include determining, for the first component of the at least portion of the slice, a signaled ALF map (e.g., a slice_alf_chroma_map_signalled syntax element, a slice_alf_chroma2_map_signalled syntax element, and / or other syntax elements described above). In some examples, process 500 can include determining, based on the chroma type array variable, a second signaled ALF map (e.g., a slice_alf_chroma_map_signalled syntax element, a slice_alf_chroma2_map_signalled syntax element, and / or other syntax elements described above) for a second component of the at least portion of the slice. In some examples, process 500 can include performing ALF filtering for the first and second components of the at least portion of the slice using the signaled ALF map and the second signaled ALF map. In some examples, process 500 can include determining, based on the chroma type array variable, a third signaled ALF map for a third component of the at least portion of the slice. In one illustrative example using the if (sps_alf_flag) syntax provided above, the signaled ALF map includes or is signaled using the slice_alf_map_signalled syntax element mentioned above, the second signaled ALF map includes or is signaled using the slice_alf_chroma_map_signalled syntax element mentioned above, and the third signaled ALF map is or is signaled using the slice_alf_chroma2_map_signalled syntax element mentioned above. In some cases, the first component is a luma component, the second component is a first chroma component, and the third component is a second chroma component. In some cases, the first component is a red component, the second component is a green component, and the third component is a blue component. In some examples, process 500 can include performing ALF processing for blocks of each component of the at least portion of the slice based on the chroma type array variable.
[0219] Figure 6 A process 600 of encoding picture and / or video data is shown, in accordance with some examples. In some aspects, the process 600 can be implemented in, or by, a system or device having a memory and one or more processors configured to perform the operations of the process 600. In some aspects, the process 600 is implemented in instructions stored in a computer-readable storage medium. For example, the instructions, when processed by one or more processors of a codec system or device (e.g., the system 100), cause the system or device to perform the operations of the process 600. In other aspects, other implementations are possible in accordance with the details provided herein.
[0220] In block 605, the process 600 includes generating adaptive loop filter (ALF) data. In one illustrative example, the ALF data includes the alf data syntax structure described herein.
[0221] In block 610, the process 600 includes determining a value of an ALF chroma filter signal flag of the ALF data. The value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream. The ALF chroma filter signal flag can also be referred to herein as an ALF flag. In one illustrative example, the ALF chroma filter signal flag can include the alf chroma filter signal flag syntax element in the alf data syntax structure described herein. In some cases, the ALF filter data can include ALF filter coefficients (e.g., f(k,l)) and / or other parameters.
[0222] In block 615, the process 600 includes generating a video bitstream that includes the ALF data. The bitstream can be generated in accordance with aspects described herein, such as those discussed with respect to Figures 1A-1C 、 Figure 7 and / or Figure 8 .
[0223] In some examples, the process 600 includes determining a value of an ALF chroma identifier (also referred to herein as a slice ALF chroma identifier). The value of the ALF chroma identifier indicates whether the ALF can be applied to one or more chroma components of a slice of the video data. In one illustrative example, the ALF chroma identifier can include the slice alf chroma idc syntax element included in the slice header() syntax structure described herein. The process 600 can include (e.g., add) the value of the ALF chroma identifier in a slice header of the video bitstream. In one illustrative example, the slice header can include the slice header() syntax structure described herein.
[0224] In some examples, process 600 includes determining a value for a chroma format identifier. The value of the chroma format identifier and the value of the ALF chroma identifier can indicate to which chroma component of one or more chroma components the ALF is applicable. In one illustrative example, the chroma format identifier can include a ChromaArrayType variable described herein. Process 600 can include (e.g., add) the value of the chroma format identifier in a slice header of the video bitstream. In some aspects, the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream. In some cases, the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least a portion of a slice of video data.
[0225] As described above, process 600 includes determining a value for a chroma format identifier (e.g., a value of a ChromaArrayType variable from a slice_header() syntax structure). In some cases, the value of the chroma format identifier may indicate one or more chroma components of at least one block of a video bitstream to be processed using luma ALF filter data. Process 600 may include (e.g., add) the value of the chroma format identifier in a slice header of the video bitstream.
[0226] In addition to the aspects described above, it will be apparent that additional aspects are possible within the scope of the details provided herein. For example, within the scope of process 500 and related processes, it is possible to repeat operations or intervene in operations. Additional variations of the above processes will also be apparent from the details described herein. A non-exhaustive list of additional aspects is provided below:
[0227] In some embodiments, the processes (or methods) described herein (including process 500, process 600, and / or other processes described herein) can be performed by a computing device or apparatus such as Figure 1A For example, process 500 and / or process 600 may be performed by Figure 1A and Figure 7 The encoding device 104 shown, another video source side device or video transmission device, Figure 1A and Figure 8 The decoding device 112 and / or another client-side device, such as a player device, a display, or any other client-side device, is shown to perform. In some cases, a computing device or apparatus can include a processor, microprocessor, microcomputer, or other component of a device configured to perform the steps of the processes described herein. In some examples, a computing device or apparatus can include a camera configured to capture video data, e.g., a video sequence, including video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or obtains the captured video data. The computing device can also include a network interface configured to communicate the video data. The network interface can be configured to communicate Internet Protocol (IP) based data or other types of data. In some examples, a computing device or apparatus can include a display for displaying output video content, such as picture samples of a video bitstream.
[0228] The processes 500 and 600 are described with respect to logical flow diagrams, the operations of which represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, the operations represent computer- runnable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer- runnable instructions include routines, programs, objects, components, data structures, procedures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.
[0229] Furthermore, the processes described herein, such as the process 500, the process 600, and / or other processes described herein, can be performed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As described above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0230] The coding techniques discussed herein can be implemented in an example video encoding and decoding system (e.g., system 100). In some examples, a system includes a source device that provides encoded video data to be decoded at a later time by a destination device. In particular, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device can comprise any of a wide variety of devices, including desktop computers, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" telephones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, and the like. In some cases, the source device and the destination device can be equipped for wireless communication.
[0231] The destination device can receive the encoded video data to be decoded via the computer-readable medium. The computer-readable medium can comprise any type of medium or device utilized to
[0232] In some examples, encoded data can be output from an output interface to a storage device. Similarly, encoded data can be accessed from the storage device by an input interface. The storage device can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In a further example, the storage device can correspond to a file server or another intermediate storage device that stores the encoded video generated by source device. The destination device can access stored video data from the storage device via streaming or download. The file server can be any type of server capable of storing encoded video data and transmitting that encoded video data to the destination device. Example file servers include web servers (e.g., for a website), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data from the file server through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on a file server. The transmission of encoded video data from the storage device can be a streaming transmission, a download transmission, or a combination thereof.
[0233] The techniques of this disclosure are not necessarily limited to wireless applications or settings. The techniques can be applied to video coding in support of any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions, such as dynamic adaptive streaming over HTTP (DASH), digital video that is encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, system can be configured to support one-way or two-way video transmission to
[0234] In one example, a source device includes a video source, a video encoder, and an output interface. The destination device can include an input interface, a video decoder, and a display device. The video encoder of the source device can be configured to apply the techniques disclosed herein. In other examples, a source device and destination device can include other components or arrangements. For example, the source device can receive video data from an external video source, such as an external camera. Also, the destination device can interface with an external display device, rather than include an integrated display device.
[0235] The example system above is merely one example. Techniques for parallel processing of video data can be performed by any digital video encoding and / or decoding device. Although generally the techniques of this disclosure are performed by a video encoding device, techniques can also be performed by a video encoder / decoder, typically referred to as a "CODEC." Furthermore, techniques of this disclosure can also be performed by a video preprocessor. Source device and destination device are merely examples of such CODEC devices, in which source device generates encoded video data for transmission to destination device. In some examples, source and destination devices can operate in a substantially symmetrical manner, such that each of the devices includes video encoding and decoding components. Hence, the example system can support one-way or two-way video transmission between video devices, e.g., for video streaming, video playback, video broadcasting, or video telephony.
[0236] The video source can include a video capture device, such as a video camera, a video archive containing previously captured video, and / or a video feed interface to receive video from a video content provider. As a further alternative, the video source can generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In some cases, if video source is a video camera, the source device and the destination device can form so-called camera phones or video phones. However, the techniques described in this disclosure can be applicable to video coding in general, and apply to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by the video encoder. The encoded video information can be output by output interface onto the computer- readable medium.
[0237] As described above, the computer-readable medium can include a transitory medium, such as a wireless broadcast or wired network transmission, or a storage medium (i.e., a non-transitory storage medium), such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, or other computer-readable medium. In some examples, a network server (not shown) can receive encoded video data from the source device and provide the encoded video data to the destination device, e.g., via network transmission. Similarly, a computing device, such as a disk stamping device, can receive encoded video data from the source device and produce a disc containing the encoded video data. Therefore, the computer-readable medium can be understood to include one or more computer-readable media of various forms, in various examples.
[0238] The input interface of destination device receives information from computer- readable medium. The information of the computer-readable medium can include syntax information defined by a video encoder, which is also used by the video decoder, that includes syntax elements that describe characteristics of and / or processing for blocks and other coded units, e.g., groups of pictures (GOPs). The display device displays the decoded video data to a user, and can comprise any of a variety of display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device. Various embodiments of the application have been described.
[0239] The specific details of the encoding device 104 and the decoding device 112 are shown in Figure 7 and Figure 8 respectively. Figure 7 is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in this disclosure. The encoding device 104 may, for example, generate the syntax structures described herein (e.g., syntax structures for VPS, SPS, PPS, or other syntax elements). The encoding device 104 can perform intra- and inter-prediction encoding of video blocks within a video slice. As previously described, intra-coding relies at least in part upon spatial prediction to reduce or remove spatial redundancy
[0240] The encoding device 104 includes a partitioning unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-prediction processing unit 46. For video block reconstruction, the encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. The filter unit 63 is intended to represent one or more loop filters such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 63 is shown as in-loop filters in Figure 7 , in other configurations, the filter unit 63 can be implemented as a post loop filter. A post processing device 57 can perform additional processing on encoded video data generated by the encoding device 104. In some instances, the techniques of this disclosure can be implemented by the encoding device 104. However, in other instances, one or more techniques of this disclosure can be implemented by the post processing device 57.
[0241] like Figure 7 As shown in , the encoding device 104 receives video data and the segmentation unit 35 segments the data into video blocks. Segmentation may also include, for example, segmentation into slices, slice segments, tiles, or other larger units according to a quadtree structure of LCUs and CUs, as well as video block segmentation. The encoding device 104 generally shows components that encode video blocks within a video slice to be encoded. A slice may be divided into multiple video blocks (and possibly into sets of video blocks called tiles). The prediction processing unit 41 may select one of multiple possible encoding modes for the current video block based on error results (e.g., encoding rate and distortion level, etc.), such as one of one or more inter-frame prediction encoding modes from multiple intra-frame prediction encoding modes. The prediction processing unit 41 may provide the resulting intra-frame or inter-frame coded block to the summer 50 to generate residual block data, and to the summer 62 to reconstruct the coded block for use as a reference picture.
[0242] Intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction encoding of the current video block relative to one or more neighboring blocks in the same frame or slice as the current block to be encoded to provide spatial compression. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction encoding of the current video block relative to one or more prediction blocks in one or more reference pictures to provide temporal compression.
[0243] Motion estimation unit 42 may be configured to determine an inter-frame prediction mode for a video slice based on a predetermined pattern for a video sequence. The predetermined pattern may designate a video slice in the sequence as a P slice, a B slice, or a GPB slice. Motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors, which estimate the motion of video blocks. For example, a motion vector may indicate the displacement of a prediction unit (PU) of a video block within a current video frame or picture relative to a prediction block within a reference picture.
[0244] A prediction block is a block that is found to closely match a PU of the video block to be encoded in terms of pixel difference, which may be determined by sum of absolute difference (SAD), sum of squared difference (SSD), or other difference metrics. In some examples, encoding device 104 may calculate values for sub-integer pixel positions of a reference picture stored in picture memory 64. For example, encoding device 104 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Thus, motion estimation unit 42 may perform a motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional pixel precision.
[0245] Motion estimation unit 42 calculates motion vectors for PUs of video blocks in inter-coded slices by comparing the position of the PUs of the video blocks to the position of prediction blocks of reference pictures. The reference pictures can be selected from either a first reference picture list (List 0) or a second reference picture list (List 1), each of which identifies one or more reference pictures stored in picture memory 64. Motion estimation unit 42 sends the calculated motion vectors to entropy encoding unit 56 and motion compensation unit 44.
[0246] Motion compensation performed by motion compensation unit 44 can involve retrieving or generating a prediction block based on the motion vectors determined by motion estimation, possibly performing interpolation to sub-pixel precision. Upon receiving a motion vector for a PU of a current video block, motion compensation unit 44 can locate the prediction block to which the motion vector points in a reference picture list. The encoding device 104 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming pixel difference values. The pixel difference values form the residual data for the block, and can include both luma and chroma difference components. Summer 50 represents the component or components that perform this subtraction operation. Motion compensation unit 44 can also generate syntax elements associated with the video block and video slice for use by the decoding device 112 in decoding the video block of the video slice.
[0247] As described above, as an alternative to inter prediction performed by motion estimation unit 42 and motion compensation unit 44, intra prediction processing unit 46 can intra predict a current block. In particular, intra prediction processing unit 46 can determine an intra prediction mode used to encode the current block. In some examples, intra prediction processing unit 46 can encode the current block using various intra prediction modes, e.g., during separate encoding passes, and can select an appropriate intra prediction mode to use from among the tested modes. For example, intra prediction processing unit 46 can use rate-distortion analysis for the various tested intra prediction modes to calculate rate-distortion values, and can select the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original, unencoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. Intra prediction processing unit 46 can calculate a ratio from the distortion and rate of various encoded blocks to determine which intra prediction mode presents the best rate-distortion values for the block.
[0248] In any case, after the intra prediction mode is selected for the block, the intra prediction processing unit 46 can provide information indicative of the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 can encode the information indicative of the selected intra prediction mode. The encoding device 104 can include configuration data defining of the coding contexts for various blocks in the transmitted bitstream, as well as indications of the most probable intra prediction modes, intra prediction mode index tables, and modified intra prediction mode index tables for each of the contexts. The bitstream configuration data can include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables).
[0249] After the prediction processing unit 41 generates the prediction block for the current video block via either inter prediction or intra prediction, the encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block can be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform. The transform processing unit 52 can convert the residual video data from a pixel domain to a transform domain, such as a frequency domain.
[0250] The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 can perform a scan of a matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 56 can perform the scan.
[0251] After quantization, the entropy encoding unit 56 entropy encodes the quantized transform coefficients. For example, the entropy encoding unit 56 can perform Context- Adaptive Variable Length Coding (CAVLC), Context- Adaptive Binary Arithmetic Coding (CABAC), Syntax-Based Context- Adaptive Binary Arithmetic Coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy encoding technique. After the entropy encoding by the entropy encoding unit 56, the encoded bitstream can be transmitted to the decoding device 112, or archived for later transmission or retrieval by the decoding device 112. The entropy encoding unit 56 can also entropy encode motion vectors and other syntax elements for the current video slice that is being encoded.
[0252] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to the residual blocks in the pixel domain for later use as reference blocks of a reference picture. Motion compensation unit 44 can compute a reference block by adding the residual block to a prediction block of one of the reference pictures within the reference picture list. Motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to compute sub-integer pixel values for use in motion estimation. Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in picture memory 64. Motion estimation unit 42 and motion compensation unit 44 can use the reference blocks as reference blocks for inter prediction of blocks in subsequent video frames or pictures.
[0253] In this manner, Figure 7 Encoding device 104 represents an example of a video encoder configured to perform any of the techniques described herein, including any of the processes or techniques described above. In some cases, some of the techniques of this disclosure can also be implemented by post-processing device 57.
[0254] Figure 8 is a block diagram illustrating an example decoding device 112. Decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and a picture memory 92. Prediction processing unit 81 includes a motion compensation unit 82 and an intra-prediction processing unit 84. In some examples, decoding device 112 can perform decoding that is reciprocal to the encoding performed by encoding device 104, as generally described with respect to FIG. 1. Figure 7 Encoding device 104 represents an example of a video encoder configured to perform any of the techniques described herein, including any of the processes or techniques described above. In some cases, some of the techniques of this disclosure can also be implemented by post-processing device 57.
[0255] During the decoding process, decoding device 112 receives an encoded video bitstream transmitted by encoding device 104, the encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements. In some embodiments, decoding device 112 can receive the encoded video bitstream from encoding device 104. In some embodiments, decoding device 112 can receive the encoded video bitstream from a network entity 79, such as a server, a media aware network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the techniques described above. Network entity 79 can or can not include encoding device 104. Some of the techniques described in this disclosure can be implemented by network entity 79 prior to network entity 79 transmitting the encoded video bitstream to decoding device 112. In some video decoding systems, network entity 79 and decoding device 112 can be parts of discrete devices, while in other cases the functionality described with respect to network entity 79 can be performed by the same device that includes decoding device 112.
[0256] Entropy decoding unit 80 of the decoding device 112 entropy decodes the bitstream to generate quantized coefficients, motion vectors, and other syntax elements. Entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or video block level. Entropy decoding unit 80 can process and parse both fixed-length syntax elements and variable-length syntax elements in one or more parameter sets, such as VPS, SPS, and PPS.
[0257] When the video slice is coded as an intra coded (I) slice, intra-prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for a video block of the current video slice based on the signalled intra-prediction mode and data from previously decoded blocks of the current frame or picture. When the video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, motion compensation unit 82 of the prediction processing unit 81 produces the prediction block for a video block of the current video slice based on the motion vectors and other syntax elements received from the entropy decoding unit 80. The prediction block can be produced from one of the reference pictures within the reference picture lists. The decoding device 112 can use default construction techniques to construct the reference frame lists, List 0 and List 1, based on reference pictures stored in picture memory 92.
[0258] Motion compensation unit 82 determines the prediction information of the video block of the current video slice by parsing the motion vectors and other syntax elements, and uses the prediction information to produce the prediction block for the current video block being decoded. For example, motion compensation unit 82 can use one or more syntax elements in the parameter sets to determine the prediction mode used to code the video block of the video slice (e.g., intra or inter prediction), the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information of one or more reference picture lists for the slice, the motion vectors for each inter coded video block of the slice, the inter prediction status of each inter coded video block of the slice, and other information used to decode the video blocks in the current video slice.
[0259] Motion compensation unit 82 can also perform interpolation based on an interpolation filter. Motion compensation unit 82 can use the interpolation filter as used by the encoding device 104 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. In this case, motion compensation unit 82 can determine the interpolation filter used by the encoding device 104 from the received syntax elements and can use the interpolation filter to produce the prediction block.
[0260] Inverse quantization unit 86 inverse quantizes or de-quantizes quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 80. The inverse quantization process can include using a quantization parameter calculated by the encoding device 104 for each video block in the video slice to determine a degree of quantization and, likewise, a degree of inverse quantization that should be applied. Inverse transform processing unit 88 applies an inverse transform (such as an inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to produce residual blocks in the pixel domain.
[0261] After motion compensation unit 82 generates the predicted block for the current video block based on the motion vectors and other syntax elements, the decoded video block is formed by decoding device 112 by summing the residual block from inverse transform processing unit 88 with the corresponding predicted block generated by motion compensation unit 82. Summer 90 represents the component or components that perform this summation operation. If desired, loop filters (either in loop or after loop) can also be used to smooth pixel transitions, or otherwise improve the video quality. Filter unit 91 is intended to represent one or more loop filters such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although filter unit 91 is shown following motion compensation unit 82 and inverse transform processing unit 88 in FIG. 10, those skilled in the art will Figure 8 recognize that in other configurations, filter unit 91 can be implemented before motion compensation unit 82 or before other components of the decoding pipeline. The decoded video blocks in a given frame or picture are stored in picture memory 92, which stores reference pictures that are used for subsequent motion compensation. Picture memory 92 also stores decoded video for presentation on a display device (such as display 120 shown in FIG. 1 1) at a later time. Figure 1A
[0262] In this way, Figure 8 Decoding device 112 of FIG. 1 1 represents an example of a video decoder configured to perform any of the techniques described herein, including the processes or techniques described above.
[0263] The techniques of this disclosure are not necessarily limited to wireless applications or settings. The techniques can be applied to video coding in support of any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions, such as dynamic adaptive streaming over HTTP (DASH), digital video that is encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, system can be configured to support one-way or two-way video transmission to
[0264] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, and various other mediums capable of storing, containing or carrying instruction(s) and / or data. A computer-readable medium can include a non-transitory medium in which data can be stored and not a carrier wave or signals per se and / or transitory electronic signals. Examples of a non-transitory medium can include, but are not limited to, a magnetic disk, a magnetic tape, or a cassette, an optical disk, e.g., a compact disk (CD) or a digital versatile disk (DVD), a flash memory, a memory or memory device, etc. Computer-readable media can store code and / or machine- readable instructions that can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0265] In some embodiments, computer-readable storage devices, media and memories can include wired or wireless signals, e.g., as the bit stream, etc. However, when referred to as being non-transitory, computer-readable storage media expressly excludes media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0266] In the above description, specific details are provided to provide a thorough understanding of the embodiments and examples provided herein. However, a person of ordinary skill in the art will understand that the embodiments can be practiced without these specific details. In some instances, techniques have been presented as examples in terms of functional blocks, including a functional block that represents a group of devices, components, or steps or routines in software or a combination of software and hardware. Additional components, besides those shown and / or described, can be used. For example, circuitry, systems, networks, processes, and other components can be shown as components in block diagram form, rather than in detail, to avoid obscuring the embodiments. In other instances, well-known structures have not been shown or described in detail to avoid unnecessarily obscuring the embodiments.
[0267] Various embodiments can be described as a process or method according to an example. Although a process or method can be described as a sequence of operations, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations can be re-arranged. A process can be terminated when its operations are completed, but could also occur intermittently during the process. Processes can correspond to methods, functions, routines, subroutines, subprograms, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0268] The processes and methods according to the above-described examples can be implemented using computer-readable media storing computer-executable instructions or otherwise made available to the computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a particular function or group of functions. Portions of computer resources used can be accessible via a network. The computer-executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information and / or data used in the methods according to the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, etc.
[0269] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks can be stored in a computer-readable or machine-readable medium such as a storage medium. A processor(s) can perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, and other devices. Functionality described herein also can be implemented in peripheral devices and / or add-in cards. As further examples, such functionality can be implemented in different circuitry among different chips or on the same chip in different processes.
[0270] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0271] In the preceding description, various aspects of the present application have been described with reference to particular embodiments thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while the illustrative embodiments of the application have been described in detail herein, it should be understood that the application concept can be otherwise variously embodied and used, and that the appended claims are intended to include such alternatives as would be within the scope of the present application, except insofar as limited by the prior art. Various features and aspects of the above-described application can be used individually or jointly. Further, embodiments can be utilized in any number of environments and applications beyond the specific examples described herein, without departing from the broader spirit and scope of the specification. Thus, the specification and drawings are to be regarded as illustrative rather than restrictive. For the purposes of summarizing the disclosure, certain aspects, advantages and features of the application have been described herein. It is to be understood that not necessarily all such advantages can be achieved in accordance with any particular embodiment. Thus, the application described herein specifically references the attainment of at least some of such advantages, but the application can not necessarily achieve all such advantages in accordance with a particular embodiment. Further, specific embodiments of the application have been described herein for the purpose of illustrating the application and not for the purpose of limiting the same, for it will be appreciated that various modifications can be made to the application as described herein to achieve at least some of the advantages of the application. Further, it is to be understood that the application is not limited in its application to the details set forth in the description contained herein above, as the best mode of putting into practice known to the inventor(s). Accordingly, the application is capable of other embodiments and of being practiced and carried out in various ways.
[0272] Those of ordinary skill in the art will appreciate that the less than (<) and greater than (>) symbols or terminology used herein can be replaced by less than or equal to (≤) and greater than or equal to (≥) symbols, respectively, without departing from the scope of the present description.
[0273] Where components are described as being "configured to" perform certain operations, such configuration can be accomplished, for example, by designing the electronic circuitry or other hardware of the component to perform the operation, by programming the component to perform the operation, or any combination thereof.
[0274] The phrase "coupled to" means any connection, direct or indirect, physical or logical, between one component and another component, and / or any connection, direct or indirect, physical or logical, between one component and another component through another intervening component.
[0275] Claim language or other language in a set of recitations indicating "at least one of A and B" or "at least one of A or B" indicates that A alone, B alone, or A and B collectively satisfy the claim. In other words, the claim language or other language in the set of recitations indicates that at least one of A and B are present in the
[0276] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0277] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be realized at least in part by a computer-readable data storage medium or media having computer code thereon for causing a processor to implement at least one of the methods specified above. The computer-readable data storage medium or media can form part of a computer program product. The computer program product can include packaging materials. The computer-readable medium can comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques can be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0278] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor; but, in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein can refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementing the techniques described herein. In addition, in some aspects, the functionality described herein can be provided within dedicated software modules or hardware modules configured for encoding and decoding, or incorporated in a combined video encoder-decoder (CODEC).
[0279] Illustrative aspects of the present disclosure include:
[0280] Aspect 1. An apparatus for decoding video data, the apparatus comprising: a memory and at least one processor (e.g., implemented in circuitry) coupled to the memory. The at least one processor is configured to: obtain a video bitstream, the video bitstream comprising adaptive loop filter (ALF) data; determine, from the ALF data, a value of an ALF chroma filter signal flag, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in the video bitstream; and process at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0281] Aspect 2. The apparatus of aspect 1, wherein the at least one processor is further configured to: obtain, from the video bitstream, a slice header of the slice of the video data; determine, from the slice header, a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether the ALF is applicable to one or more chroma components of the slice; and process the at least a portion of the slice of the video data based on the ALF chroma identifier from the slice header.
[0282] Aspect 3. The apparatus of aspect 2, wherein the at least one processor is further configured to: determine, from the slice header, a value of a chroma format identifier, the value of the chroma format identifier and the value of the ALF chroma identifier indicating which of the one or more chroma components the ALF is applicable to.
[0283] Aspect 4. The apparatus of any of aspects 1 to 3, wherein a value of the ALF chroma filter signal flag indicates that the chroma ALF filter data is signaled in the video bitstream, and wherein the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least a portion of a slice of the video data.
[0284] Aspect 5. The apparatus of any of aspects 1 to 4, wherein the at least one processor is configured to: obtain, based on a value of the ALF chroma filter signal flag, chroma ALF filter data to be used for processing at least a portion of a slice of video data; and apply the chroma ALF filter data to at least a portion of a slice of the video data.
[0285] Aspect 6. The apparatus of any of aspects 1 to 5, wherein the at least one processor is configured to infer a value of the ALF chroma filter signal flag to be zero when the value of the ALF chroma filter signal flag is not present in the ALF data.
[0286] Aspect 7. The apparatus of any of aspects 1 to 6, wherein the at least one processor is configured to: obtain, based on a value of the ALF chroma filter signal flag, luma ALF filter data to be used for one or more chroma components of at least one block of the video bitstream; and apply the luma ALF filter data to the one or more chroma components of at least one block of the video bitstream.
[0287] Aspect 8. The apparatus of any of aspects 1 to 7, wherein the at least one processor is configured to: obtain, from the video bitstream, a slice header of a slice of the video data; determine, from the slice header, a value of a chroma format identifier; and process, based on the value of the chroma format identifier from the slice header, one or more chroma components of at least one block of the video bitstream using the luma ALF filter data.
[0288] Aspect 9. The apparatus of any of aspects 1 to 8, wherein the at least one processor is further configured to: process a value of the ALF chroma filter signal flag from the ALF data to determine that the chroma ALF filter data is signaled in the video bitstream.
[0289] Aspect 10. The apparatus of any of aspects 1 to 9, wherein the at least one processor is further configured to: determine an ALF application parameter set (APS) identifier for a first color component of at least a portion of a slice; and determine an ALF map for the first color component of at least a portion of the slice.
[0290] Aspect 11. The apparatus of any of aspects 1 to 10, wherein the at least one processor is further configured to enable ALF filtering for at least two non-luma components of the at least portion of the slice based on a component of the at least portion of the slice that includes a shared characteristic.
[0291] Aspect 12. The apparatus of aspect 11, wherein the at least two non-luma components of the at least portion of the slice include a red component, a green component, and a blue component of the at least portion of the slice.
[0292] Aspect 13. The apparatus of aspect 11, wherein the at least two non-luma components of the at least portion of the slice include a chroma component of the at least portion of the slice.
[0293] Aspect 14. The apparatus of any of aspects 1 to 13, wherein the at least portion of the slice of video data includes video data in a 4:4:4 format.
[0294] Aspect 15. The apparatus of any of aspects 1 to 14, wherein the at least one processor is configured to enable ALF filtering for at least two non-luma components of the at least portion of the slice based on the at least portion of the slice including video data in a non-4:2:0 format.
[0295] Aspect 16. The apparatus of any of aspects 1 to 15, wherein the at least one processor is configured to determine a chroma type array variable for the at least portion of the slice, determine an ALF chroma application parameter set (APS) identifier for a first component of the at least portion of the slice based on the chroma type array variable for the at least portion of the slice, and determine a signaled ALF map for the first component of the at least portion of the slice.
[0296] Aspect 17. The apparatus of aspect 16, wherein the at least one processor is further configured to determine a second signaled ALF map for a second component of the at least portion of the slice based on the chroma type array variable.
[0297] Aspect 18. The apparatus of aspect 17, wherein the at least one processor is configured to perform ALF filtering on the first component and the second component of the at least portion of the slice using the signaled ALF map and the second signaled ALF map.
[0298] Aspect 19. The apparatus of any of aspects 17 or 18, wherein the at least one processor is further configured to determine a third signaled ALF map for a third component of the at least portion of the slice based on the chroma type array variable.
[0299] Aspect 20. The apparatus of aspect 19, wherein the first component is a luma component, wherein the second component is a first chroma component, and wherein the third component is a second chroma component.
[0300] Aspect 21. The apparatus of aspect 19, wherein the first component is a red component, wherein the second component is a green component, and wherein the third component is a blue component.
[0301] Aspect 22. The apparatus of any of aspects 16 to 21, wherein the at least one processor is configured to perform ALF processing on blocks of each component of at least a portion of a slice based on a chroma type array variable.
[0302] Aspect 23. The apparatus of any of aspects 1 to 22, wherein the apparatus comprises a mobile device.
[0303] Aspect 24. The apparatus of any of aspects 1 to 23, further comprising a display configured to display one or more images.
[0304] Aspect 25. A method of decoding video data, comprising: obtaining a video bitstream, the video bitstream comprising adaptive loop filter (ALF) data; determining, from the ALF data, a value of an ALF chroma filter signal flag, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in the video bitstream; and processing at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0305] Aspect 26. The method of aspect 25, further comprising: obtaining, from the video bitstream, a slice header of a slice of the video data; determining, from the slice header, a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether ALF can be applied to one or more chroma components of the slice; and processing at least the portion of the slice based on the ALF chroma identifier from the slice header.
[0306] Aspect 27. The method of aspect 26, further comprising: determining, from the slice header, a value of a chroma format identifier, the value of the chroma format identifier and the value of the ALF chroma identifier indicating which of the one or more chroma components ALF can be applied to.
[0307] Aspect 28. The method of any of aspects 25 to 27, wherein the value of the ALF chroma filter signal flag indicates that the chroma ALF filter data is signaled in the video bitstream, and wherein the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least the portion of the slice.
[0308] Aspect 29. The method of any of aspects 25 to 28, further comprising: based on a value of the ALF chroma filter signal flag, obtaining chroma ALF filter data to be used for processing at least a portion of a slice of video data; and applying the chroma ALF filter data to at least a portion of a slice of the video data.
[0309] Aspect 30. The method of any of aspects 25 to 29, further comprising: when a value of the ALF chroma filter signal flag is not present in the ALF data, inferring the value of the ALF chroma filter signal flag to be zero.
[0310] Aspect 31. The method of any of aspects 25 to 30, further comprising: based on a value of the ALF chroma filter signal flag, obtaining luma ALF filter data to be used for one or more chroma components of at least one block of a video bitstream; and applying the luma ALF filter data to the one or more chroma components of at least one block of the video bitstream.
[0311] Aspect 32. The method of any of aspects 25 to 31, further comprising: obtaining a slice header of a slice of video data from a video bitstream; determining a value of a chroma format identifier from the slice header; and based on the value of the chroma format identifier from the slice header, using the luma ALF filter data to process one or more chroma components of at least one block of the video bitstream.
[0312] Aspect 33. The method of any of aspects 25 to 32, further comprising: processing a value of the ALF chroma filter signal flag from the ALF data to determine that chroma ALF filter data is signaled in the video bitstream.
[0313] Aspect 34. The method of any of aspects 25 to 33, further comprising: determining an ALF application parameter set (APS) identifier for a first color component of at least a portion of a slice; and determining an ALF map for the first color component of at least a portion of the slice.
[0314] Aspect 35. The method of any of aspects 25 to 34, further comprising: based on a component of at least a portion of a slice that includes a shared characteristic, enabling ALF filtering for at least two non-luma components of at least a portion of the slice.
[0315] Aspect 36. The method of aspect 35, wherein the at least two non-luma components of at least a portion of the slice include a red component, a green component, and a blue component of at least a portion of the slice.
[0316] Aspect 37. The method of aspect 35, wherein the at least two non-luma components of the at least portion of the slice comprise chroma components of the at least portion of the slice.
[0317] Aspect 38. The method of any of aspects 25 to 37, wherein the at least portion of the slice comprises video data in a 4:4:4 format.
[0318] Aspect 39. The method of any of aspects 25 to 38, further comprising enabling ALF filtering for at least two non-luma components of the at least portion of the slice based on the at least portion of the slice comprising video data in a non-4:2:0 format.
[0319] Aspect 40. The method of any of aspects 25 to 39, further comprising determining a chroma type array variable for the at least portion of the slice, determining an ALF chroma application parameter set (APS) identifier for a first component of the at least portion of the slice based on the chroma type array variable for the at least portion of the slice, and determining a signaled ALF map for the first component of the at least portion of the slice.
[0320] Aspect 41. The method of aspect 40, further comprising determining a second signaled ALF map for a second component of the at least portion of the slice based on the chroma type array variable.
[0321] Aspect 42. The method of aspect 41, further comprising performing ALF filtering on the first component and the second component of the at least portion of the slice using the signaled ALF map and the second signaled ALF map.
[0322] Aspect 43. The method of any of aspects 41 or 42, further comprising determining a third signaled ALF map for a third component of the at least portion of the slice based on the chroma type array variable.
[0323] Aspect 44. The method of aspect 43, wherein the first component is a luma component, wherein the second component is a first chroma component, and wherein the third component is a second chroma component.
[0324] Aspect 45. The method of aspect 43, wherein the first component is a red component, wherein the second component is a green component, and wherein the third component is a blue component.
[0325] Aspect 46. The method of any of aspects 25 to 45, further comprising performing ALF processing on blocks of each component of the at least portion of the slice based on the chroma type array variable.
[0326] Aspect 47. An apparatus for encoding video data, the apparatus comprising: a memory; and at least one processor (e.g., implemented in circuitry) coupled to the memory. The at least one processor is configured to: generate adaptive loop filter (ALF) data; determine a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generate the video bitstream including the ALF data.
[0327] Aspect 48. The apparatus of aspect 47, wherein the at least one processor is further configured to: determine a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether the ALF can be applied to one or more chroma components of a slice of the video data; and include the value of the ALF chroma identifier in a slice header of the video bitstream.
[0328] Aspect 49. The apparatus of aspect 48, wherein the at least one processor is further configured to: determine a value of a chroma format identifier, the value of the chroma format identifier and the value of the ALF chroma identifier indicating which of the one or more chroma components the ALF can be applied to; and include the value of the chroma format identifier in the slice header of the video bitstream.
[0329] Aspect 50. The apparatus of any of aspects 47 to 49, wherein the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream, and wherein the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least a portion of a slice.
[0330] Aspect 51. The apparatus of any of aspects 47 to 50, wherein the at least one processor is configured to: determine a value of a chroma format identifier, wherein the value of the chroma format identifier indicates one or more chroma components of at least one block of the video bitstream to be processed using luma ALF filter data; and include the value of the chroma format identifier in a slice header of the video bitstream.
[0331] Aspect 52. The apparatus of any of aspects 47 to 51, wherein the apparatus comprises a mobile device.
[0332] Aspect 53. The apparatus of any of aspects 47 to 52, further comprising a display configured to display one or more images.
[0333] Aspect 54. A method of encoding video data, comprising: generating adaptive loop filter (ALF) data; determining a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generating the video bitstream including the ALF data.
[0334] Aspect 55. The method of aspect 54, further comprising: determining a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether the ALF can be applied to one or more chroma components of a slice of the video data; and including the value of the ALF chroma identifier in a slice header of the video bitstream.
[0335] Aspect 56. The method of aspect 55, further comprising: determining a value of a chroma format identifier, the value of the chroma format identifier and the value of the ALF chroma identifier indicating which of the one or more chroma components the ALF can be applied to; and including the value of the chroma format identifier in the slice header of the video bitstream.
[0336] Aspect 57. The method of any of aspects 54 to 56, wherein the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream, and wherein the chroma ALF filter data is signaled in an adaptation parameter set (APS) for processing at least a portion of a slice.
[0337] Aspect 58. The method of any of aspects 54 to 57, further comprising: determining a value of a chroma format identifier, wherein the value of the chroma format identifier indicates one or more chroma components of at least one block of the video bitstream to be processed using luma ALF filter data; and including the value of the chroma format identifier in a slice header of the video bitstream.
[0338] Aspect 59. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the operations of aspects 1 to 46.
[0339] Aspect 60. An apparatus comprising means for performing any of the operations of aspects 1 to 46.
[0340] Aspect 61. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the operations of aspects 47 to 58.
[0341] Aspect 62. An apparatus comprising means for performing any of the operations of aspects 47 to 58.
[0342] Aspect 63. A computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the operations of aspects 1-58.
[0343] Aspect 64. An apparatus comprising means for performing any of the operations of aspects 1-58.< / highlightend> < / highlight> < / highlightend> < / highlight> < / highlightend> < / highlight> < / highlightend> < / highlight> < / highlightend> < / highlight> < / highlightend> < / highlight> < / highlightend> < / highlight> < / highlightend> < / highlight>
Claims
1. A device for decoding video data, the device comprising: Memory; and at least one processor coupled to the memory and configured to: Obtaining a video bitstream, wherein the video bitstream includes adaptive loop filter (ALF) data; determining a value of an ALF chroma filter signal flag from the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in the video bitstream; and processing at least a portion of a slice of video data based on a value of the ALF chroma filter signal flag, in, Processing at least a portion of the slice of video data comprises: After determining to signal the chroma ALF filter data in the video bitstream based on the value of the ALF chroma filter signal flag, ALF filtering of at least one chroma component for the at least portion of the slice is enabled in response to the at least portion of the slice including video data in a non-4:2:0 format.
2. The device according to claim 1, wherein The at least one processor is further configured to: obtaining a slice header for the slice of video data from the video bitstream; determining a value of an ALF chroma identifier from the slice header, the value of the ALF chroma identifier indicating whether ALF can be applied to one or more chroma components of the slice; as well as The at least portion of the slice of video data is processed based on the ALF chroma identifier from the slice header.
3. The device according to claim 2, wherein The at least one processor is further configured to: A value of a chroma format identifier is determined from the slice header, the value of the chroma format identifier and the value of the ALF chroma identifier indicating to which chroma component of the one or more chroma components ALF is applicable.
4. The device according to claim 3, wherein The value of the ALF chroma filter signal flag indicates that the chroma ALF filter data is signaled in the video bitstream, and wherein the chroma ALF filter data is signaled in an adaptation parameter set APS for processing the at least portion of the slice of video data.
5. The device according to claim 4, wherein The at least one processor is configured to: obtaining the chroma ALF filter data to be used for processing the at least portion of the slice of video data based on a value of the ALF chroma filter signal flag; as well as The chroma ALF filter data is applied to the at least portion of the slice of video data.
6. The device according to claim 1, wherein The at least one processor is configured to infer that the value of the ALF chroma filter signal flag is zero when the value of the ALF chroma filter signal flag is not present in the ALF data.
7. The device according to claim 1, wherein The at least one processor is configured to: obtaining luma ALF filter data for one or more chroma components of at least one block of the video bitstream based on a value of the ALF chroma filter signal flag; as well as The luma ALF filter data is applied to the one or more chroma components of the at least one block of the video bitstream.
8. The device according to claim 1, wherein The at least one processor is configured to: obtaining a slice header for the slice of video data from the video bitstream; determining a value of a chroma format identifier from the slice header; as well as One or more chroma components of at least one block of the video bitstream are processed using luma ALF filter data based on a value of the chroma format identifier from the slice header.
9. The device according to claim 1, wherein The at least one processor is further configured to: The value of the ALF chroma filter signal flag from the ALF data is processed to determine signaling of the chroma ALF filter data in the video bitstream.
10. The device according to claim 9, wherein The at least one processor is further configured to: determining an ALF application parameter set (APS) identifier for a first color component of the at least portion of the slice; as well as An ALF map for the first color component of the at least portion of the slice is determined.
11. The device according to claim 9, wherein The at least one processor is further configured to enable ALF filtering of at least two non-luminance components of the at least portion of the slice based on the at least portion of the components of the slice that include a shared characteristic.
12. The device according to claim 11, wherein The at least two non-luminance components of the at least portion of the slice include a red component, a green component, and a blue component of the at least portion of the slice.
13. The device according to claim 11, wherein The at least two non-luminance components of the at least portion of the slice include the chroma components of the at least portion of the slice.
14. The device according to claim 13, wherein At least a portion of the slice of video data includes video data in a 4:4:4 format.
15. The device according to claim 14, wherein The at least one processor is configured to enable ALF filtering for at least two non-luminance components of the at least portion of the slice based on the at least portion of the slice comprising video data in a non-4:2:0 format.
16. The device according to claim 9, wherein The at least one processor is configured to: determining a chroma type array variable for said at least portion of said slice; determining, based on the chroma type array variable of the at least portion of the slice, an ALF chroma application parameter set (APS) identifier for a first component of the at least portion of the slice; as well as A signaled ALF map for the first component of the at least portion of the slice is determined.
17. The device according to claim 16, wherein The at least one processor is further configured to: Based on the chroma-type array variable, a second signaled ALF map for a second component of the at least portion of the slice is determined.
18. The device according to claim 17, wherein The at least one processor is configured to perform ALF filtering on the first component and the second component of the at least portion of the slice using the signaled ALF map and the second signaled ALF map.
19. The device according to claim 17, wherein The at least one processor is further configured to: Based on the chroma-type array variable, a third signaled ALF map for a third component of the at least portion of the slice is determined.
20. The device according to claim 19, wherein The first component is a luma component, wherein the second component is a first chroma component, and wherein the third component is a second chroma component.
21. The apparatus according to claim 19, wherein The first component is a red component, wherein the second component is a green component, and wherein the third component is a blue component.
22. The apparatus according to claim 19, wherein The at least one processor is configured to perform ALF processing on blocks of each component of the at least portion of the slice based on the chroma-type array variable.
23. The device according to claim 1, wherein The apparatus comprises a mobile device.
24. The apparatus of claim 1, further comprising a display configured to display one or more images.
25. A method for decoding video data, comprising: Obtaining a video bitstream, wherein the video bitstream includes adaptive loop filter (ALF) data; determining a value of an ALF chroma filter signal flag from the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in the video bitstream; and processing at least a portion of a slice of video data based on a value of the ALF chroma filter signal flag, in, Processing at least a portion of the slice of video data comprises: After determining to signal the chroma ALF filter data in the video bitstream based on the value of the ALF chroma filter signal flag, ALF filtering of at least one chroma component for the at least portion of the slice is enabled in response to the at least portion of the slice including video data in a non-4:2:0 format.
26. The method of claim 25, further comprising: obtaining a slice header for the slice of video data from the video bitstream; determining a value of an ALF chroma identifier from the slice header, the value of the ALF chroma identifier indicating whether ALF can be applied to one or more chroma components of the slice; as well as The at least portion of the slice is processed based on the ALF chroma identifier from the slice header.
27. The method of claim 26, further comprising: A value of a chroma format identifier is determined from the slice header, the value of the chroma format identifier and the value of the ALF chroma identifier indicating to which chroma component of the one or more chroma components ALF is applicable.
28. The method according to claim 27, wherein The value of the ALF chroma filter signal flag indicates that the chroma ALF filter data is signaled in the video bitstream, and wherein the chroma ALF filter data is signaled in an adaptation parameter set APS for processing the at least part of the slice.
29. The method of claim 28, further comprising: obtaining the chroma ALF filter data to be used for processing the at least portion of the slice of video data based on the value of the ALF chroma filter signal flag; and The chroma ALF filter data is applied to the at least portion of the slice of video data.
30. The method of claim 25, further comprising: When the value of the ALF chroma filter signal flag is not present in the ALF data, the value of the ALF chroma filter signal flag is inferred to be zero.
31. The method of claim 25, further comprising: obtaining luma ALF filter data for one or more chroma components of at least one block of the video bitstream based on a value of the ALF chroma filter signal flag; and The luma ALF filter data is applied to the one or more chroma components of the at least one block of the video bitstream.
32. The method of claim 25, further comprising: obtaining a slice header for the slice of video data from the video bitstream; determining a value of a chroma format identifier from the slice header; as well as One or more chroma components of at least one block of a video bitstream are processed using luma ALF filter data based on a value of a chroma format identifier from the slice header.
33. The method of claim 25, further comprising: The value of the ALF chroma filter signal flag from the ALF data is processed to determine signaling of the chroma ALF filter data in the video bitstream.
34. The method of claim 33, further comprising: determining an ALF application parameter set (APS) identifier for a first color component of the at least portion of the slice; and An ALF map for the first color component of the at least portion of the slice is determined.
35. The method of claim 33, further comprising: ALF filtering is enabled for at least two non-luminance components of the at least portion of the slice based on the at least portion of the components including the shared characteristic.
36. The method according to claim 35, wherein The at least two non-luminance components of the at least portion of the slice include a red component, a green component, and a blue component of the at least portion of the slice.
37. The method according to claim 35, wherein The at least two non-luminance components of the at least portion of the slice include the chroma components of the at least portion of the slice.
38. The method of claim 37, wherein: The at least portion of the slice includes video data in a 4:4:4 format.
39. The method of claim 38, further comprising: ALF filtering is enabled for at least two non-luminance components of the at least portion of the slice based on the at least portion of the slice comprising video data in a non-4:2:0 format.
40. The method of claim 33, further comprising: determining a chroma type array variable for said at least portion of said slice; determining an ALF chroma application parameter set (APS) identifier for a first component of the at least portion of the slice based on the chroma type array variable of the at least portion of the slice; and A signaled ALF map for the first component of the at least portion of the slice is determined.
41. The method of claim 40, further comprising: Based on the chroma-type array variable, a second signaled ALF map for a second component of the at least portion of the slice is determined.
42. The method of claim 41 , further comprising: ALF filtering is performed on the first component and the second component of the at least portion of the slice using the signaled ALF map and the second signaled ALF map.
43. The method of claim 41 , further comprising: Based on the chroma-type array variable, a third signaled ALF map for a third component of the at least portion of the slice is determined.
44. An apparatus for encoding video data, the apparatus comprising: Memory; and at least one processor coupled to the memory and configured to: Generate adaptive loop filter ALF data; determining a value of an ALF chroma filter signal flag for the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generating the video bitstream including the ALF data, wherein ALF filtering of at least one chroma component of at least a portion of a slice of video data in the video bitstream is enabled in response to the at least a portion of the slice comprising video data in a non-4:2:0 format.
45. A method for encoding video data, comprising: Generate adaptive loop filter ALF data; determining a value of an ALF chroma filter signal flag for the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generating the video bitstream including the ALF data, wherein ALF filtering of at least one chroma component of at least a portion of a slice of video data in the video bitstream is enabled in response to the at least a portion of the slice comprising video data in a non-4:2:0 format.
46. The method of claim 45, further comprising: determining a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether ALF can be applied to one or more chroma components of the slice of video data; and The value of the ALF chroma identifier is included in a slice header of the video bitstream.
47. The method of claim 46, further comprising: determining a value of a chroma format identifier, the value of the chroma format identifier and the value of the ALF chroma identifier indicating to which chroma component of the one or more chroma components ALF is applicable; and The value of the chroma format identifier is included in a slice header of the video bitstream.
48. The method of claim 45, wherein a value of the ALF chroma filter signal flag indicates that the chroma ALF filter data is signaled in the video bitstream, and wherein the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing the at least portion of the slice.
49. The method of claim 45, further comprising: determining a value of a chroma format identifier, wherein the value of the chroma format identifier indicates one or more chroma components of at least one block of the video bitstream to be processed using luma ALF filter data; and The value of the chroma format identifier is included in a slice header of the video bitstream.
50. An apparatus for decoding video data, the apparatus comprising means for performing the method of any one of claims 25-43.
51. An apparatus for encoding video data, the apparatus comprising means for performing the method of any one of claims 45-49.
52. A computer program product comprising computer readable instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 25 to 43.
53. A computer program product comprising computer readable instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 45 to 49.
Citation Information
Patent Citations
An adaptive loop filter
GB201903187D0