Adaptive Loop Filter Processing for Color Format Support
By incorporating an ALF chroma flag for flexible ALF filtering in video coding standards, the limitations of existing standards are addressed, resulting in improved operational efficiency and output performance for video coding devices.
Patent Information
- Application Number
- JP2022561171
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-14
- Filing Date
- 2021-04-15
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2041-04-15
AI Technical Summary
Existing video coding standards lack flexibility in Adaptive Loop Filter (ALF) filtering for chroma components, particularly in video formats where color components have similar characteristics, such as RGB or 4:4:4 formats.
The introduction of additional flags for ALF filtering of multiple color components, specifically an ALF chroma flag, is signaled in the Adaptive Parameter Set (APS) and slice header data, allowing for more flexible ALF filtering across different color components.
This approach enhances the operational efficiency of video coding devices by improving flexibility and output performance, especially in formats where chroma and luma components have similar characteristics.
Smart Images

Figure 0007696920000026 
Figure 0007696920000027 
Figure 0007696920000028
Abstract
Description
Technical Field
[0001] This application relates to video coding. For example, aspects of this application relate to systems, devices, methods, and computer-readable media (referred to as "systems and techniques") for improving loop filters such as an Adaptive Loop Filter (ALF). In some examples, the systems and techniques can enable the coding (e.g., encoding and / or decoding) of video data using different color formats (e.g., 4:4:4 color format, 4:2:0 color format, and / or other color formats).
Background Art
[0002] Many devices and systems enable video data to be processed and output for consumption. Digital video data includes large amounts of data to meet the needs of consumers and video providers. For example, consumers of video data desire the highest quality video with high fidelity, resolution, frame rate, etc. As a result, the large amounts of video data required to meet these needs place a burden on communication networks and devices that process and store video data.
[0003] To compress video data, various video coding techniques can be used. Video coding is performed according to one or more video coding standards. For example, video coding standards include, among others, Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Moving Picture Experts Group (MPEG) coding, VP9, Alliance of Open Media (AOMedia) Video 1 (AV1). Video coding generally utilizes prediction methods (e.g., inter prediction, intra prediction, etc.) that exploit the redundancy present in video images or sequences. An important objective of video coding techniques is to compress video data in a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. As continuously evolving video services become available, coding techniques with better coding efficiency are required.
Summary of the Invention
Means for Solving the Problems
[0004] Systems and techniques for coding (e.g., encoding and / or decoding) image and / or video content are described. In some video coding standards (e.g., the Essential Video Coding (EVC) standard), an Adaptive Loop Filter (ALF) identifies filters pre-stored in memory or signaled (e.g., via an Adaptive Parameter Set (APS)), and filters the luma component using an adaptive filter bank and a classifier for use in ALF filtering. The chroma component of the bitstream is filtered with a single filter (e.g., a 5×5 filter) using coefficients for filters signaled once per APS. However, the operation of such video coding standards lacks flexibility and may not be efficient for videos where different color channels have similar characteristics (e.g., video data having a Red-Green-Blue (RGB) format with a red component, a green component, and a blue component per pixel, video data having a luma component and a chroma component per pixel in 4:4:4 format video data, or other video data). In such formats where color components have similar characteristics, color component data not undergoing ALF filtering in the video coding standards described above can benefit from ALF filtering. However, some syntax elements of video coding standards signal an ALF chroma identifier in the APS (e.g., using one or more APS syntax elements). Using APS syntax elements to control ALF filtering for the chroma component lacks flexibility.
[0005] The examples described in this specification add flags associated with ALF filtering of multiple color components (e.g., an added ALF chroma flag in addition to an existing ALF luma flag) to APS signaling. In some cases, the ALF chroma identifier used to indicate chroma ALF filtering is included in slice header data (e.g., using the syntax structure described below) rather than in ALF data (e.g., signaled using the alf_data syntax structure as described in this specification). Using slice header data for chroma ALF filter signaling can improve the operation of a video coding device (e.g., an encoding device, a decoding device, or a combined encoding / decoding device) due to improved flexibility in ALF filtering. Using slice header data for chroma ALF filter signaling can also improve video output performance (e.g., for chroma data in a video format where chroma and luma have similar characteristics). In some examples, ALF flag signaling can be used to identify an ALF map for ALF filtering for some color components of video data (e.g., blocks of video data that carry data specific to some color components) in a slice of a picture of the video data to which ALF filtering is applied. For example, a decoder receiving a bitstream that includes video data having a 4:4:4 format can identify the presence of an ALF flag (e.g., an ALF chroma filter signal flag) in the bitstream. The ALF flag can indicate that chroma ALF filtering is available for a slice of a picture of the video data. Additional information in the bitstream can indicate an ALF map (e.g., slice_alf_chroma_map_signalled, slice_alf_chroma2_map_signalled, etc.) that provides information used in ALF filtering for at least a portion of the slice.
[0006] According to one exemplary example, an apparatus for decoding video data is provided. The apparatus includes a memory and at least one processor coupled to the memory (e.g., configured in a circuit). The at least one processor is configured to obtain a video bitstream, where the video bitstream includes adaptive loop filter (ALF) data, determine a value of an ALF chroma filter signal flag from the ALF data, where the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream, and process at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0007] According to another exemplary example, a method for decoding video data is provided. The method includes obtaining a video bitstream, where the video bitstream includes adaptive loop filter (ALF) data, determining a value of an ALF chroma filter signal flag from the ALF data, where the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream, and processing at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0008] In another example, when executed by one or more processors, a non-transitory computer-readable medium storing instructions that cause the one or more processors to obtain a video bitstream, where the video bitstream includes adaptive loop filter (ALF) data, determine a value of an ALF chroma filter signal flag from the ALF data, where the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream, and process at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag is provided.
[0009] In another example, an apparatus for decoding video data is provided. The apparatus includes means for obtaining a video bitstream, wherein the video bitstream includes adaptive loop filter (ALF) data, means for determining a value of an ALF chroma filter signal flag from the ALF data, wherein the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream, and means for processing at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0010] Some aspects of the methods, apparatuses, and computer-readable media described above include obtaining a slice header of a slice of video data from a video bitstream, determining a value of an ALF chroma identifier from the slice header, wherein the value of the ALF chroma identifier indicates whether ALF can be applied to one or more chroma components of the slice, and processing at least a portion of the slice based on the ALF chroma identifier from the slice header.
[0011] Some aspects of the methods, apparatuses, and computer-readable media described above include determining a value of a chroma format identifier from the slice header, wherein the value of the chroma format identifier and the value of the ALF chroma identifier indicate to which chroma component(s) of one or more chroma components ALF is applicable.
[0012] In some aspects, the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream. In some aspects, the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of the slice.
[0013] Some aspects of the methods, apparatuses, and computer-readable media described above include obtaining chroma ALF filter data to be used to process at least a portion of a slice of video data based on the value of an ALF chroma filter signal flag, and applying the chroma ALF filter data to at least a portion of a slice of video data.
[0014] Some aspects of the methods, apparatuses, and computer-readable media described above include presuming that the value of the ALF chroma filter signal flag is 0 when the value of the ALF chroma filter signal flag does not exist in the ALF data.
[0015] Some aspects of the methods, apparatuses, and computer-readable media described above include obtaining luma ALF filter data to be used for one or more chroma components of at least one block of a video bitstream based on the value of an ALF chroma filter signal flag, and applying the luma ALF filter data to one or more chroma components of at least one block of a video bitstream.
[0016] Some aspects of the methods, apparatuses, and computer-readable media described above include obtaining a slice header of a slice of video data from a video bitstream, determining a value of a chroma format identifier from the slice header, and processing one or more chroma components of at least one block of a video bitstream using luma ALF filter data based on the value of the chroma format identifier from the slice header.
[0017] Some aspects of the methods, apparatuses, and computer-readable media described above include processing the value of an ALF chroma filter signal flag from ALF data to determine that chroma ALF filter data is signaled in a video bitstream.
[0018] Some aspects of the methods, apparatuses, and computer-readable media described above include determining an ALF application parameter set (APS) identifier for a first color component of at least a portion of a slice and determining an ALF map for the first color component of at least a portion of the slice.
[0019] Some aspects of the methods, apparatuses, and computer-readable media described above include enabling ALF filtering for at least two non-luma components of at least a portion of a slice based on the components of at least a portion of the slice including a shared characteristic.
[0020] In some aspects, at least two non-luma components of at least a portion of a slice include a red component, a green component, and a blue component of at least a portion of the slice.
[0021] In some aspects, at least two non-luma components of at least a portion of a slice include a chroma component of at least a portion of the slice.
[0022] In some aspects, at least a portion of a slice includes 4:4:4 format video data.
[0023] Some aspects of the methods, apparatuses, and computer-readable media described above include enabling ALF filtering for at least two non-luma components of at least a portion of a slice based on at least a portion of the slice including non-4:2:0 format video data.
[0024] Some aspects of the methods, apparatuses, and computer-readable media described above include determining a chroma type array variable for at least a portion of a slice, determining an ALF chroma application parameter set (APS) identifier for a first component of at least a portion of the slice based on the chroma type array variable for at least a portion of the slice, and determining a signaled ALF map for the first component of at least a portion of the slice.
[0025] Some aspects of the methods, apparatuses, and computer-readable media described above include determining a second signaled ALF map for a second component of at least a portion of the slice based on the chroma type array variable.
[0026] Some aspects of the methods, apparatuses, and computer-readable media described above include performing ALF filtering on a first component and a second component of at least a portion of the slice using the signaled ALF map and the second signaled ALF map.
[0027] Some aspects of the methods, apparatuses, and computer-readable media described above include determining a third signaled ALF map for a third component of at least a portion of the slice based on the chroma type array variable.
[0028] In some aspects, the first component is a luminance component, the second component is a first chroma component, and the third component is a second chroma component.
[0029] In some aspects, the first component is a red component, the second component is a green component, and the third component is a blue component.
[0030] Some aspects of the methods, apparatuses, and computer-readable media described above include performing ALF processing on a block for each component of at least a portion of the slice based on the chroma type array variable.
[0031] According to another exemplary example, an apparatus for encoding video data is provided. The apparatus includes a memory and at least one processor coupled to the memory (e.g., configured in a circuit). The at least one processor is configured to generate adaptive loop filter (ALF) data, determine a value of an ALF chroma filter signal flag of the ALF data, where the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in a video bitstream, and generate a video bitstream including the ALF data.
[0032] According to another exemplary example, a method for encoding video data is provided. The method includes generating adaptive loop filter (ALF) data, determining a value of an ALF chroma filter signal flag of the ALF data, where the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in a video bitstream, and generating a video bitstream including the ALF data.
[0033] In another example, when executed by one or more processors, a non-transitory computer-readable medium storing instructions that cause the one or more processors to generate adaptive loop filter (ALF) data, determine a value of an ALF chroma filter signal flag of the ALF data, where the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in a video bitstream, and generate a video bitstream including the ALF data is provided.
[0034] In another example, an apparatus for encoding video data is provided. The apparatus includes means for generating adaptive loop filter (ALF) data, means for determining a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream, and means for generating a video bitstream including the ALF data.
[0035] Some aspects of the method, apparatus, and computer-readable medium for encoding video data described above include determining a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether the ALF can be applied to one or more chroma components of a slice of video data, and including the value of the ALF chroma identifier in a slice header of the video bitstream.
[0036] Some aspects of the method, apparatus, and computer-readable medium for encoding video data described above include determining a value of a chroma format identifier, the value of the chroma format identifier and the value of the ALF chroma identifier indicating to which chroma component among one or more chroma components the ALF is applicable, and including the value of the chroma format identifier in a slice header of the video bitstream.
[0037] In some aspects, the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream. In some aspects, the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of a slice.
[0038] Some aspects of the methods, apparatuses, and computer-readable media described above for encoding video data include determining a value of a chroma format identifier, where the value of the chroma format identifier indicates one or more chroma components of at least one block of a video bitstream to be processed using luma ALF filter data, and including the value of the chroma format identifier in a slice header of the video bitstream.
[0039] In some aspects, the apparatus comprises a mobile device (e.g., a mobile phone or so-called "smartphone", a tablet computer, or other types of mobile devices), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a television, a vehicle (or a computing device of a vehicle), or other devices. In some aspects, the apparatus includes at least one camera for capturing one or more images or video frames. For example, the apparatus can include one camera (e.g., an RGB camera) or multiple cameras for capturing one or more images including video frames and / or one or more videos. In some aspects, the apparatus includes a display for displaying one or more images, videos, notifications, or other displayable data. In some aspects, the apparatus includes a transmitter configured to transmit one or more video frames and / or syntax data to at least one device via a transmission medium. In some aspects, the processor includes a neural processing unit (NPU), a central processing unit (CPU), a graphics processing unit (GPU), or other processing devices or components.
[0040] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the entire specification of this patent, any or all of the drawings, and the appropriate portions of each claim.
[0041] The foregoing will be more apparent when considered in conjunction with the following specification, claims, and accompanying drawings, along with other features and embodiments.
[0042] Exemplary embodiments of the present application will be described in detail below with reference to the following figures.
Brief Description of the Drawings
[0043]
Figure 1A
Figure 1B
Figure 1C
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Best Mode for Carrying Out the Invention
[0044] Some aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for the purpose of explanation, specific details are set forth in order to provide a complete understanding of the embodiments of the present application. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and the description are not intended to be limiting.
[0045] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the exemplary embodiments provides an explanation to enable those skilled in the art to implement the exemplary embodiments. It should be understood that various changes may be made to the functions and configurations of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0046] A video coding device implements video compression techniques to efficiently encode and decode video data. The video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra prediction), temporal prediction (e.g., inter-frame prediction or inter prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques for reducing or removing redundancy inherent in the video sequence. A video encoder can divide each picture of the original video sequence into rectangular regions called video blocks or coding units. The blocks can include coding tree blocks (CTBs), prediction blocks, transform blocks, and / or other suitable blocks. Generally, a reference to a "block" may refer to such a video block (e.g., a CTB, coding block, prediction block, transform block, or other appropriate block or sub-block as would be understood by one of ordinary skill in the art), unless otherwise specified. Further, each of these blocks may also be referred to interchangeably herein as a "unit" (e.g., coding tree unit (CTU), coding unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may refer to a coding logic unit encoded in the bitstream, and a block may refer to a portion of the video frame buffer targeted by the process. In some standards, a coding tree block (CTB) constitutes a CTU and is structured to carry individual color components of the video data. For example, a CTU may include a first CTB for the luma component of the CTU, a second CTB for the chroma blue (Cb) component of the CTU, and a third CTB for the chroma red (Cr) component of the CTU.
[0047] Video blocks can be encoded using specific prediction modes. In the case of the inter prediction mode, the video encoder can search for blocks similar to the block being encoded within a frame (or picture) located at another temporal location, called a reference frame or reference picture. The video encoder can limit this search to a certain spatial displacement from the block to be encoded. Using a two-dimensional (2D) motion vector including a horizontal displacement component and a vertical displacement component, the position of the best match can be determined. In the case of the intra prediction mode, the video encoder can form a predicted block using spatial prediction techniques based on data from previously encoded adjacent blocks within the same picture.
[0048] The video encoder can determine a prediction error. For example, the prediction error can be determined as the difference between the pixel values in the encoded block and the predicted block. The prediction error may also be called a residual. The video encoder can also apply a transform to the prediction error using transform coding (e.g., in the form of a discrete cosine transform (DCT), a discrete sine transform (DST), or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and the motion vectors can be represented using syntax elements and, together with the control information, can form an encoded representation of the video sequence. In some cases, the video encoder can entropy code the syntax elements, thereby further reducing the number of bits required for their representation.
[0049] The video decoder can use the syntax elements and control information described above to construct prediction data (e.g., predicted blocks) for decoding the current frame. For example, the video decoder can add the predicted block and the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis functions using the quantization coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.
[0050] In some cases, one or more adaptive loop filters (ALFs) may be applied individually to color components in the video data to improve the quality of the output video. For example, the ALF may be applied to a picture or a block of a picture after the picture or block has been reconstructed using inter prediction or intra prediction. In some cases, the ALF filtering may be used to correct or straighten artifacts introduced during the reconstruction of the picture or block.
[0051] In some video formats, some color components contain additional data compared to other color components. For example, video data having a 4:2:0 format includes luma data having a higher resolution than the associated chroma data. In such video formats, the ALF filtering may be applied only to the higher resolution luma data. However, in other formats such as red-green-blue (RGB) data and 4:4:4 format data (where all color components have the same sampling rate), the different color components have similar characteristics. Some video coding standards (e.g., the EVC standard) are structured such that the ALF filtering is applied to the luma component and non-luma components (e.g., chroma components) do not receive the ALF filtering. In such cases, the output performance may be improved by applying the ALF to two or more color components (e.g., in addition to being able to apply the ALF filtering to the luma component, applying the ALF filtering to one or more chroma components). The aspects described herein provide additional flexibility and efficient signaling for formats in which ALF filtering of multiple color components results in an improvement in the output image.
[0052] For example, aspects described herein can include applying ALF filter processing to multiple color components of a video bitstream to improve performance. In one example, an RGB format video bitstream can include a picture that is divided into slices that include multiple CTBs. Each slice can include separate CTBs for a red component, a green component, and a blue component. In another example, a video bitstream that includes a luma component and chroma components (e.g., in a luma (Y)-chroma blue (Cb)-chroma red (Cr) format, such as called the YCbCr format) can include a picture that is divided into slices. Each slice can include a CTB for the luma component and CTBs for two chroma components (e.g., a CTB for the Cb component and a CTB for the Cr component). Some video coding standards emphasize ALF filter processing of a single color component (e.g., a luma CTB), but the examples described herein provide slice-header based signaling to enable flexible ALF filter processing of additional color components of video data (e.g., ALF filter processing of either or both chroma CTBs in addition to a luma CTB).
[0053] In some examples, the ALF chroma filter signal flag is added to the ALF data within the video bitstream (e.g., within the parameter set, within the header data such as the slice header, etc.) (e.g., within the alf_data syntax structure). In one exemplary example, the ALF chroma filter signal flag can include the alf_chroma_filter_signal_flag syntax element in the alf_data syntax structure. The ALF chroma filter signal flag operating with the ALF luma filter signal flag can indicate whether chroma filter data is signaled or not. In some cases, the ALF chroma filter signal flag, instead of in the ALF data (e.g., instead of in the alf_data syntax structure), can be used together with a slice ALF chroma identifier (signaled as, e.g., the slice_alf_chroma_idc syntax element) signaled in the slice header to indicate ALF filter processing for additional color components (e.g., non-luma components such as chroma components). For example, the ALF chroma filter signal flag (e.g., the alf_chroma_filter_signal_flag syntax element within the alf_data syntax structure) can indicate that ALF is available for one or more chroma components, and the slice ALF chroma identifier (e.g., slice_alf_chroma_idc in the slice header) can have values (e.g., values from 0 to 3) each indicating a different ALF chroma option. In one example, the first value can indicate that the ALF filter processing should be applied to the first chroma component, the second value can indicate that the ALF filter processing should be applied to the second chroma component, the third value can indicate that the ALF filter processing should be applied to both the first chroma component and the second chroma component, and the fourth value can indicate that the ALF filter processing should not be applied to either the first chroma component or the second chroma component.
[0054] The techniques described herein can be applied to any of existing video codecs (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs), MPEG5 Efficient Video Coding (EVC) (e.g., implemented in ETM5.0), Versatile Video Coding (VVC), Joint Exploration Model (JEM), VP9, AV1, and / or can be efficient coding tools for any developing video coding standard and / or future video coding standards.
[0055] FIG. 1A is a block diagram illustrating an example of a system 100 that includes an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device (also referred to as a client device). The source device and / or the receiving device may include an electronic device such as a mobile or fixed telephone handset (e.g., a smartphone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, a server device in a server system including one or more server devices (e.g., a video streaming server system, or other suitable server system), a head-mounted display (HMD), a head-up display (HUD), smart glasses (e.g., virtual reality (VR) glasses, augmented reality (AR) glasses, or other smart glasses), or any other suitable electronic device.
[0056] The components of system 100 can include, can be included in, and / or can be implemented using one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), electronic circuits or other electronic hardware, and / or can include, and / or can be implemented using, computer software, firmware, or any combination thereof to perform the various operations described herein.
[0057] Although system 100 is shown as including several components, those skilled in the art will appreciate that system 100 can include more or fewer components than shown in FIG. 1A. For example, in some instances, system 100 can include one or more memory devices other than storage 108 and storage 118 (e.g., one or more random access memory (RAM) components, read only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), one or more processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices) communicating with and / or electrically connected to the one or more memory devices, one or more wireless interfaces (e.g., including one or more transceivers and baseband processors per wireless interface) for performing wireless communication, one or more wired interfaces (e.g., serial interfaces such as universal serial bus (USB) inputs, lightning connectors, and / or other wired interfaces) for performing communication via one or more hardwired connections, and / or other components not shown in FIG. 1A.
[0058] The coding techniques described in this specification are applicable to video coding in various multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcast or transmission, encoding of digital video for storage on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, system 100 can support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.
[0059] The encoding device 104 (or encoder) can be used to encode video data using a video coding standard or video coding protocol to generate an encoded video bitstream. Examples of video coding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) including its scalable video coding (SVC) extension and multi-view video coding (MVC) extension, and high efficiency video coding (HEVC) or ITU-T H.265. There are various extensions of HEVC to handle multi-layer video coding, including range extensions and screen content coding extensions, 3D video coding extensions (3D-HEVC) and multi-view extensions (MV-HEVC), as well as scalable extensions (SHVC). HEVC and its extensions have been developed by the Video Coding Joint Team (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG), as well as the Joint Collaborative Team on 3D Video Coding Extension Development (JCT-3V).
[0060] MPEG and ITU-T VCEG have also formed a Joint Exploration Video Team (JVET) to explore and develop new video coding tools for the next-generation video coding standard, called Versatile Video Coding (VVC). The reference software is called the VVC Test Model (VTM). The goal of VVC is to achieve a significant improvement in compression performance over the existing HEVC standard and to support the deployment of higher-quality video services and emerging applications such as, for example, 360° omnidirectional immersive multimedia, high dynamic range (HDR) video, etc. VP9 and Alliance of Open Media (AOMedia) Video 1 (AV1) are other video coding standards to which the techniques described herein may be applied.
[0061] Many of the embodiments described herein may be implemented using a video codec such as MPEG5 EVC, VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein may also be applicable to other video coding standards such as MPEG4 or other MPEG standards, Joint Photographic Experts Group (JPEG) (or other coding standards for still images), VP9, AV1, their extensions, or other suitable coding standards that are already available or not yet available or undeveloped. Thus, while the techniques and systems described herein may sometimes be described with respect to a particular video coding standard, one of ordinary skill in the art will understand that the description should not be construed as applicable only to that particular standard.
[0062] Referring to FIG. 1A, video source 102 may provide video data to encoding device 104. Video source 102 may be part of a source device, or may be part of a device other than the source device. Video source 102 may include a video capture device (e.g., a video camera, a camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider that provides video data, a video feed interface that receives video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.
[0063] Video data from video source 102 may include one or more input pictures. A picture may sometimes be referred to as a "frame". A picture or frame is in some cases a still image that is part of a video. In some examples, the data from video source 102 may be a still image that is not part of a video. In some video coding specifications, a video sequence can include a series of pictures. A picture may include three sample arrays denoted as S L , S Cb , and S Cr . S L is a two-dimensional array of luminance samples, S Cb is a two-dimensional array of Cb chrominance samples, and S Cr is a two-dimensional array of Cr chrominance samples. Chrominance samples may sometimes be referred to herein as "chroma" samples. In other cases, the picture may be monochrome and may contain only an array of luminance samples. A pixel may sometimes refer to a point in a picture that includes luminance and chroma samples. For example, a given pixel may have a luminance sample value from the S L array, a Cb chrominance sample value from the S Cb array, and an S CrIt can include Cr chrominance sample values from the array.
[0064] The encoder engine 106 (or encoder) of the encoding device 104 encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more coded video sequences. A coded video sequence (CVS) starts from an access unit (AU) having a random access point picture with some characteristics in the base layer and extends to just before the next AU having a random access point picture with some characteristics in the base layer, including a series of AUs. For example, some characteristics of the random access point picture that starts a CVS may include a random access skip reading (RASL) picture flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (having a RASL flag equal to 0) does not start a CVS. An access unit (AU) includes one or more coded pictures and control information corresponding to the coded pictures sharing the same output time. The coded slice of a picture is encapsulated at the bitstream level into a data unit called a network abstraction layer (NAL) unit. For example, in some video standards, the video bitstream may include one or more CVSs including NAL units. Each of the NAL units has a NAL unit header. In one example, the header is 1 byte in H.264 / AVC (excluding multi-layer extensions) and 2 bytes in HEVC. The syntax elements in the NAL unit header take designated bits and are thus recognizable in all kinds of systems and transport layers, especially transport streams, real-time transport protocol (RTP), file formats, etc.
[0065] There are two classes of NAL units in some video standards, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain coded picture data that forms the coded video bitstream. For example, a sequence of bits that forms the coded video bitstream exists in a VCL NAL unit. A VCL NAL unit can contain one slice or a slice segment (described below) of the coded picture data, and a non-VCL NAL unit contains control information regarding one or more coded pictures. In some cases, a NAL unit may be called a packet. A HEVC AU contains a VCL NAL unit that contains coded picture data and, if any, a non-VCL NAL unit corresponding to the coded picture data. The non-VCL NAL unit can contain a parameter set that has high-level information regarding the coded video bitstream, among other information. For example, the parameter set can include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). In some cases, each slice or other part of the bitstream can reference a single active PPS, SPS, and / or VPS to enable the decoding device 112 to access the information used to decode the slice or other part of the bitstream.
[0066] A NAL unit may contain a sequence of bits that form an encoded representation of video data, such as an encoded representation of a picture in a video (e.g., an encoded video bitstream, a CVS of a bitstream, etc.). In some cases, the encoder engine 106 can generate an encoded representation of a picture by dividing each picture into a plurality of slices. A slice is independent of other slices such that the information within that slice is coded without dependence on data from other slices within the same picture. A slice includes one or more slice segments, including an independent slice segment and, if present, one or more dependent slice segments that depend on a previous slice segment.
[0067] In some examples, the encoder engine 106 can divide each picture into sub-pictures, slices, and tiles, such as those described in the VVC standard. FIG. 1B is a figure from the VVC standard showing an example of a picture 121 divided into slices and tiles. As shown, picture 121 is divided into one or more tile rows and one or more tile columns. A tile can be defined as a sequence of CTUs that cover a rectangular region of a picture. In some cases, the CTUs within a tile are scanned in raster scan order within that tile. A slice can include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. For example, each vertical slice boundary can also be a vertical tile boundary. As described below, each CTU can include a number of coding tree blocks (CTBs). A sub-picture can include one or more slices that collectively cover a rectangular region of a picture. For example, each sub-picture boundary can also be a slice boundary, and each vertical sub-picture boundary can also be a vertical tile boundary. In some cases, all CTUs within a sub-picture belong to the same tile. In some cases, all CTUs within a tile belong to the same sub-picture.
[0068] In some video standards, a slice is partitioned into luminance samples and chrominance samples of coding tree blocks (CTBs). One or more CTBs of luminance samples and chrominance samples, together with the syntax for the samples, are referred to as a coding tree unit (CTU). As described herein, video data can be structured as CTBs of different color components. Adaptive loop filter (ALF) processing for different color components can be applied to CTBs for each color component. A CTU may also be referred to as a "tree block" or a "largest coding unit" (LCU). A CTU is the basic processing unit for encoding in some standards. A CTU can be split into multiple coding units (CUs) of various sizes. A CU includes a luminance sample array and a chrominance sample array called a coding block (CB).
[0069] A luminance CB and a chrominance CB can be further split into prediction blocks (PBs). A PB is a block of samples of a luminance component or a chrominance component that uses the same motion parameters for inter prediction or intra block copy prediction (when available or enabled for use). A luminance PB and one or more chrominance PBs, together with the associated syntax, form a prediction unit (PU). In the case of inter prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and is used for inter prediction of the luminance PB and one or more chrominance PBs. Motion parameters may also be referred to as motion information. A CB can also be partitioned into one or more transform blocks (TBs). A TB represents a square block of samples of a color component to which a residual transform (e.g., in some cases, the same 2D transform) is applied to code the prediction residual signal. A transform unit (TU) represents the TBs of luminance samples and chrominance samples, as well as the corresponding syntax elements. Transform coding is described in more detail below.
[0070] The size of a CU corresponds to the size of the coding mode and can be square in shape. For example, the size of a CU can be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other suitable size up to the size of the corresponding CTU. The phrase "N×N" is used herein to refer to the pixel dimensions of a video block when converted to vertical and horizontal dimensions (e.g., 8 pixels × 8 pixels). The pixels in a block can be arranged in rows and columns. In some embodiments, a block may not have the same number of pixels in the horizontal direction as in the vertical direction. The syntax data associated with a CU can describe, for example, the partitioning of the CU into one or more PUs. The partitioning mode can differ between whether the CU is coded in an intra prediction mode or an inter prediction mode. A PU can be partitioned such that its shape is non-square. The syntax data associated with a CU can also describe, for example, the partitioning of the CU into one or more TUs according to a CTU. A TU can have a square or non-square shape.
[0071] According to some video coding standards, the transform can be performed using a transform unit (TU). A TU can be different for different CUs. A TU can be sized based on the size of the PUs within a given CU. A TU can be the same size or smaller than a PU. In some examples, the residual samples corresponding to a CU can be re-partitioned into smaller units using a quadtree structure known as a residual quadtree (RQT). The leaf nodes of the RQT can correspond to TUs. The pixel difference values associated with a TU can be transformed to generate transform coefficients. The transform coefficients can be quantized by the encoder engine 106.
[0072] When the pictures of video data are partitioned into CUs, the encoder engine 106 predicts each PU using a prediction mode. The prediction unit or prediction block is subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. The prediction mode can include intra prediction (or intra-picture prediction) or inter prediction (or inter-picture prediction). Intra prediction utilizes the correlation between spatially adjacent samples within a picture. For example, when using intra prediction, each PU is predicted from adjacent picture data within the same picture using, for example, DC prediction to find the average value of the PU, plane prediction to fit a flat surface to the PU, direction prediction to extrapolate from adjacent data, or any other suitable type of prediction. Inter prediction uses the temporal correlation between pictures to derive motion compensation prediction for a block of picture samples. For example, when using inter prediction, each PU is predicted using motion compensation prediction from the picture data in one or more reference pictures (before or after the current picture in output order). The decision of whether to code a picture area using inter-picture prediction or intra-picture prediction can be made, for example, at the CU level.
[0073] The encoder engine 106 and the decoder engine 116 (described in more detail below) can be configured to operate according to a given video coding standard (e.g., EVC). According to some video coding standards, a video coder (such as the encoder engine 106 and / or the decoder engine 116) divides a picture into a plurality of coding tree units (CTUs) (a CTB for luma samples and one or more CTBs for chroma samples, together with the syntax for samples, are called CTUs). The video coder can divide the CTUs according to a tree structure such as a quadtree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concept of multiple division types, such as the distinction between CUs, PUs, and TUs in some standards. The QTBT structure includes two levels, a first level divided according to quadtree division and a second level divided according to binary tree division. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).
[0074] In the MTT division structure, the block can be divided using quadtree division, binary tree division, and one or more types of triple tree division. The triple tree division is a division in which the block is split into three sub-blocks. In some examples, the triple tree division splits the block into three sub-blocks without dividing the original block through the center. The division types in the MTT (e.g., quadtree, binary tree, and triple tree) can be symmetric or asymmetric.
[0075] In some examples, the video coder can use a single QTBT structure or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video coder can use two or more QTBT structures or MTT structures, such as one QTBT structure or MTT structure for the luminance component and another QTBT structure or MTT structure for both chrominance components (or two QTBT structures and / or MTT structures for each respective chrominance component).
[0076] The video coder may be configured to use a quadtree partition, QTBT partition, MTT partition, or other partition structure for each of SOME STANDARDS. For purposes of illustration, the description herein may refer to the QTBT partition. However, it should be understood that the techniques of the present disclosure may also be applicable to video coders configured to use a quadtree partition, or other type of partition.
[0077] In some examples, one or more slices of a picture are assigned a slice type. Slice types include an intracoded slice (I slice), an intercoded P slice, and an intercoded B slice. An I slice (intracoded frame that can be decoded independently) is a slice of a picture that is coded only by intra prediction, and thus can be decoded independently because an I slice only needs data within the frame to predict any prediction unit or prediction block of the slice. A P slice (unidirectional prediction frame) is a slice of a picture that can be coded using intra prediction and unidirectional inter prediction. Each prediction unit or prediction block within a P slice is coded using either intra prediction or inter prediction. When inter prediction is applied, the prediction unit or prediction block is predicted by only one reference picture, and thus the reference samples are from only one reference region of one frame. A B slice (bidirectional prediction frame) is a slice of a picture that can be coded using intra prediction and inter prediction (e.g., either bidirectional or unidirectional). A prediction unit or prediction block of a B slice may be predicted bidirectionally from two reference pictures, where each picture contributes to one reference region and the sets of samples from the two reference regions are weighted (e.g., using equal weights or different weights) to generate the prediction signal for the bidirectional prediction block. As described above, slices of one picture are coded independently. In some cases, a picture may be coded as just one slice.
[0078] As described above, intra-picture prediction utilizes the correlation between spatially adjacent samples within a picture. There are multiple intra prediction modes (also referred to as "intra modes"). In some examples, the intra prediction of a luminance block includes 35 modes, including a planar mode, a DC mode, and 33 angular modes (e.g., a diagonal intra prediction mode and angular modes adjacent to the diagonal intra prediction mode). The 35 intra prediction modes are indexed as shown in Table 1 below. In other examples, more intra modes can be defined that include prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with the angular modes may be different from those used in some standards.
[0079] [Table 1]
[0080] Inter-picture prediction uses the temporal correlation between pictures to derive motion compensation prediction for blocks of an image sample. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the position of the current block, and Δy specifies its vertical displacement. In some cases, the motion vector (Δx, Δy) can be of integer sample precision (also called integer precision), in which case the motion vector points to the integer pel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector (Δx, Δy) can be of fractional sample precision (also called fractional pel precision or non-integer precision) in order to more accurately capture the motion of the underlying object without being restricted to the integer pel grid of the reference frame. The precision of the motion vector is represented by the quantization level of the motion vector. For example, the quantization level can be of integer precision (e.g., 1 pixel) or fractional pel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive the prediction signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values at fractional positions. The previously decoded reference picture is indicated by a reference index (refIdx) to a reference picture list. The motion vector and the reference index are sometimes called motion parameters. Two types of inter-picture prediction can be performed, including single prediction and dual prediction.
[0081] When using inter prediction that employs dual prediction, two sets of motion parameters (Δx0, Δy0, refIdx0 and Δx1, Δy1, refIdx1) are used to generate two motion compensation predictions (from the same reference picture or, in some cases, different reference pictures). For example, when using dual prediction, each prediction block uses two motion compensation prediction signals and generates B prediction units. The two motion compensation predictions are combined to obtain the final motion compensation prediction. For example, the two motion compensation predictions can be combined by averaging. In another example, weighted prediction may be used, in which case different weights can be applied to each motion compensation prediction. The reference pictures that can be used in dual prediction are stored in two separate lists denoted as list 0 and list 1. The motion parameters can be derived at the encoder using a motion estimation process.
[0082] When using inter prediction that employs single prediction, one set of motion parameters (Δx0, Δy0, refIdx0) is used to generate a motion compensation prediction from a reference picture. For example, when using single prediction, each prediction block uses at most one motion compensation prediction signal and generates P prediction units.
[0083] A PU may contain data related to the prediction process (e.g., motion parameters or other suitable data). For example, when a PU is encoded using intra prediction, the PU may contain data describing the intra prediction mode of the PU. As another example, when a PU is encoded using inter prediction, the PU may contain data defining the motion vector of the PU. The data defining the motion vector of the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution of the motion vector (e.g., integer precision, 1 / 4 pixel precision, or 1 / 8 pixel precision), the reference picture pointed to by the motion vector, the reference index, the reference picture list of the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.
[0084] After performing the prediction using intra prediction and / or inter prediction, the encoding device 104 can perform transformation and quantization. For example, following the prediction, the encoder engine 106 can calculate the residual values corresponding to the PUs. The residual values can include pixel difference values between the current block (PU) of the coded pixels and the prediction block (e.g., the predicted version of the current block) used to predict the current block. For example, after generating the prediction block (e.g., using inter prediction or intra prediction), the encoder engine 106 can generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel difference values that quantify the difference between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.
[0085] Any residual data that may remain after the prediction is performed is transformed using a block transform that can be based on a discrete cosine transform (DCT), a discrete sine transform (DST), an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., kernels of size 32×32, 16×16, 8×8, 4×4, or other suitable sizes) can be applied to the residual data within each CU. In some examples, the TUs can be used for the transform process and quantization process implemented by the encoder engine 106. A given CU having one or more PUs can also include one or more TUs. As will be described in more detail below, the residual values may be transformed into transform coefficients using a block transform and may be quantized and scanned using the TUs to generate serialized transform coefficients for entropy coding.
[0086] In some embodiments, following intra prediction coding or inter prediction coding that uses the PU of the CU, the encoder engine 106 may calculate residual data for the TU of the CU. The PU may include pixel data in a spatial region (or pixel region). As described above, the residual data may correspond to the pixel difference value between the pixels of the non-encoded picture and the predicted value corresponding to the PU. The encoder engine 106 may form one or more TUs including the residual data for the CU (including the PU), and may transform the TUs to generate transform coefficients for the CU. The TU may include coefficients in the transform region after applying block transform.
[0087] The encoder engine 106 may perform quantization of the transform coefficients. Quantization achieves further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization may reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient having an n-bit value may be truncated to an m-bit value during quantization, where n is greater than m.
[0088] When quantization is performed, the encoded video bitstream includes quantization transform coefficients, prediction information (e.g., prediction mode, motion vectors, block vectors, etc.), segmentation information, and any other suitable data such as other syntax data. Different elements of the encoded video bitstream can be entropy encoded by the encoder engine 106. In some examples, the encoder engine 106 can scan the quantization transform coefficients using a predefined scan order to generate a serialized vector that can be entropy encoded. In some examples, the encoder engine 106 can perform adaptive scanning. After scanning the quantization transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 can entropy encode the vector. For example, the encoder engine 106 can use context-adaptive variable-length coding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval segmentation entropy coding, or another suitable entropy coding technique.
[0089] The output unit 110 of the encoding device 104 can send NAL units that constitute the encoded video bitstream data to the decoding device 112 of the receiving device via the communication link 120. The input unit 114 of the decoding device 112 can receive the NAL units. The communication link 120 can include a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The wireless network may include any wireless interface or combination of wireless interfaces and may include any suitable wireless network (e.g., the Internet or other wide area network, packet-based network, WiFi (trademark), radio frequency (RF), ultra-wideband (UWB), WiFi-Direct, cellular, long term evolution (LTE), WiMax (trademark), etc.). The wired network may include any wired interface (e.g., fiber, Ethernet, powerline Ethernet, Ethernet via coaxial cable, digital subscriber line (DSL), etc.). The wired network and / or the wireless network can be implemented using various devices such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the receiving device.
[0090] In some examples, the encoding device 104 may store the encoded video bitstream data in the storage 108. The output unit 110 may retrieve the encoded video bitstream data from the encoder engine 106 or from the storage 108. The storage 108 may include any of a variety of data storage media that are distributed or locally accessible. For example, the storage 108 may include a hard drive, a storage disk, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. The storage 108 may also include a decoded picture buffer (DPB) for storing reference pictures for use in inter prediction. In a further example, the storage 108 may correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. In such a case, a receiving device including the decoding device 112 can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing the encoded video data and transmitting the encoded video data to the receiving device. Exemplary file servers can include a web server (e.g., for a website), an FTP server, a network-attached storage (NAS) device, or a local disk drive. The receiving device can access the encoded video data through any standard data connection including an Internet connection. This can include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, which is suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage 108 can be a streaming transmission, a download transmission, or a combination thereof.
[0091] The input section 114 of the decoding device 112 receives the encoded video bitstream data and can provide the video bitstream data to the decoder engine 116 or to the storage 118 for later use by the decoder engine 116. For example, the storage 118 can include a DPB for storing reference pictures for use in inter prediction. A receiving device including the decoding device 112 can receive the encoded video data to be decoded via the storage 108. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the receiving device. The communication medium for transmitting the encoded video data can comprise any wireless communication medium or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other device that can be useful for facilitating communication from a source device to the receiving device.
[0092] The decoder engine 116 can decode the encoded video bitstream data by entropy decoding and extracting elements of one or more coded video sequences that make up the encoded video data (e.g., using an entropy decoder). The decoder engine 116 can perform rescaling and inverse transformation on the encoded video bitstream data. The residual data is passed to the prediction stage of the decoder engine 116. The decoder engine 116 predicts a block of pixels (e.g., a PU). In some examples, the prediction is added to the output of the inverse transformation (residual data).
[0093] The video decoding device 112 may output the decoded video to a video destination device 119 that may include a display or other output device for presenting the decoded video data to a content consumer. In some embodiments, the video destination device 119 may be part of a receiving device that includes the decoding device 112. In some embodiments, the video destination device 119 may be part of a separate device other than the receiving device.
[0094] In some embodiments, the video encoding device 104 and / or the video decoding device 112 may be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary to implement the coding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated as part of a combined encoder / decoder (codec) in their respective devices.
[0095] The exemplary system shown in FIG. 1A is one exemplary example that may be used herein. Techniques for processing video data using the techniques described herein may be performed by any digital video encoding and / or decoding device. In general, the techniques of the present disclosure are performed by a video encoding device or a video decoding device, but the techniques may also be performed by a combined video encoder-decoder, typically referred to as a “codec”. Additionally, the techniques of the present disclosure may also be performed by a video pre-processor. The source device and the receiving device are merely examples of such coding devices that generate coded video data for transmission by the source device to the receiving device. In some examples, the source device and the receiving device may operate substantially symmetrically such that each of the devices includes video encoding and decoding components. Thus, the exemplary system may support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0096] As described previously, some video coding standards define a bitstream that includes a group of NAL units that includes VCL NAL units and non-VCL NAL units. VCL NAL units contain coded picture data that forms the coded video bitstream. For example, a sequence of bits that forms the coded video bitstream is present in a VCL NAL unit. Non-VCL NAL units can include parameter sets that have high-level information regarding the coded video bitstream, among other information. For example, the parameter sets can include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). Examples of the purposes of the parameter sets include providing bitrate efficiency, error resilience, and system layer interfaces. Each slice references a single active PPS, SPS, and VPS in order to access information that a decoding device 112 can use to decode the slice. An identifier (ID) may be coded for each parameter set and includes a VPS ID, an SPS ID, and a PPS ID. The SPS includes the SPS ID and the VPS ID. The PPS includes the PPS ID and the SPS ID. Each slice header includes the PPS ID. Using the IDs, the active parameter sets can be identified for a given slice.
[0097] The PPS contains information applicable to all slices within a given picture. As a result, all slices within a picture refer to the same PPS. Slices in different pictures may also refer to the same PPS. The SPS contains information applicable to all pictures within the same coded video sequence (CVS) or bitstream. As previously explained, a coded video sequence is a series of access units (AUs) starting from a random access point picture (e.g., an instantaneous decoding refresh (IDR) picture or a broken link access (BLA) picture, or other suitable random access point picture) with some characteristics in the base layer, up to just before the next AU with a random access point picture with some characteristics in the base layer (or the end of the bitstream). The information in the SPS may not change for each picture within the coded video sequence. Pictures within a coded video sequence may use the same SPS. The VPS contains information applicable to all layers within a coded video sequence or bitstream. The VPS contains a syntax structure with syntax elements applicable to the entire coded video sequence. In some embodiments, the VPS, SPS, or PPS may be sent in-band along with the encoded bitstream. In some embodiments, the VPS, SPS, or PPS may be sent out-of-band in a separate transmission from the NAL unit containing the coded video data.
[0098] Various chroma formats can be used for video. Chroma format syntax elements can be used to specify chroma sampling. For example, the syntax element chroma_format_idc specifies the chroma sampling relative to luma sampling, e.g., in the VVC standard and / or the EVC standard. Optionally, the value of chroma_format_idc is assumed to be within the range from 0 to 2 inclusive.
[0099] Depending on the value of chroma_format_idc, the values of the variables SubWidthC and SubHeightC are assigned as specified in Clause 6.2 of VVC, and the variable ChromaArrayType is assigned. For example, the values of the variables SubWidthC and SubHeightC can be assigned as follows.
[0100]
Table 2
[0101] In some examples, the variable ChromaArrayType is assigned as follows. - When chroma_format_idc is equal to 0, ChromaArrayType is set equal to 0. - Otherwise, ChromaArrayType is set equal to chroma_format_idc.
[0102] The variables SubWidthC and SubHeightC are specified in Table 1 below according to the chroma format sampling structure specified through chroma_format_idc. Other values of chroma_format_idc, SubWidthC, and SubHeightC may be specified by ISO / IEC in the future.
[0103]
Table 3
[0104] In monochrome sampling, there is only one sample array, which is nominally regarded as a luma array. In 4:2:0 sampling, each of the two chroma arrays (e.g., Cb and Cr) has half the height and half the width of the luma array. In 4:2:2 sampling, each of the two chroma arrays has the same height as the luma array and half the width. In 4:4:4 sampling, each of the two chroma arrays has the same height and width as the luma array.
[0105] In the field of video coding, filtering can be applied to enhance the quality of a decoded or reconstructed video signal. In some cases, the filter may be applied as a post-filter, and the filtered frame is not used for prediction of future frames. In some cases, the filter may be applied as an in-loop filter, and the filtered frame is used to predict one or more future frames. For example, the in-loop filter can filter a picture after reconstruction is performed on the picture (e.g., after adding the residual to the prediction) and before the picture is output and / or before the picture is stored in a picture buffer (e.g., a decoded picture buffer). The filter can be designed, for example, by minimizing the error between the original signal and the decoded and filtered signal. Examples of filters include a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter.
[0106] FIG. 1C shows an exemplary implementation of a filter unit 122 that can be used to process a video picture or block using ALF filter processing according to an example of the present specification. In some cases, the filter unit 122 can be implemented as the filter unit 63 of FIG. 7 and / or the filter unit 91 of FIG. 8. For example, the filter units 63 and 91 can execute the techniques of the present disclosure in cooperation with other components of the video encoding device 104 or the video decoding device 112 in some cases. In some examples, the filter units 63 and 91 can be post-processing units that can execute the techniques of the present disclosure, for example, outside the video encoding device 104 and the video decoding device 112 (for example, after the decoded video is output from the video decoding device 112).
[0107] In the example of FIG. 1C, the filter unit 122 includes a deblocking filter 124, a sample adaptive offset (SAO) filter 126, and an adaptive loop filter (ALF) / geometry transform-based adaptive loop filter (GALF) filter 128. The SAO filter 126 can be configured to determine, for example, offset values for samples of a block. The deblocking filter 124 can be used to compensate for the use of block structure units in the coding process. The ALF filter 128 can be used to minimize the error (for example, mean squared error) between the original sample and the decoded sample by using an adaptive filter that can be a Wiener-based adaptive filter or other suitable adaptive filter. The ALF filter 128 can be configured to determine parameters for filtering the current block based on, for example, signaling parameters for filtering the color components of the current block. As further described herein, the signaling parameters can be based on the sampling format, which can improve device performance by providing ALF filter processing for multiple color components (for example, RGB or 4:4:4 format luma-chroma-chroma format video data) in some video signals.
[0108] In some examples, following the deblocking filter 124, an ALF filter 128 (e.g., involving block-based filter application) may be applied. Optionally, for the luma component of a block, one of 25 filters may be selected for each block (e.g., for each 4×4 block or other sized blocks) through a classification process based on local statistical estimators such as gradients and directions. To benefit from the symmetric characteristics of the filter, the ALF utilized can employ a filter coefficient transformation process. Further details regarding the ALF design are provided below. General loop filters are further described below with respect to FIGS. 7 and 8.
[0109] The filter unit 122 may include fewer filters than that shown in FIG. 1C and / or may include additional filters and / or other components. Additionally, the specific filters shown in FIG. 1C may be implemented in a different order. To smooth pixel transitions or otherwise improve video quality, other loop filters (either within or after the coding loop) may be used. When within the coding loop, the decoded video blocks within a given frame or picture may be stored in a decoded picture buffer (DPB). The DPB stores reference pictures used for subsequent motion compensation (e.g., for inter prediction). The DPB may be part of additional memory that stores decoded video for later presentation on a display device such as the display of the video destination device 119 in FIG. 1A, or may be separate from such memory.
[0110] As described above, different sample formats for color components may result in different performance results when filter processing is applied. Some video coding standards emphasize the ALF filter processing for the luma component (in which case the ALF filter processing is not applied to non-luma components such as chroma components) due to, for example, the predominance of video data having a 4:2:0 format. However, other formats such as 4:4:4 format data RGB data can benefit from improved image quality when the ALF filter processing is applied to two or more color components (e.g., two or three color components in a luma-chroma format such as the RGB format or the YUV format and the YCbCr format).
[0111] Various filter shapes can be used. In some implementations, two diamond filter shapes (as shown in FIGS. 2A and 2B) are used. In some examples, the 5×5 diamond of FIG. 2A and the 7×7 diamond of FIG. 2B may be used to filter luma samples, and the 5×5 diamond shape of FIG. 2A may be used for chroma samples.
[0112] In some cases, block classification may be performed. For example, for the luma component (luma sample) of a pixel, each 4×4 block can be classified into one of 25 classes. The classification index C is derived as follows based on its directionality D and the quantized value of activity
[0113]
Number
[0114] is derived as follows.
[0115]
Number
[0116] In some examples, D and
[0117]
Number
[0118] To calculate, the gradients in the horizontal, vertical, and two diagonal directions are first calculated using a 1D Laplacian.
[0119]
Number
[0120] Here, the indices i and j refer to the coordinates of the top - left sample within a 4×4 block, and R(i,j) indicates the reconstructed sample at the coordinates (i,j).
[0121] To reduce the complexity of block classification, subsampled 1D Laplacian calculations are applied. As shown in FIGS. 3A, 3B, 3C, and 3D showing the subsampled Laplacian calculations, the same subsampled positions are used for gradient calculations in all directions, including the vertical gradient (in FIG. 3A), the horizontal gradient (in FIG. 3B), and the diagonal gradients (in FIGS. 3C and 3D).
[0122] The maximum and minimum values of the horizontal and vertical gradients are
[0123]
Number
[0124] set as
[0125] The maximum and minimum values of the two diagonal gradients are
[0126]
Number
[0127] is set as
[0128] To derive the value of the direction D, these values are compared with each other and with two threshold values t1 and t2. Step 1.
[0129]
Number
[0130] and
[0131]
Number
[0132] If both are true, D is set to 0. Step 2.
[0133]
Number
[0134] If so, continue from Step 3; otherwise, continue from Step 4. Step 3.
[0135]
Number
[0136] If so, D is set to 2; otherwise, D is set to 1. Step 4.
[0137]
Number
[0138] If so, D is set to 4; otherwise, D is set to 3.
[0139] The activity value A is
[0140] [Number]
[0141] calculated as, and the value A is further quantized to the range from 0 to 4 including both end values, and the quantization values are
[0142] [Number]
[0143] shown as.
[0144] For the chroma components in the picture, the classification method is not applied (for example, a single set of ALF coefficients is applied for each chroma component).
[0145] In some cases, a geometric transformation of the filter coefficients may be applied. For example, in some examples, before filtering each 4×4 luminance block, a geometric transformation such as rotation or diagonal and vertical flipping is applied to the filter coefficients f(k,l) according to the gradient values calculated for that block. In some cases, this is equivalent to applying these transformations to the samples in the filter support region. The transformations can make the different blocks to which the ALF is applied more similar by aligning their directions.
[0146] Three geometric transformations including diagonal, vertical flipping and rotation, namely, Diagonal: f D (k,l) = f(l,k), Vertical flipping: f V (k,l) = f(k,K - l - 1), Rotation: f R (k,l) = f(K - l - 1,k) It may be introduced, where K is the size of the filter, and 0 ≦ k, l ≦ K - 1 are the coefficient coordinates such that the location (0, 0) is at the upper left corner and the location (K - 1, K - 1) is at the lower right corner. The transformation is applied to the filter coefficient f(k, l) according to the gradient value calculated for that block. The relationship between the transformation and the four gradients in the four directions is summarized in the following table.
[0147]
Table 4
[0148] In some aspects, an Adaptive Parameter Set (APS) can be used to signal filter parameters or filter data (e.g., filter coefficients and / or other parameters) in a bitstream, such as in an APS NAL unit. The APS can have a related type, such as an ALF type or a Luma Mapping with Chroma Scaling (LMCS) type (defined in, for example, the VVC standard or other video coding standards). The APS can include a set of luma filter parameters, one or more sets of chroma filter parameters, or a combination thereof. The APS can be used in various video coding standards such as VVC, EVC, etc. In some cases, the signaling of the APS can be restricted. For example, a tile group (e.g., a group of one or more tiles such as those shown in Figure 1B) can signal only the index of the APS used for that tile group (e.g., in the tile group header).
[0149] In some examples, each APS may be identified by a unique identifier (e.g., adaptation_parameter_set_id) used to reference the current APS information from other syntax elements. The APS may be shared across a picture and may vary depending on different parts of the picture (e.g., by different tile groups within the picture). When the tile_group_alf_enabled_flag is equal to 1, the APS is referenced by the tile group header, and the ALF parameters carried in the APS are carried out-of-band (e.g., provided by external means other than the video encoding device) and may provide benefits in some situations.
[0150] As described herein, in addition to existing signaling (e.g., flags) indicating ALF filter processing for other color components (e.g., luma components), aspects are described in which the APS can be used to indicate whether ALF filter processing is available for some color components (e.g., chroma components) (e.g., using one or more flags or other syntax). To increase the flexibility and available output quality associated with ALF filter processing for additional color components (e.g., ALF filter processing for chroma components in addition to luma components), additional signaling can be implemented using slice headers.
[0151] For example, examples where filter applicability can be controlled at a block level (e.g., CTB level) are described herein. In some examples, a first flag is signaled to indicate whether ALF is applied to the luma component of a block (e.g., luma CTB), and a second flag is signaled to indicate whether ALF is available for one or more chroma CTBs (such as Cb CTB and / or Cr CTB) of a block. Such flags can be signaled, for example, in ALF data (e.g., the alf_data syntax structure described below). In some aspects, in the case of chroma CTB signaling, the flag can be signaled to indicate whether ALF is available for being applied using the alf_chroma_ctb_present_flag syntax element, as detailed below.
[0152] On the decoder side, when ALF is enabled for a block (e.g., CTB), each sample R(i,j) within the block (e.g., within the CTB or within the coding block of the CTB such as a CU) is filtered to yield a sample value R'(i,j) as shown below, where L indicates the filter length, f m,n represents the filter coefficient, and f(k,l) indicates the decoded filter coefficient.
[0153]
Number
[0154] In some cases, fixed filters may be used. For example, some ALF designs are initialized with a set of fixed filters provided to the decoder as side information. In some cases, there are a total of 64 7×7 filters (e.g., each filter contains 13 coefficients). For each classification class, a mapping is applied to define which 16 fixed filters from the 64 filters can be used for the current class. The selection index (0 to 15) for each class is signaled as the fixed filter index. When adaptively derived filters are used, the difference between the fixed filter coefficients and the adaptive filter coefficients is signaled.
[0155] In some cases, temporal filters may be used. For example, to further benefit from the temporal correlation of video data, the ALF design can utilize the reuse of ALF coefficients signaled previously in the APS NAL unit. Each APS is identified by a unique adaptation_parameter_set_id used to reference the current APS information from other syntax elements (e.g., from the tile group header). In some cases, all signaled APSs with unique set identifier values are stored in an APS buffer (e.g., having a size of up to 32 (entries)). To enable random access (RA) coding configurations, the choices of the encoder's APS adaptation_parameter_set_id usage are restricted. For example, to maintain temporal scalability, only temporal filters from the same or lower temporal layer may be used.
[0156] An example of the adaptive loop filter data syntax is as follows.
[0157]
Table 5A
[0158]
Table 5B
[0159] As described above, the ALF filter processing in some video coding standards (e.g., MPEG5 EVC) involves filtering the luma component using an APS filter bank and a classifier. The adaptive filter bank and classifier can indicate a filter process selected from pre-stored or signaled filters in the APS filter bank. In some such examples, the chroma component can be filtered with a single 5×5 filter whose coefficients are signaled once per APS.
[0160] As further described above, such coding operations can be inefficient for coding RGB format video, 4:4:4 chroma format video, or video having other formats where the color components (e.g., luma and chroma components) have similar characteristics. For example, in the 4:4:4 chroma format, all three color components are present at full resolution. When all three color components thus share the full resolution, the two non-luma components (e.g., chroma components) can benefit from a more advanced filter process (e.g., where a luma type ALF is applied to all three components). As described herein, luma type ALF filter processing may refer to ALF filter processing for higher resolution or more complex data (e.g., when compared to the small filters whose coefficients are signaled once per APS as described above). The examples described herein can signal ALF filter data at the slice header level when the chroma ALF flag is set to enable luma type ALF filter processing for the chroma component, which can enable more flexible ALF filter processing of chroma data.
[0161] Another problem that existed in the ALF filter processing design is that some video coding standards utilize the alf_chroma_idc syntax element that is signaled in the APS. The alf_chroma_idc syntax element is used to control the filtering process (whether filtering is on or off for some video data) for some chroma component data. In some such standards, an alf_chroma_idc equal to 0 specifies that the chroma adaptive loop filter set is not signaled and should not be applied to the Cb and Cr color components, an alf_chroma_idc greater than 0 indicates that the chroma ALF set is signaled, an alf_chroma_idc equal to 1 indicates that the chroma ALF set is applied to the Cb color component, an alf_chroma_idc equal to 2 indicates that the chroma ALF set is applied to the Cr color component, and an alf_chroma_idc equal to 3 indicates that the chroma ALF set is applied to both the Cb and Cr color components. When such signals are shown in the APS without the option to change the settings for subsequent slices that are the subject of the shared APS signaling, there is no option to adjust the ALF filter processing when it is necessary to target the changes that occur at the slice level. Thus, such APS signaling degrades performance by preventing the use of proper chroma ALF filter processing.
[0162] Systems, methods, and computer-readable media are described for improving filter processing (e.g., adaptive loop filter (ALF), deblocking, and / or other filter processing) and enabling coding of video data having different color formats (e.g., 4:4:4 color format, 4:2:0 color format, and / or other color formats). Examples described herein enable additional flexibility in ALF filter processing of color components (e.g., chroma components) and improve coding efficiency and / or output performance for some video coding devices and networks, thereby improving existing techniques (e.g., techniques based on video coding standards). Such improvements can be applied to any video coding standard, such as those that emphasize the luma ALF filter processing described above. As one possible implementation, the following changes to MPEG5 EVC are proposed, and examples are described below in the context of the EVC standard. It will be apparent that similar implementations can be used with other standards having the characteristics described above.
[0163] As part of the improved flexibility described above, in some aspects, the signaling of chroma filters is controlled by a separate flag signaled in the APS. As shown below, the slice ALF chroma identifier (e.g., the slice_alf_chroma_idc syntax element) is moved from the ALF data (e.g., the alf_data syntax structure) to the slice header (e.g., the slice_header() syntax structure) to enable more flexible signaling and more frequent changes to the settings from the slice ALF chroma identifier described below to adapt the available chroma ALF filter processing to the video data format (e.g., to the 4:4:4 format, etc.). Exemplary syntax is described below, along with the portions modified to show the aspects described herein.
[0164]
Table 6
[0165]
Table 7
[0166] As described above, for the ALF data (alf_data syntax structure) and the slice_header syntax structure, the aspects described herein can use the ALF data to signal both the flag for luma ALF and the flag for chroma ALF. As indicated by the above syntax, the alf_chroma_filter_signal_flag syntax element specifies whether chroma filter data is signaled in the bitstream (e.g., in APS). The filter data can optionally include filter coefficients. The slice_alf_chroma_idc syntax element in the slice header data of the bitstream (e.g., the slice_header( ) syntax element) specifies whether ALF is applied to one or more chroma components (e.g., the Cb color component and / or the Cr color component) in the slice. In some examples, when the ALF chroma filter signal flag value (e.g., the value of the alf_chroma_filter_signal_flag syntax element) does not contain an explicit value in the bitstream, the value can be inferred as 0. In such examples, the decoder can determine that the ALF chroma filter signal flag value is 0 from the ALF data even when an explicit flag value is not signaled in the bitstream. The associated slice ALF chroma identifier (e.g., the slice_alf_chroma_idc syntax element) can indicate when ALF is applied to chroma data, such as after the ALF chroma filter signal flag value (e.g., the value of the alf_chroma_filter_signal_flag syntax element) is used to identify that chroma ALF is signaled in the bitstream. In the described implementation, a non-zero value of slice_alf_chroma_idc (e.g., the ALF chroma indication from the slice header) indicates that ALF can be applied to the chroma components of the video data. Specific non-zero values (e.g., 1, 2, 3, etc.) can be used to map which specific color components can have ALF applied to them.
[0167] For example, the ChromaArrayType variable describes the format of the video signal (e.g., whether the signal has chroma at all, such as a monochrome format or a format with one or more chroma components). The slice_alf_chroma_idc syntax element specifies the application of ALF to the chroma components. slice_alf_chroma_idc = 0 specifies that ALF is not applied to the chroma components. slice_alf_chroma_idc = 0 can be signaled even from ChromaArrayType = 2 (indicating that there are chroma components in the video signal). However, when ChromaArrayType == 0 (indicating that the video signal is only luma), slice_alf_chroma_idc cannot be signaled as it has a value greater than 0, which would indicate that chroma filtering is to be performed.
[0168] Figure 4 shows aspects of a process 400 for providing ALF support for different color formats to a decoding device (e.g., decoding device 112) according to several examples. As shown in Figure 4, operation 402 of process 400 involves the decoding device obtaining an encoded bitstream (also referred to as a bitstream or video bitstream). The encoded bitstream can include both ALF data (e.g., the alf_data( ) syntax structure described above) and slice header data (e.g., the slice_header( ) syntax structure described above) for the encoded video data of the encoded bitstream. As part of processing the encoded video data, in operation 404, the decoding device determines whether ALF is enabled for the encoded bitstream. If ALF is not enabled (e.g., not available for processing either the luma color component or the chroma color component), the decoding proceeds without ALF filter processing in operation 405 until the settings change to enable ALF for a portion of the encoded video data (e.g., for one or more slices of the video data).
[0169] When ALF is available, the ALF luma filter signal flag (e.g., the alf_luma_filter_signal_flag syntax element from the alf_data( ) syntax structure above) and the ALF chroma filter signal flag (e.g., the alf_chroma_filter_signal_flag syntax element from the alf_data( ) syntax structure above) may be included in the ALF data as part of the encoded bitstream (e.g., added by the encoding device 104). As described herein, in some cases, the values of such flags may be inferred. For example, in one implementation, a value of 1 may be explicitly signaled for the ALF luma filter signal flag and / or the ALF chroma filter signal flag, and a value of 0 may be inferred when the value is not present in the bitstream. In operation 406, the decoding device can process the ALF data to determine whether the ALF chroma filter signal flag has a value of 0. If the value of the ALF chroma filter signal flag is not 0 (e.g., the value of the flag is 1 or some other value not equal to 1), the decoding device can determine in operation 407 that the chroma filter data is signaled in the APS (instead of being indicated in the slice header data).
[0170] In the implementation form of FIG. 4, when it is determined that the ALF chroma filter signal flag is 0, the decoding device can determine whether ALF is available for application to any chroma component by determining whether the value of the slice ALF chroma indication (e.g., slice_alf_chroma_idc or ChromaArrayType) is a non-zero value (e.g., a value of 1 or another value not equal to 1) in operation 408. If it is determined that the value of the slice ALF chroma indication (e.g., slice_alf_chroma_idc or ChromaArrayType) is not a non-zero value (e.g., it is determined to be a value of 0), the decoding device can determine in operation 409 that ALF should not be applied to the chroma color component of the current slice of the video data (e.g., associated with the slice data including the slice ALF chroma identifier). Based on this determination, the decoding device cannot apply ALF to the chroma color component of the current slice. If it is determined by the decoding device that the slice ALF chroma identifier is a non-zero value (e.g., a value of 1 or another value), the decoding device can determine in operation 410 that ALF is available (or can be applied) for application to one or more chroma color components (e.g., the Cb component and / or the Cr component) of a slice of the video data or a portion thereof (e.g., a block of the slice such as a CTU, CU, CTB, CB, etc.). Then, the decoding device can proceed to apply ALF to one or more of the chroma components based on the specific value of the slice ALF chroma identifier, as further described below.
[0171] In some examples, the signaling and applicability of the ALF parameter may depend on a chroma format identifier (id) indicating the chroma format of video data, such as a 4:2:0 format, a 4:4:4 format, a 4:2:2 format, or other chroma formats. In some cases, a chroma type array variable (e.g., ChromaArrayType) may be used as a chroma format id to indicate the chroma format of video data. In some examples, when video is coded with a ChromaArrayType equal to 3 (corresponding to the 4:4:4 format) or not equal to 4:2:0, ALF filtering is enabled for two non-luma components (e.g., the Cb component and the Cr component) of the coded video data. In some examples, each chroma component can refer to a separate APS to access an optimal filter bank when the video data is coded using a chroma format other than the 4:2:0 format (e.g., using the 4:4:4 format or the 4:2:2 format). In some examples, when video data is coded using a format other than the 4:2:0 format, an ALF classifier is executed on the non-luma components to generate an index for a specific filter. In some examples, for each chroma component, an independent block-based applicability map (e.g., when block-level flags are signaled) is coded for video data having a format other than the 4:2:0 format. In some examples, for video data having the 4:2:0 format and / or for video data having the 4:2:2 format, chroma filtering may be performed using a single filter without a classifier, no map is signaled, and in that case, any block may be conditionally filtered.
[0172] An example of the syntax structure and semantics that can be used as part of operation 410 to determine how available chroma ALF should be applied is as follows (changes to the EVC standard are "<highlight>」 symbol and 「 <highlightend>Underlined text indicating new text and strikethrough-marked text indicating old deleted text, between the "」" symbol).
[0173]
Table 8
[0174]
Table 9
[0175] alf_ctb_flag[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] equal to 1 specifies that the adaptive loop filter is applied to the coding tree block of the luma component of the coding tree unit at the luma location (xCtb, yCtb). alf_ctb_flag[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] equal to 0 specifies that the adaptive loop filter is not applied to the coding tree block of the luma component of the coding tree unit at the luma location (xCtb, yCtb). When alf_ctb_flag[ xCtb >> CtbLog2SizeY ][ yCtb >> CtbLog2SizeY ] does not exist, it is assumed to be equal to slice_alf_enabled_flag.
[0176]
Table 10
[0177] For any coding tree unit having a luma coding tree block location (rx, ry), when rx = 0 PicWidthInCtbsY - 1 and ry = 0 PicHeightInCtbsY - 1, the following applies. - When alf_ctb_flag[rx][ry] is equal to 1, the specified coding tree block luma type filter processing process takes recPicture L , alfPicture L , the reference APS identifier slice_alf_luma_aps_id, and the luma coding tree block location (xCtb, yCtb) set equal to (rx << CtbLog2SizeY, ry << CtbLog2SizeY) as inputs and the output is the modified filtered picture alfPicture L . - When ChromaArrayType is equal to 3, the coding tree block luma type filter processing process for the chroma samples of the Cb chroma component and the Cr chroma component is called as follows. - When alf_ctb_chroma_flag[rx][ry] is equal to 1, the specified coding tree block luma type filter processing process takes recPictureCb, alfPictureCb, the reference APS identifier slice_alf_chroma_aps_id, and the chroma coding tree block location (xCtb, yCtb) set equal to (rx << CtbLog2SizeY, ry << CtbLog2SizeY) as inputs and the output is the modified filtered picture alfPictureCb. - When alf_ctb_chroma2_flag[rx][ry] is equal to 1, the specified coding tree block luma type filter processing process takes recPictureCr, alfPictureCr, the reference APS identifier slice_alf_chroma2_aps_id, and the chroma coding tree block location (xCtb, yCtb) set equal to (rx << CtbLog2SizeY, ry << CtbLog2SizeY) as inputs and the output is the modified filtered picture alfPictureCr. - Otherwise, when ChromaArrayType is in the range from 1 to 2 inclusive of both end values and slice_alf_chroma_idc is greater than 0, the following applies. - When sliceChromaAlfEnableFlag is equal to 1, the specified coding tree block chroma type filtering process is called with recPicture set equal to recPictureCb, alfPicture set equal to alfPictureCb, reference APS identifier slice_alf_chroma_aps_id, and chroma coding tree block location (xCtbC, yCtbC) set equal to ((rx << CtbLog2SizeY) / SubWidthC, (ry << CtbLog2SizeY) / SubHeightC) as inputs, and the output is the modified filtered picture alfPictureCb. - When sliceChroma2AlfEnableFlag is equal to 1, the specified coding tree block chroma type filtering process is called with recPicture set equal to recPictureCr, alfPicture set equal to alfPictureCr, reference APS identifier slice_alf_chroma_aps_id, and chroma coding tree block location (xCtbC, yCtbC) set equal to ((rx << CtbLog2SizeY) / SubWidthC, (ry << CtbLog2SizeY) / SubHeightC) as inputs, and the output is the modified filtered picture alfPictureCr.
[0178] The above is an example of an implementation as a modification to EVC, but similar modifications can be made to the modes in other video coding standards. The above syntax structure (e.g., the if( sps_alf_flag ) syntax structure) can be part of the slice header syntax structure (e.g., the slice_header() syntax structure) in the EVC encoded bitstream. If the decoder device determines that ALF is generally available (e.g., for the luma component and the chroma component) and available for signaling at the slice level (e.g., not in APS as described as part of operation 407), the above syntax structure can be processed by the decoder device to determine the specific chroma component to which ALF should be applied.
[0179] For example, as shown in the above exemplary slice header syntax, when the slice_alf_enable_flag is true (e.g., a value of 1), the decoding device can check whether the ChromaArrayType variable has a value of 1 or 2 when slice_alf_chroma_idc (e.g., the ALF chroma indication from the slice header) is greater than 0. If this statement is true (e.g., ChromaArrayType is 1 or 2), an ALF application parameter set (APS) identifier is applied to the first color component (e.g., luma component, red component, green component, blue component, etc.) of the slice data. The slice_alf_map_signalled syntax element indicates the ALF filter value to be used when processing the first color component with ALF. A ChromaArrayType of 3 with sliceChromaAlfEnabledFlag provides an ALF APS identifier for the second color component (e.g., the first chroma component such as the Cb component or Cr component, red component, green component, blue component, etc.) of the slice of video data. Similarly, a ChromaArrayType of 3 with sliceChroma2AlfEnabledFlag provides an ALF APS identifier for the third color component (e.g., the second chroma component such as the Cb component or Cr component). The identifier (e.g., slice_alf_chroma_aps_id for the second component or slice_alf_chroma2_aps_id for the third color component, red component, green component, blue component, etc.) is used together with the signalled map for the second color component (e.g., slice_alf_chroma_map_signalled) or slice_alf_chroma2_map_signalled for the third color component to identify the ALF parameters to be used when filtering the corresponding data of the slice of video data.
[0180] In some cases, there are multiple syntax elements that control the application of ALF to the chroma components of video data. For example, the above-mentioned slice_alf_chroma_aps_id syntax element is a group flag that specifies ALF application to chroma by a number. For example, the slice_alf_chroma_aps_id syntax element can indicate that for the current slice, the first chroma component (e.g., Cb) is filtered, or the second chroma component (e.g., Cr) is filtered, or both the first chroma component and the second chroma component (e.g., Cb and Cr) are filtered, or neither chroma component is filtered. In the case of the slice_alf_chroma_aps_id syntax element, different chroma components can share the same identifier (ID), for example, when the chroma format is 4:2:0. In the case of the 4:4:4 format, different color components may have different IDs (e.g., aps_id). The Alf_map syntax element (e.g., slice_alf_chroma_map_signalled) specifies whether the ALF filter is ON or OFF for a given block (e.g., a luma CTB). The slice_alf_chroma_map_signalled syntax element specifies that an additional syntax element of the ALF ON / OFF map (e.g., the Alf_map syntax element) is signaled for the first chroma component (e.g., Cb). In some cases, the slice_alf_chroma_map_signalled syntax element is signaled for non-4:0:0 video, 4:2:0 video. The slice_alf_chroma2_map_signalled syntax element specifies that an additional syntax element of the ALF ON / OFF map is signaled for the second chroma component (e.g., Cr). In some cases, the slice_alf_chroma2_map_signalled syntax element is signaled for non-4:0:0 video, 4:2:0 video.
[0181] Figure 5 shows a process 500 for decoding image and / or video data according to several examples. In some aspects, process 500 may be implemented in or by a system or apparatus having a memory and one or more processors configured to execute the operations of process 500. In some aspects, process 500 is implemented in instructions stored on a computer-readable storage medium. For example, when the instructions are processed by one or more processors of a coding system or apparatus (e.g., system 100), the system or apparatus is caused to execute the operations of process 500. In other aspects, other implementations are possible according to the details provided herein.
[0182] In block 505, process 500 includes obtaining a video bitstream. The video bitstream includes adaptive loop filter (ALF) data. In one exemplary example, the ALF data may be signaled using the alf_data syntax structure described herein.
[0183] In block 510, process 500 includes determining a value of an ALF chroma filter signal flag from ALF data. The value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream. The ALF chroma filter signal flag may also be referred to herein as the ALF flag. In one exemplary example, the ALF chroma filter signal flag can include an alf_chroma_filter_signal_flag syntax element in the alf_data syntax structure described herein. Optionally, the ALF filter data can include ALF filter coefficients (e.g., f(k, l)) and / or other parameters. In some examples, process 500 can include inferring that the value of the ALF chroma filter signal flag is 0 when the value of the ALF chroma filter signal flag does not exist in the ALF data. In some examples, process 500 can include processing the value of the ALF chroma filter signal flag from the ALF data to determine that chroma ALF filter data is signaled in the video bitstream.
[0184] In block 515, process 500 includes processing at least a portion of a slice of video data based on the value of the ALF chroma filter signal flag. A portion of the slice can include a block of the slice (e.g., CTB, CB, CTU, CB, etc.), multiple blocks of the slice (e.g., two or more CTBs, CBs, CTUs, CBs, etc.), or the entire slice. In some aspects, at least a portion of the slice includes 4:4:4 format video data or non-4:2:0 format video data. Various examples of processing at least a portion of a slice of video data are described herein.
[0185] In some examples, process 500 can include obtaining a slice header of a slice of video data (e.g., the slice_header( ) syntax structure described herein) from a video bitstream. Process 500 can include determining a value of an ALF chroma identifier (also referred to herein as a slice ALF chroma identifier) from the slice header. The value of the ALF chroma identifier indicates whether ALF can be applied to one or more chroma components of the slice. In one exemplary example, the ALF chroma identifier can include a slice_alf_chroma_idc syntax element included in the slice_header( ) syntax structure described herein. Process 500 can further include processing at least a portion of the slice based on the ALF chroma identifier from the slice header. In some cases, process 500 can include determining a value of a chroma format identifier from the slice header (e.g., from the slice_header( ) syntax structure). For example, the value of the chroma format identifier and the value of the ALF chroma identifier indicate which chroma components among one or more chroma components ALF is applicable to. In one exemplary example, the chroma format identifier can include a ChromaArrayType variable described herein. In some aspects, the value of an ALF chroma filter signal flag (e.g., the alf_chroma_filter_signal_flag syntax element in the alf_data syntax structure) indicates that chroma ALF filter data is signaled in the video bitstream (thus, for example, ALF is available for one or more chroma components). In some cases, the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of the slice. In some examples, based on the value of the ALF chroma filter signal flag, process 500 can include obtaining chroma ALF filter data to be used for processing at least a portion of the slice.Process 500 can further include applying chroma ALF filter data to at least a portion of a slice of video data.
[0186] In some examples, based on the value of the ALF chroma filter signal flag, process 500 can include obtaining luma ALF filter data to be used for one or more chroma components of at least one block of the video bitstream. Process 500 can further include applying the luma ALF filter data to one or more chroma components of at least one block of the video bitstream.
[0187] In some examples, process 500 can include obtaining a slice header of a slice of video data from the video bitstream. As described above, process 500 can include determining the value of the chroma format identifier (e.g., the value of the ChromaArrayType variable from the slice_header( ) syntax structure) from the slice header. Based on the value of the chroma format identifier from the slice header, process 500 can include processing one or more chroma components of at least one block of the video bitstream using the luma ALF filter data.
[0188] As described above, process 500 can include processing the value of the ALF chroma filter signal flag from the ALF data to determine that chroma ALF filter data is signaled in the video bitstream. In some examples, process 500 can include determining an ALF application parameter set (APS) identifier for a first color component of at least a portion of a slice. Process 500 can further include determining an ALF map for a first color component of at least a portion of a slice. In one exemplary example, the ALF map can include or be signaled using the slice_alf_chroma_map_signalled syntax element, the slice_alf_chroma2_map_signalled syntax element, and / or other syntax elements described above. In some examples, process 500 can include enabling ALF filter processing for at least two non-luma components of at least a portion of a slice based on the components of at least a portion of the slice having a shared characteristic. In some cases, the at least two non-luma components of at least a portion of the slice include the red, green, and blue components of at least a portion of the slice. In some cases, the at least two non-luma components of at least a portion of the slice include one or more chroma components (e.g., Cb component and / or Cr component) of at least a portion of the slice.
[0189] In some examples, process 500 can include enabling ALF filter processing for at least two non-luma components of at least a portion of a slice based on the at least a portion of the slice including non-4:2:0 format video data.
[0190] As described above, in some cases, at least a portion of the slice includes 4:4:4 format video data. In some examples, process 500 can include determining a chroma type array variable for at least a portion of the slice. Process 500 can include determining an ALF chroma application parameter set (APS) identifier for a first component of at least a portion of the slice based on the chroma type array variable for at least a portion of the slice. Process 500 can include determining a signaled ALF map (e.g., slice_alf_chroma_map_signalled syntax element, slice_alf_chroma2_map_signalled syntax element, and / or other syntax elements described above) for a first component of at least a portion of the slice. In some examples, process 500 can include determining a second signaled ALF map (e.g., slice_alf_chroma_map_signalled syntax element, slice_alf_chroma2_map_signalled syntax element, and / or other syntax elements described above) for a second component of at least a portion of the slice based on the chroma type array variable. In some examples, process 500 can include performing ALF filtering on a first component and a second component of at least a portion of the slice using the signaled ALF map and the second signaled ALF map. In some examples, process 500 can include determining a third signaled ALF map for a third component of at least a portion of the slice based on the chroma type array variable.In one exemplary example using the if( sps_alf_flag ) syntax given above, the signaled ALF map is signaled including or using the slice_alf_map_signalled syntax element described above, the second signaled ALF map is signaled including or using the slice_alf_chroma_map_signalled syntax element described above, and the third signaled ALF map is signaled including or using the slice_alf_chroma2_map_signalled syntax element described above. In some cases, the first component is the luma component, the second component is the first chroma component, and the third component is the second chroma component. In some cases, the first component is the red component, the second component is the green component, and the third component is the blue component. In some examples, process 500 can include performing ALF processing on a block for each component of at least a portion of a slice based on a chroma type array variable.
[0191] FIG. 6 shows a process 600 for encoding image and / or video data according to some examples. In some aspects, process 600 can be implemented in or by a system or apparatus having a memory and one or more processors configured to perform the operations of process 600. In some aspects, process 600 is implemented in instructions stored on a computer-readable storage medium. For example, when the instructions are processed by one or more processors of a coding system or apparatus (e.g., system 100), the system or apparatus is caused to perform the operations of process 600. In other aspects, other implementations are possible in accordance with the details provided herein.
[0192] In block 605, process 600 includes generating adaptive loop filter (ALF) data. In one exemplary example, the ALF data includes the alf_data syntax structure described herein.
[0193] In block 610, process 600 includes determining the value of the ALF chroma filter signal flag for the ALF data. The value of the ALF chroma filter signal flag indicates whether the chroma ALF filter data is signaled in the video bitstream. The ALF chroma filter signal flag may also be referred to herein as the ALF flag. In one exemplary example, the ALF chroma filter signal flag can include the alf_chroma_filter_signal_flag syntax element in the alf_data syntax structure described herein. Optionally, the ALF filter data can include ALF filter coefficients (e.g., f(k, l)) and / or other parameters.
[0194] In block 615, process 600 includes generating a video bitstream that includes the ALF data. The bitstream can be generated according to the aspects described herein, such as those described with respect to FIGS. 1, 7, and / or 8.
[0195] In some examples, process 600 includes determining the value of an ALF chroma identifier (also referred to herein as a slice ALF chroma identifier). The value of the ALF chroma identifier indicates whether ALF can be applied to one or more chroma components of a slice of video data. In one exemplary example, the ALF chroma identifier can include the slice_alf_chroma_idc syntax element included in the slice_header( ) syntax structure described herein. Process 600 can include (e.g., add) the value of the ALF chroma identifier to the slice header of the video bitstream. In one exemplary example, the slice header can include the slice_header( ) syntax structure described herein.
[0196] In some examples, process 600 includes determining a value of a chroma format identifier. The value of the chroma format identifier and the value of the ALF chroma identifier can indicate which of one or more chroma components the ALF is applicable to. In one exemplary example, the chroma format identifier can include the ChromaArrayType variable described herein. Process 600 can include (e.g., add) the value of the chroma format identifier to the slice header of the video bitstream. In some aspects, the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream. In some cases, the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of a slice of video data.
[0197] As described above, process 600 includes determining a value of a chroma format identifier (e.g., the value of the ChromaArrayType variable from the slice_header( ) syntax structure). In some cases, the value of the chroma format identifier can indicate one or more chroma components of at least one block of the video bitstream to be processed using the luma ALF filter data. Process 600 can include (e.g., add) the value of the chroma format identifier to the slice header of the video bitstream.
[0198] In addition to the aspects described above, it will be apparent that additional aspects are possible within the scope of the details provided herein. For example, repeated operations or intervening operations are possible within the scope of process 500 and related processes. Additional variations to the processes described above will also be apparent from the details described herein. A non-exhaustive list of additional aspects is provided below.
[0199] In some implementations, the processes (or methods) described herein (including process 500, process 600, and / or other processes described herein) may be executed by a computing device or apparatus such as system 100 shown in FIG. 1A. For example, process 500 and / or process 600 may be executed by the encoding device 104 shown in FIGS. 1A and 7, by another video source side device or video transmitting device, by the decoding device 112 shown in FIGS. 1A and 8, and / or by another client side device such as a player device, a display, or any other client side device. In some cases, the computing device or apparatus may include a processor, a microprocessor, a microcomputer, or other components of the device configured to perform the steps of the processes described herein. In some examples, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) including video frames. In some examples, a camera or other capture device for capturing video data is separate from the computing device, in which case the computing device receives or acquires the captured video data. The computing device may further include a network interface configured to communicate video data. The network interface may be configured to communicate Internet Protocol (IP)-based data or other types of data. In some examples, the computing device or apparatus may include a display for displaying output video content such as samples of pictures of a video bitstream.
[0200] Processes 500 and 600 are described with respect to a logical flow diagram, and their operations represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0201] In addition, the processes described herein (e.g., Process 500, Process 600, and / or other processes described herein) may be executed under the control of one or more computer systems configured with executable instructions, or implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes collectively on one or more processors, by hardware, or by a combination thereof. As described above, the code may be stored on a computer-readable storage medium or a machine-readable storage medium in the form of, for example, a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable storage medium or the machine-readable storage medium may be non-transitory.
[0202] The coding techniques described herein may be implemented in an exemplary video encoding and decoding system (e.g., system 100). In some examples, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device may comprise any of a wide range of devices including desktop computers, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device and the destination device may be equipped for wireless communication.
[0203] The destination device may receive the encoded video data to be decoded via a computer-readable medium. The computer-readable medium may comprise any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium may comprise a communication medium for enabling the source device to transmit the encoded video data directly to the destination device in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device. The communication medium may comprise any wireless communication medium or wired communication medium such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device that may be useful for facilitating communication from the source device to the destination device.
[0204] In some examples, the encoded data can be output from an output interface to a storage device. Similarly, the encoded data can be accessed from the storage device by an input interface. The storage device can include any of a variety of data storage media that are distributed or locally accessible, such as a hard drive, Blu-ray disk, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In further examples, the storage device can correspond to a file server or another intermediate storage device that can store the encoded video generated by a source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device. Exemplary file servers can include a web server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device can access the encoded video data through any standard data connection including an Internet connection. This can include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, which is suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device can be a streaming transmission, a download transmission, or a combination thereof.
[0205] The techniques of the present disclosure are not necessarily limited to wireless applications or settings. The techniques can be applied to video coding that supports any of various multimedia applications such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0206] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device can include an input interface, a video decoder, and a display device. The video encoder of the source device can be configured to apply the techniques disclosed herein. In other examples, the source device and the destination device can include other components or configurations. For example, the source device can receive video data from an external video source such as an external camera. Similarly, the destination device can interface with an external display device rather than including an integrated display device.
[0207] The above exemplary system is merely one example. Techniques for processing video data in parallel can be performed by any digital video encoding and / or decoding device. In general, the techniques of the present disclosure are performed by a video encoding device, but the techniques can also be performed by a video encoder / decoder, commonly referred to as a "codec". Additionally, the techniques of the present disclosure can also be performed by a video preprocessor. The source device and the destination device are merely examples of such coding devices that generate encoded video data for the source device to transmit to the destination device. In some examples, the source device and the destination device can operate substantially symmetrically such that each device includes video encoding and decoding components. Thus, the exemplary system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0208] The video source can include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source can generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In some cases, when the video source is a video camera, the source device and the destination device can form a so-called camera phone or video phone. However, as described above, the techniques described in the present disclosure can generally be applicable to video coding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by a video encoder. The encoded video information can be output onto a computer-readable medium by an output interface.
[0209] As described above, a computer-readable medium can include a transient medium such as a wireless broadcast or a wired network transmission, or a storage medium (i.e., a non-transient storage medium) such as a hard disk, a flash drive, a compact disk, a digital video disk, a Blu-ray disk, or other computer-readable media. In some examples, a network server (not shown) can receive encoded video data from a source device via, for example, a network transmission and provide the encoded video data to a destination device. Similarly, a computing device of a media manufacturing facility such as a disk stamping facility can receive encoded video data from a source device and manufacture a disk that includes the encoded video data. Thus, a computer-readable medium can be understood to include one or more computer-readable media in various forms in various examples.
[0210] An input interface of a destination device receives information from a computer-readable medium. The information of the computer-readable medium can include syntax information defined by a video encoder that includes blocks and other coded units, such as characteristics and / or processing of a picture group (GOP), and / or syntax elements that describe the processing, and the syntax information is also used by a video decoder. A display device displays the decoded video data to a user and can include any of various display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device. Various embodiments of the present application have been described.
[0211] The specific details of the symbolization device 104 and the decoding device 112 are shown in FIGS. 7 and 8, respectively. FIG. 7 is a block diagram showing an exemplary symbolization device 104 that may implement one or more of the techniques described in the present disclosure. The symbolization device 104 may generate, for example, a syntax structure (e.g., a syntax structure of a VPS, SPS, PPS, or other syntax element) described herein. The symbolization device 104 may perform intra-prediction coding and inter-prediction coding of video blocks within a video slice. As previously explained, intra-coding at least partially relies on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-coding at least partially relies on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. The intra mode (I mode) may refer to any of several spatial-based compression modes. Inter modes such as unidirectional prediction (P mode) or bi-prediction (B mode) may refer to any of several temporal-based compression modes.
[0212] The symbolization device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, an adder 50, an inverse transformation processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra prediction processing unit 46. For video block reconstruction, the symbolization device 104 also includes an inverse quantization unit 58, an inverse transformation processing unit 60, and an adder 62. The filter unit 63 is intended to represent one or more loop filters such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 63 is shown in FIG. 7 as an in-loop filter, in other configurations, the filter unit 63 may be implemented as a post-loop filter. The post-processing device 57 may perform additional processing on the encoded video data generated by the symbolization device 104. The techniques of the present disclosure may be implemented by the symbolization device 104 in some cases. However, in other cases, one or more of the techniques of the present disclosure may be implemented by the post-processing device 57.
[0213] As shown in FIG. 7, the encoding device 104 receives video data, and the segmentation unit 35 segments the data into video blocks. Segmenting may also include segmenting into slices, slice segments, tiles, or other larger units, and may include, for example, video block segmentation according to the quadtree structure of LCU and CU. The encoding device 104 generally shows components that encode video blocks within a video slice to be encoded. A slice may be divided into a plurality of video blocks (and, in some cases, a set of video blocks called tiles). The prediction processing unit 41 may select, for the current video block, one of a plurality of possible coding modes, such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes, based on error results (such as coding rate and level of distortion). The prediction processing unit 41 may provide the obtained intra-coded block or inter-coded block to the adder 50 to generate residual block data and to the adder 62 to reconstruct an encoded block for use as a reference picture.
[0214] The intra prediction processing unit 46 within the prediction processing unit 41 may perform intra prediction coding of the current video block with respect to one or more adjacent blocks in the same frame or slice as the current block to be coded for spatial compression. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter prediction coding of the current video block with respect to one or more prediction blocks in one or more reference pictures for temporal compression.
[0215] The motion estimation unit 42 may be configured to determine an inter-prediction mode for a video slice according to a predetermined pattern for a video sequence. The predetermined pattern may specify video slices in the sequence as P slices, B slices, or GPB slices. The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion of a video block. The motion vector may indicate, for example, the displacement of a prediction unit (PU) of a video block in the current video frame or picture relative to a prediction block in a reference picture.
[0216] The prediction block is a block that has been found to exactly match the PU of the video block to be coded with respect to pixel differences that may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some examples, the encoding device 104 may calculate values for sub-integer pixel positions of the reference picture stored in the picture memory 64. For example, the encoding device 104 may interpolate values at 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation unit 42 may perform motion search for full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0217] The motion estimation unit 42 calculates a motion vector for the PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a prediction block in a reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1), each of which identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy encoding unit 56 and the motion compensation unit 44.
[0218] The motion compensation performed by the motion compensation unit 44 may involve fetching or generating a prediction block based on a motion vector determined by motion estimation that, in some cases, performs interpolation to sub-pixel accuracy. Upon receiving a motion vector for the PU of the current video block, the motion compensation unit 44 may identify the position of the prediction block pointed to by the motion vector within the reference picture list. The encoding device 104 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the currently encoded video block to form a pixel difference value. The pixel difference value forms the residual data for the block and may include both a luma difference component and a chroma difference component. The adder 50 represents one or more components that perform this subtraction operation. The motion compensation unit 44 may also generate syntax elements associated with the video block and the video slice for use by the decoding device 112 when decoding the video block of the video slice.
[0219] As described above, the intra prediction processing unit 46 can perform intra prediction of the current block as an alternative to inter prediction executed by the motion estimation unit 42 and the motion compensation unit 44. Specifically, the intra prediction processing unit 46 can determine an intra prediction mode to be used for encoding the current block. In some examples, the intra prediction processing unit 46 may encode the current block using various intra prediction modes, for example, between different encoding paths, and the intra prediction processing unit 46 may select an appropriate intra prediction mode to be used from the tested modes. For example, the intra prediction processing unit 46 may calculate rate distortion values using rate distortion analysis for various tested intra prediction modes, and may select an intra prediction mode having the best rate distortion characteristics from among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block. The intra prediction processing unit 46 can calculate a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate distortion value for the block.
[0220] In any case, after selecting an intra prediction mode for a block, the intra prediction processing unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data the definition of encoding contexts for various blocks, as well as the indication of the most accurate intra prediction mode, the intra prediction mode index table, and the modified intra prediction mode index table to be used for each of the contexts. The bitstream configuration data may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables).
[0221] After the prediction processing unit 41 generates a prediction block for the current video block via either inter prediction or intra prediction, the encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 uses a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to transform the residual video data into residual transform coefficients. The transform processing unit 52 may convert the residual video data from a pixel domain to a transform domain such as a frequency domain.
[0222] The transform processing unit 52 may send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be adjusted by adjusting the quantization parameter. In some examples, the quantization unit 54 may perform a scan of a matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 56 may perform the scan.
[0223] Following quantization, entropy encoding unit 56 entropy-encodes the quantized transform coefficients. For example, entropy encoding unit 56 may perform context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding technique. Following the entropy encoding by entropy encoding unit 56, the encoded bitstream may be transmitted to decoding device 112 or may be archived for later transmission or retrieval by decoding device 112. Entropy encoding unit 56 may also entropy-encode motion vectors and other syntax elements for the current video slice being coded.
[0224] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel region for later use as a reference block of the reference picture. Motion compensation unit 44 may calculate the reference block by adding the residual block to a prediction block of one of the reference pictures in the reference picture list. Motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-pixel values for use in motion estimation. Adder 62 adds the reconstructed residual block to the motion compensation prediction block generated by motion compensation unit 44 to generate a reference block for storage in picture memory 64. The reference block may be used by motion estimation unit 42 and motion compensation unit 44 as a reference block for inter-predicting blocks in subsequent video frames or pictures.
[0225] In this way, the encoding device 104 of FIG. 7 represents an example of a video encoder configured to execute any of the techniques described herein, including any of the processes or techniques described above. In some cases, some of the techniques of the present disclosure may also be implemented by the post-processing device 57.
[0226] FIG. 8 is a block diagram showing an exemplary decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra prediction processing unit 84. The decoding device 112 may execute a decoding path generally inverse to the encoding path described with respect to the encoding device 104 from FIG. 7 in some examples.
[0227] During the decoding process, the decoding device 112 receives an encoded video bitstream representing video blocks of the encoded video slices and associated syntax elements sent by the encoding device 104. In some embodiments, the decoding device 112 may receive the encoded video bitstream from the encoding device 104. In some embodiments, the decoding device 112 may receive the encoded video bitstream from a network entity 79 such as a server, a media recognition network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the techniques described above. The network entity 79 may or may not include the encoding device 104. Some of the techniques described in this disclosure may be implemented by the network entity 79 before the network entity 79 transmits the encoded video bitstream to the decoding device 112. In some video decoding systems, the network entity 79 and the decoding device 112 may be part of separate devices, but in other cases, the functions described with respect to the network entity 79 may be performed by the same device that includes the decoding device 112.
[0228] The entropy decoding unit 80 of the decoding device 112 entropy decodes the bitstream to generate quantization coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 transfers the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 may receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 may process and parse both fixed-length syntax elements and variable-length syntax elements in one or more parameter sets such as VPS, SPS, and PPS.
[0229] When a video slice is coded as an intra-coded (I) slice, the intra prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video blocks of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current frame or picture. When a video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for the video blocks of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be generated from one of the reference pictures within the reference picture list. The decoding device 112 can construct the reference frame lists, namely, list 0 and list 1, using a default construction technique based on the reference pictures stored in the picture memory 92.
[0230] The motion compensation unit 82 determines prediction information for the video blocks of the current video slice by parsing the motion vector and other syntax elements, and uses the prediction information to generate a prediction block for the currently decoded video block. For example, the motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra prediction or inter prediction) used to code the video blocks of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information for one or more reference picture lists for the slice, the motion vector for each inter-coded video block of the slice, the inter prediction status for each inter-coded video block of the slice, and other information for decoding the video blocks within the current video slice.
[0231] The motion compensation unit 82 may also perform interpolation based on an interpolation filter. The motion compensation unit 82 may use an interpolation filter as used by the encoding device 104 during the encoding of a video block to calculate interpolation values for sub-integer pixels of a reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the encoding device 104 from the received syntax elements and may use that interpolation filter to generate a prediction block.
[0232] The inverse quantization unit 86 inverse quantizes or de-quantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include using quantization parameters calculated by the encoding device 104 for each video block in a video slice to determine the degree of quantization and, similarly, to determine the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to generate a residual block in the pixel domain.
[0233] After the motion compensation unit 82 generates a prediction block for the current video block based on the motion vector and other syntax elements, the decoding device 112 forms a decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82. The adder 90 represents one or more components that perform this addition operation. If desired, a loop filter (either within the coding loop or after the coding loop) may also be used to smooth pixel transitions or improve video quality in other ways. The filter unit 91 is intended to represent one or more loop filters such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 91 is shown in FIG. 8 as being an in-loop filter, in other configurations, the filter unit 91 may be implemented as a post-loop filter. The decoded video blocks within a given frame or picture are stored in the picture memory 92, which stores the reference pictures used for subsequent motion compensation. The picture memory 92 also stores the decoded video for later presentation on a display device such as the video destination device 119 shown in FIG. 1A.
[0234] In this way, the decoding device 112 of FIG. 8 represents an example of a video decoder configured to perform any of the techniques described herein, including the processes or techniques described above.
[0235] The techniques of the present disclosure are not necessarily limited to wireless applications or settings. The techniques can be applied to video coding that supports any of various multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0236] As used herein, the term "computer-readable medium" includes, without limitation, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium includes non-transitory media that can store data and does not include carrier waves and / or transient electronic signals that propagate wirelessly or via a wired connection. Examples of non-transitory media can include, without limitation, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. The computer-readable medium may store code and / or machine-executable instructions that can represent any combination of procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, transferred, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0237] In some embodiments, a computer-readable storage device, medium, and memory can include a cable signal or wireless signal including a bitstream, etc. However, when mentioned, non-transitory computer-readable storage media inherently and expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals.
[0238] To provide a complete understanding of the embodiments and examples provided in this specification, specific details are provided in the above description. However, it will be understood by those skilled in the art that the embodiments can be practiced without these specific details. For the sake of clarity, in some cases, the present technology may be presented as including individual functional blocks that include a device, device components, and steps or routines in a method embodied in software, or a combination of hardware and software. Additional components other than those shown in the figures and / or described in this specification may be used. For example, in order not to obscure the embodiments with unnecessary details, circuits, systems, networks, processes, or other components may be shown as components in the form of block diagrams. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details in order to avoid obscuring the embodiments.
[0239] Individual embodiments may be described above as a process or method shown as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. A flowchart may describe operations as a sequential process, but many of the operations may be executed in parallel or simultaneously. In addition, the order of the operations may be rearranged. A process ends when its operations are completed, but may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its end may correspond to the function returning to the calling function or main function.
[0240] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored or otherwise available from a computer-readable medium. Such instructions can cause, for example, a general-purpose computer, a special-purpose computer, or a processing device to execute a particular function or group of functions, or otherwise configure a general-purpose computer, special-purpose computer, or processing device to execute a particular function or group of functions, and can include instructions and data. Portions of the computer resources used may be accessible via a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store the instructions, the information used, and / or the information created during the methods according to the examples described include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-connected storage devices, and the like.
[0241] Devices implementing the processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and can take any of various form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks can be stored on a computer-readable or machine-readable medium. A processor can execute the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, and the like. The functions described herein can also be embodied in a peripheral device or an add-in card. Such functions can also be implemented, as a further example, on a circuit board within different chips or different processes executing in a single device.
[0242] Commands, media for conveying such commands, computing resources for executing the commands, and other structures for supporting such computing resources are exemplary means for providing the functions described in the present disclosure.
[0243] In the above description, aspects of the present application have been described with respect to their specific embodiments, but those skilled in the art will recognize that the present application is not limited thereto. Accordingly, although exemplary embodiments of the present application have been described in detail herein, it is to be understood that the inventive concept can be embodied and adopted in various other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. It is to be understood that the various features and aspects of the application examples described above can be used individually or together. Further, the embodiments can be utilized in any number of environments and application examples other than those described herein without departing from the broader spirit and scope of this specification. Accordingly, this specification and the drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, the method has been described in a particular order. It is to be understood that in alternative embodiments, the method can be performed in an order different from that described.
[0244] Those skilled in the art will understand that the symbols or terms “less than” (“<”) and “greater than” (">") used herein can be replaced, without departing from the scope of this specification, by the symbols “less than or equal to” (“≦”) and “greater than or equal to” (“≧”), respectively.
[0245] When an element is described as being "configured to" perform some operations, such a configuration can be achieved, for example, by designing an electronic circuit or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor, or other suitable electronic circuit) to perform the operations, or by any combination thereof.
[0246] The phrase "coupled to" refers to any element that is physically connected, either directly or indirectly, to another element, and / or that communicates, either directly or indirectly, with another element (e.g., is connected to another element via a wired or wireless connection and / or other suitable communication interface).
[0247] Claims language or other language reciting a set "of at least one" and / or "one or more" of a set indicates that one member of the set or multiple members (in any combination) of the set satisfy the claim. For example, claims language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claims language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "of at least one" of a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claims language reciting "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and can include as additional items those not listed in the set of A and B.
[0248] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0249] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, including but not limited to general-purpose computers, wireless communication device handsets, or integrated circuit devices having multiple applications including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. When implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product that may include packaging material. The computer-readable medium may comprise a memory or data storage media, such as random access memory (RAM), such as synchronous dynamic random access memory (SDRAM), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. The techniques may be realized at least in part by a computer-readable communication medium that conveys or communicates program code in the form of an instruction or data structure as a propagated signal or wave, accessible, readable, and / or executable by a computer.
[0250] The program code can be executed by a processor that can include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated logic circuits or discrete logic circuits. Such a processor can be configured to execute any of the techniques described in this disclosure. The general-purpose processor can be a microprocessor, but alternatively, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration. Thus, the term "processor" as used herein can refer to any of the above structures, any combination of the above structures, or any other structure or device suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein may be provided within dedicated software modules or hardware modules configured for encoding and decoding, or may be incorporated into a composite video encoder decoder (codec).
[0251] Exemplary aspects of the present disclosure include the following.
[0252] Aspect 1. An apparatus for decoding video data, comprising a memory and at least one processor coupled to the memory (e.g., implemented in a circuit). The at least one processor is configured to obtain a video bitstream, wherein the video bitstream includes adaptive loop filter (ALF) data; determine a value of an ALF chroma filter signal flag from the ALF data, wherein the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream; and process at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0253] Aspect 2. The apparatus of Aspect 1, wherein the at least one processor is further configured to obtain a slice header of a slice of the video data from the video bitstream; determine a value of an ALF chroma identifier from the slice header, wherein the value of the ALF chroma identifier indicates whether ALF can be applied to one or more chroma components of the slice; and process at least a portion of a slice of the video data based on the ALF chroma identifier from the slice header.
[0254] Aspect 3. The apparatus of Aspect 2, wherein the at least one processor is further configured to determine a value of a chroma format identifier from the slice header, wherein the value of the chroma format identifier and the value of the ALF chroma identifier indicate which chroma component(s) of one or more chroma components ALF is applicable to.
[0255] Aspect 4. The apparatus according to any one of Aspects 1 to 3, wherein the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream, and the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of a slice of the video data.
[0256] Aspect 5. At least one processor is configured to obtain chroma ALF filter data to be used to process at least a portion of a slice of video data based on the value of an ALF chroma filter signal flag, and apply the chroma ALF filter data to at least a portion of a slice of video data, for a device according to any one of Aspects 1 to 4.
[0257] Aspect 6. At least one processor is configured to presume that the value of the ALF chroma filter signal flag is 0 when the value of the ALF chroma filter signal flag does not exist in the ALF data, for a device according to any one of Aspects 1 to 5.
[0258] Aspect 7. At least one processor is configured to obtain luma ALF filter data to be used for one or more chroma components of at least one block of a video bitstream based on the value of an ALF chroma filter signal flag, and apply the luma ALF filter data to one or more chroma components of at least one block of a video bitstream, for a device according to any one of Aspects 1 to 6.
[0259] Aspect 8. At least one processor is configured to obtain a slice header of a slice of video data from a video bitstream, determine a value of a chroma format identifier from the slice header, and process one or more chroma components of at least one block of a video bitstream using luma ALF filter data based on the value of the chroma format identifier from the slice header, for a device according to any one of Aspects 1 to 7.
[0260] Aspect 9. At least one processor is further configured to process the value of an ALF chroma filter signal flag from the ALF data to determine that chroma ALF filter data is signaled in the video bitstream, for a device according to any one of Aspects 1 to 8.
[0261] Aspect 10. The apparatus according to any one of Aspects 1 to 9, wherein at least one processor is further configured to determine an ALF application parameter set (APS) identifier for a first color component of at least a portion of a slice and to determine an ALF map for the first color component of at least a portion of the slice.
[0262] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein at least one processor is further configured to enable ALF filtering for at least two non-luma components of at least a portion of a slice based on the fact that the components of at least a portion of the slice include a shared characteristic.
[0263] Aspect 12. The apparatus of Aspect 11, wherein the at least two non-luma components of at least a portion of the slice include a red component, a green component, and a blue component of at least a portion of the slice.
[0264] Aspect 13. The apparatus of Aspect 11, wherein the at least two non-luma components of at least a portion of the slice include a chroma component of at least a portion of the slice.
[0265] Aspect 14. The apparatus according to any one of Aspects 1 to 13, wherein at least a portion of a slice of video data includes 4:4:4 format video data.
[0266] Aspect 15. The apparatus according to any one of Aspects 1 to 14, wherein at least one processor is configured to enable ALF filtering for at least two non-luma components of at least a portion of a slice based on the fact that at least a portion of the slice includes non-4:2:0 format video data.
[0267] Aspect 16. A device according to any of Aspects 1 to 15, wherein at least one processor is configured to determine a chroma type array variable for at least a portion of a slice, determine an ALF chroma application parameter set (APS) identifier for a first component of at least a portion of the slice based on the chroma type array variable for at least a portion of the slice, and determine a signaled ALF map for the first component of at least a portion of the slice.
[0268] Aspect 17. The device of Aspect 16, wherein at least one processor is further configured to determine a second signaled ALF map for a second component of at least a portion of the slice based on the chroma type array variable.
[0269] Aspect 18. The device of Aspect 17, wherein at least one processor is configured to perform ALF filtering on the first and second components of at least a portion of the slice using the signaled ALF map and the second signaled ALF map.
[0270] Aspect 19. The device according to either of Aspects 17 or 18, wherein at least one processor is further configured to determine a third signaled ALF map for a third component of at least a portion of the slice based on the chroma type array variable.
[0271] Aspect 20. The device of Aspect 19, wherein the first component is a luminance component, the second component is a first chroma component, and the third component is a second chroma component.
[0272] Aspect 21. The device of Aspect 19, wherein the first component is a red component, the second component is a green component, and the third component is a blue component.
[0273] Aspect 22. A device according to any of Aspects 16 to 21, wherein at least one processor is configured to perform ALF processing on a block for each component of at least a portion of the slice based on the chroma type array variable.
[0274] Aspect 23. An apparatus according to any one of Aspects 1 to 22, comprising a mobile device.
[0275] Aspect 24. An apparatus according to any one of Aspects 1 to 23, further comprising a display configured to display one or more images.
[0276] Aspect 25. A method of decoding video data, comprising: obtaining a video bitstream, wherein the video bitstream includes adaptive loop filter (ALF) data; determining a value of an ALF chroma filter signal flag from the ALF data, wherein the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream; and processing at least a portion of a slice of the video data based on the value of the ALF chroma filter signal flag.
[0277] Aspect 26. The method of Aspect 25, further comprising: obtaining a slice header of a slice of video data from the video bitstream; determining a value of an ALF chroma identifier from the slice header, wherein the value of the ALF chroma identifier indicates whether ALF can be applied to one or more chroma components of the slice; and processing at least a portion of the slice based on the ALF chroma identifier from the slice header.
[0278] Aspect 27. The method of Aspect 26, further comprising: determining a value of a chroma format identifier from the slice header, wherein the value of the chroma format identifier and the value of the ALF chroma identifier indicate which chroma component of one or more chroma components ALF is applicable to.
[0279] Aspect 28. A method according to any of aspects 25 to 27, wherein the value of the ALF chroma filter signal flag indicates that the chroma ALF filter data is signaled in the video bitstream, and the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of a slice.
[0280] Aspect 29. A method according to any of aspects 25 to 28, further comprising obtaining chroma ALF filter data to be used for processing at least a portion of a slice of video data based on the value of the ALF chroma filter signal flag, and applying the chroma ALF filter data to at least a portion of the slice of video data.
[0281] Aspect 30. A method according to any of aspects 25 to 29, further comprising presuming that the value of the ALF chroma filter signal flag is 0 when the value of the ALF chroma filter signal flag does not exist in the ALF data.
[0282] Aspect 31. A method according to any of aspects 25 to 30, further comprising obtaining luma ALF filter data to be used for one or more chroma components of at least one block of the video bitstream based on the value of the ALF chroma filter signal flag, and applying the luma ALF filter data to one or more chroma components of at least one block of the video bitstream.
[0283] Aspect 32. A method according to any of aspects 25 to 31, further comprising obtaining a slice header of a slice of video data from the video bitstream, determining a value of a chroma format identifier from the slice header, and processing one or more chroma components of at least one block of the video bitstream using the luma ALF filter data based on the value of the chroma format identifier from the slice header.
[0284] Method according to any of aspects 25 to 32, further comprising the step of processing the value of the ALF chroma filter signal flag from the ALF data to determine that the chroma ALF filter data is signaled in the video bitstream.
[0285] Method according to any of aspects 25 to 33, further comprising the step of determining an ALF application parameter set (APS) identifier for a first color component of at least a portion of the slice and the step of determining an ALF map for the first color component of at least a portion of the slice.
[0286] Method according to any of aspects 25 to 34, further comprising the step of enabling ALF filter processing for at least two non-luma components of at least a portion of the slice based on the fact that at least a portion of the components of the slice include a common characteristic.
[0287] Method of aspect 35, wherein at least two non-luma components of at least a portion of the slice include a red component, a green component, and a blue component of at least a portion of the slice.
[0288] Method of aspect 35, wherein at least two non-luma components of at least a portion of the slice include a chroma component of at least a portion of the slice.
[0289] Method according to any of aspects 25 to 37, wherein at least a portion of the slice includes 4:4:4 format video data.
[0290] Method according to any of aspects 25 to 38, further comprising the step of enabling ALF filter processing for at least two non-luma components of at least a portion of the slice based on the fact that at least a portion of the slice includes non-4:2:0 format video data.
[0291] Aspect 40. A method according to any of aspects 25 to 39, further comprising: determining a chroma type array variable for at least a portion of a slice; determining an ALF chroma application parameter set (APS) identifier for a first component of at least a portion of the slice based on the chroma type array variable for at least a portion of the slice; and determining a signaled ALF map for the first component of at least a portion of the slice.
[0292] Aspect 41. The method of aspect 40, further comprising determining a second signaled ALF map for a second component of at least a portion of the slice based on the chroma type array variable.
[0293] Aspect 42. The method of aspect 41, further comprising performing ALF filtering on the first and second components of at least a portion of the slice using the signaled ALF map and the second signaled ALF map.
[0294] Aspect 43. The method according to any of aspects 41 or 42, further comprising determining a third signaled ALF map for a third component of at least a portion of the slice based on the chroma type array variable.
[0295] Aspect 44. The method of aspect 43, wherein the first component is a luma component, the second component is a first chroma component, and the third component is a second chroma component.
[0296] Aspect 45. The method of aspect 43, wherein the first component is a red component, the second component is a green component, and the third component is a blue component.
[0297] Aspect 46. A method according to any of aspects 25 to 45, further comprising performing ALF processing on a block for each component of at least a portion of the slice based on the chroma type array variable.
[0298] Aspect 47. An apparatus for encoding video data, the apparatus comprising a memory and at least one processor coupled to the memory (e.g., implemented in a circuit). The at least one processor is configured to generate adaptive loop filter (ALF) data, determine a value of an ALF chroma filter signal flag of the ALF data, wherein the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in a video bitstream, and generate a video bitstream including the ALF data.
[0299] Aspect 48. The apparatus of aspect 47, wherein the at least one processor is further configured to determine a value of an ALF chroma identifier, wherein the value of the ALF chroma identifier indicates whether the ALF can be applied to one or more chroma components of a slice of video data, and include the value of the ALF chroma identifier in a slice header of the video bitstream.
[0300] Aspect 49. The apparatus of aspect 48, wherein the at least one processor is further configured to determine a value of a chroma format identifier, wherein the value of the chroma format identifier and the value of the ALF chroma identifier indicate to which chroma component of one or more chroma components the ALF is applicable, and include the value of the chroma format identifier in a slice header of the video bitstream.
[0301] Aspect 50. The apparatus according to any one of aspects 47 to 49, wherein the value of the ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in the video bitstream, and the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of the slice.
[0302] Aspect 51. An apparatus according to any of aspects 47 to 50, wherein at least one processor is configured to determine a value of a chroma format identifier, the value of the chroma format identifier indicating one or more chroma components of at least one block of a video bitstream to be processed using luma ALF filter data, and to include the value of the chroma format identifier in a slice header of the video bitstream.
[0303] Aspect 52. An apparatus according to any of aspects 47 to 51, comprising a mobile device.
[0304] Aspect 53. An apparatus according to any of aspects 47 to 52, further comprising a display configured to display one or more images.
[0305] Aspect 54. A method of encoding video data, the method comprising: generating adaptive loop filter (ALF) data; determining a value of an ALF chroma filter signal flag of the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in a video bitstream; and generating a video bitstream including the ALF data.
[0306] Aspect 55. The method of aspect 54, further comprising: determining a value of an ALF chroma identifier, the value of the ALF chroma identifier indicating whether ALF can be applied to one or more chroma components of a slice of video data; and including the value of the ALF chroma identifier in a slice header of the video bitstream.
[0307] Aspect 56. A method according to aspect 55, further comprising: determining a value of a chroma format identifier, wherein the value of the chroma format identifier and the value of an ALF chroma identifier indicate to which chroma component among one or more chroma components ALF is applicable; and including the value of the chroma format identifier in a slice header of a video bitstream.
[0308] Aspect 57. A method according to any one of aspects 54 to 56, wherein a value of an ALF chroma filter signal flag indicates that chroma ALF filter data is signaled in a video bitstream, and the chroma ALF filter data is signaled in an adaptive parameter set (APS) for processing at least a portion of a slice.
[0309] Aspect 58. A method according to any one of aspects 54 to 57, further comprising: determining a value of a chroma format identifier, wherein the value of the chroma format identifier indicates one or more chroma components of at least one block of a video bitstream to be processed using luma ALF filter data; and including the value of the chroma format identifier in a slice header of the video bitstream.
[0310] Aspect 59. A computer-readable storage medium that, when executed by one or more processors, stores instructions that cause the one or more processors to perform any of the operations of aspects 1 to 46.
[0311] Aspect 60. An apparatus comprising means for performing any of the operations of aspects 1 to 46.
[0312] Aspect 61. A computer-readable storage medium that, when executed by one or more processors, stores instructions that cause the one or more processors to perform any of the operations of aspects 47 to 58.
[0313] Aspect 62. An apparatus comprising means for performing any of the operations of aspects 47 to 58.
[0314] Aspect 63. A computer-readable storage medium that stores instructions which, when executed by one or more processors, cause the one or more processors to execute any of the operations of Aspects 1 to 58.
[0315] Aspect 64. An apparatus comprising means for performing any of the operations of Aspects 1 to 58.
Explanation of Signs
[0316] 35 Division Unit 41 Prediction Processing Unit 42 Motion Estimation Unit 44 Motion Compensation Unit 46 Intra Prediction Processing Unit 50 Adder 52 Transformation Processing Unit 54 Quantization Unit 56 Entropy Encoding Unit 57 Post-Processing Device 58 Inverse Quantization Unit 60 Inverse Transformation Processing Unit 62 Adder 63 Filter Unit 64 Picture Memory 79 Network Entity 80 Entropy Decoding Unit 81 Prediction Processing Unit 82 Motion Compensation Unit 84 Intra Prediction Processing Unit 86 Inverse Quantization Unit 88 Inverse Transformation Processing Unit 90 Adder 91 Filter Unit 92 Picture Memory 100 System 102 Video Source 104 Encoding Device, Video Encoding Device 106 Encoder Engine 108 Storage 110 Output unit 112 Decoding device, video decoding device 114 Input unit 116 Decoder engine 118 Storage 119 Video destination device 120 Communication link 121 Picture 122 Filter unit 124 Deblocking filter 126 Sample Adaptive Offset (SAO) filter, SAO filter 128 Adaptive Loop Filter (ALF) / Geometry Transformation-based Adaptive Loop Filter (GALF) filter, ALF filter 400 Process 500 Process 600 Process< / highlightend> < / highlight>
Claims
1. An apparatus for decoding video data, comprising: a memory; at least one processor coupled to the memory; wherein the at least one processor is configured to: obtain a video bitstream, the video bitstream including adaptive loop filter (ALF) data; determine a value of an ALF chroma filter signal flag from the ALF data, the value of the ALF chroma filter signal flag indicating whether chroma ALF filter data is signaled in the video bitstream; process at least a portion of a slice of video data including 4:4:4 format video data or 4:2:2 format video data based on the value of the ALF chroma filter signal flag; and processing at least a portion of the slice of video data including 4:4:4 format video data or 4:2:2 format video data includes, after determining that the chroma ALF filter data is signaled in the video bitstream based on the value of the ALF chroma filter signal flag, enabling ALF filter processing for at least two chroma components of at least the portion of the slice, in response to at least the portion of the slice including 4:4:4 format video data or 4:2:2 format video data. The apparatus according to claim 1.
2. Enabling ALF filter processing for at least two chroma components of at least the portion of the slice includes enabling each chroma component to reference a separate adaptive parameter set (APS) to access a filter bank for performing ALF filter processing on the chroma component. The apparatus according to claim 1.
3. At least a portion of the slice of video data includes 4:4:4 format video data, and enabling ALF filter processing for the at least two chroma components is in accordance with at least said portion of the slice including 4:4:4 format video data, the apparatus according to claim 1.
4. The at least one processor obtains a slice header of the slice of video data from the video bitstream, determines a value of an ALF chroma identifier from the slice header, the value of the ALF chroma identifier indicating whether ALF can be applied to one or more chroma components of the slice, determines a value of a chroma format identifier from the slice header, the value of the chroma format identifier and the value of the ALF chroma identifier indicating to which chroma component among the one or more chroma components ALF is applicable, processes at least said portion of the slice of video data based on the ALF chroma identifier from the slice header and is further configured to perform, the apparatus according to claim 1.
5. The at least one processor obtains a slice header of the slice of video data from the video bitstream, determines a value of a chroma format identifier from the slice header, and processes one or more chroma components of at least one block of the video bitstream using luma ALF filter data based on the value of the chroma format identifier from the slice header and is configured to perform, the apparatus according to claim 1.
6. The at least one processor Determine a separate ALF map for each chroma component of at least said portion of said slice The apparatus according to claim 1, further configured as such. **Claim 7** Said at least one processor Determine a chroma type array variable for at least said portion of said slice, said chroma type array variable indicating the chroma format of video data, Determine an ALF chroma application parameter set (APS) identifier for a first component of at least said portion of said slice based on said chroma type array variable for at least said portion of said slice, Determine a signaled ALF map for said first component of at least said portion of said slice The apparatus according to claim 1, configured as such. **Claim 8** Said at least one processor Determine a second signaled ALF map for a second component of at least said portion of said slice based on said chroma type array variable The apparatus according to claim 7, further configured as such. **Claim 9** Said at least one processor is configured to perform ALF filtering on said first component and said second component of at least said portion of said slice using said signaled ALF map and said second signaled ALF map, the apparatus according to claim 8. **Claim 10** Said at least one processor Determine a third signaled ALF map for a third component of at least said portion of said slice based on said chroma type array variable The apparatus according to claim 8, further configured as such. **Claim 11** The apparatus according to claim 10, wherein the first component is a luma component, the second component is a first chroma component, and the third component is a second chroma component.
12. The apparatus according to claim 1, comprising a mobile device.
13. The apparatus according to claim 1, further comprising a display configured to display one or more images.
14. The apparatus according to claim 1, wherein the slice of the video data is MPEG5 Essential Video Coding video data.
15. A method for decoding video data, comprising: obtaining a video bitstream, wherein the video bitstream includes Adaptive Loop Filter (ALF) data; determining a value of an ALF chroma filter signal flag from the ALF data, wherein the value of the ALF chroma filter signal flag indicates whether chroma ALF filter data is signaled in the video bitstream; processing at least a portion of a slice of video data including 4:4:4 format video data or 4:2:2 format video data based on the value of the ALF chroma filter signal flag; and processing at least a portion of the slice of video data including 4:4:4 format video data or 4:2:2 format video data includes enabling ALF filter processing for at least two chroma components of at least the portion of the slice in response to determining that the chroma ALF filter data is signaled in the video bitstream based on the value of the ALF chroma filter signal flag and that at least the portion of the slice includes 4:4:4 format video data or 4:2:2 format video data.
Citation Information
Patent Citations
Adaptive loop filtering for chromatic components
JP2014533012A
Temporal parameter prediction in nonlinear adaptive loop filters.
JP2022526633A