Systems and methods for signaling downsampling offset information in video coding
Patent Information
- Application Number
- JP2023011093
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-02
- Filing Date
- 2023-01-27
- Publication Date
- 2025-12-24
AI Technical Summary
Existing video encoding standards, such as HEVC and emerging standards like VVC, lack efficient methods for signaling downsampling offset information, leading to potential phase shifts and alignment issues during image rendering.
The proposed technique involves signaling phase indication information through a syntax element that specifies the horizontal and vertical positions of luminance sampling relative to a display window, using specific mathematical operations to adjust the position values, ensuring accurate alignment and reducing phase shifts during downsampling.
This approach enhances the accuracy of video rendering by aligning downsampled images correctly, minimizing perceptible phase shifts and improving the overall quality of video playback across various devices and systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 418,494, filed October 22, 2022, which is incorporated by reference in its entirety.
[0002] FIELD This disclosure relates to video coding, and more particularly to techniques for signaling downsampling information for coded video. [Background technology]
[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, laptop or desktop computers, tablet computers, digital recording devices, digital media players, video gaming devices, cellular telephones, including so-called smart phones, medical imaging devices, and the like. Digital video can be encoded according to a video coding standard. Video coding standards define the format of a compliant bitstream that encapsulates the encoded video data. A compliant bitstream is a data structure that can be received and decoded by a video decoding device to generate reconstructed video data. Video coding standards can incorporate video compression techniques. Examples of video coding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High-Efficiency Video Coding (HEVC). HEVC is described in High Efficiency Video Coding (HEVC), Rec. ITU-T H.265 (December 2016), which is incorporated herein by reference and is referred to herein as ITU-T H.265. Extensions and improvements to ITU-T H.265 are currently being considered for the development of next-generation video coding standards. For example, the ITU-T Video Coding Experts Group (VCEG) and ISO / IEC (Moving Picture Experts Group (MPEG), collectively referred to as the Joint Video Exploration Team (JVET)) are working on a standardized video coding technique that has compression capabilities significantly exceeding those of the current HEVC standard.The Joint Exploration Model 7 (JEM 7), Algorithm Description of Joint Exploration Test Model 7 (JEM 7), ISO / IEC JTC1 / SC29 / WG11 Document: JVET-G1001, July 2017, Torino, IT, incorporated herein by reference, describes coding features that have been studied in a collaborative test model by JVET as having the potential to improve video coding technology beyond the capabilities of ITU-T H.265. It should be noted that the coding features of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM may collectively refer to the algorithms included in JEM 7 and the implementation of the JEM reference software. Additionally, in response to the "Joint Call for Proposals on Video Compression with Capabilities beyond HEVC" jointly issued by VCEG and MPEG, several specifications for video coding tools were proposed by various groups at the 10th Meeting of ISO / IEC JTC1 / SC29 / WG11 16-20 April 2018, San Diego, CA. From these descriptions of video coding tools, a draft video coding specification was developed and described in "Versatile Video Coding (Draft 1)", 10th Meeting of ISO / IEC JTC1 / SC29 / WG11 16-20 April 2018, San Diego, CA, document JVET-J1001-v2, which is incorporated herein by reference and referred to as JVET-J1001. The current development of the next generation video coding standard by VCEG and MPEG is referred to as the Versatile Video Coding (VVC) project.“Versatile Video Coding (Draft 10)”, 20th Meeting of ISO / IEC JTC1 / SC29 / WG11 7-16 October 2020, Teleconference, document JVET-T2001-v2, which is incorporated herein by reference and referred to as JVET-T2001, represents the current iteration of the draft video coding specification corresponding to the VVC project.
[0004] Video compression techniques allow for reducing data requirements for storing and transmitting video data. Video compression techniques can reduce data requirements by exploiting inherent redundancy in a video sequence. Video compression techniques may sequentially subdivide a video sequence into smaller portions (i.e., groups of pictures in a video sequence, pictures in groups of pictures, regions in pictures, sub-regions in regions, etc.). Intra-prediction coding techniques (e.g., spatial prediction techniques within a picture) and inter-prediction techniques (i.e., techniques between pictures (temporal)) may be used to generate difference values between a unit of video data to be coded and a reference unit of video data. The difference values may be referred to as residual data. The residual data may be coded as quantized transform coefficients. Syntax elements may associate the residual data with the reference coding units (e.g., intra-prediction mode index and motion information). The residual data and syntax elements may be entropy coded. The entropy coded residual data and syntax elements may be included in a data structure that forms a compliant bitstream. Summary of the Invention
[0005] Generally, this disclosure describes various techniques for encoding video data. Specifically, this disclosure describes techniques for signaling downsampling offset information for encoded video data. It should be noted that although the techniques of this disclosure are described with respect to ITU-T H.264, ITU-T H.265, JEM, and JVET-T2001, the techniques of this disclosure are generally applicable to video coding. For example, the encoding techniques described herein can be incorporated into video coding systems (including video coding systems based on future video coding standards) that include video block structures, intra-prediction techniques, inter-prediction techniques, transform techniques, filtering techniques, and / or entropy coding techniques other than those included in ITU-T H.265, JEM, and JVET-T2001. Thus, references to ITU-T H.264, ITU-T H.265, JEM, and / or JVET-T2001 are for illustrative purposes and should not be construed as limiting the scope of the technology described herein. Furthermore, it should be noted that the incorporation by reference of a document herein is for illustrative purposes and should not be construed as limiting or creating ambiguity with respect to the terms used herein. For example, if an incorporated reference provides a definition of a term that differs from that of another incorporated reference and / or as that term is used herein, that term should be construed to broadly include each corresponding definition and / or to include each specific definition instead.
[0006] In one example, a method of signaling phase indication information for video image data includes signaling a phase indication information message corresponding to an encoded video sequence; signaling within the phase indication information message a syntax element that specifies a horizontal position of a luminance sampling location relative to a display window, where the horizontal position is represented as a signaled value of subtracting 128 from the specified horizontal position, dividing the difference by the phase denominator, and adding 0.5 to the quotient; and signaling within the phase indication information message one or more syntax elements that specify a vertical position of a luminance sampling location relative to the display window, where the vertical position is represented as a signaled value of subtracting 128 from the specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0007] In one example, a device includes one or more processors configured to: signal a phase indication information message corresponding to a coded video sequence; signal within the phase indication information message a syntax element that specifies a horizontal position of a luminance sampling location relative to a display window, where the horizontal position is represented as a signaled value of subtracting 128 from the specified horizontal position, dividing the difference by the phase denominator, and adding 0.5 to the quotient; and signal within the phase indication information message one or more syntax elements that specify a vertical position of a luminance sampling location relative to the display window, where the vertical position is represented as a signaled value of subtracting 128 from the specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0008] In one example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to signal a phase indication information message corresponding to an encoded video sequence; signal within the phase indication information message a syntax element that specifies a horizontal position of a luminance sampling location relative to a display window, where the horizontal position is represented as a signaled value of subtracting 128 from the specified horizontal position, dividing the difference by the phase denominator, and adding 0.5 to the quotient; and signal within the phase indication information message one or more syntax elements that specify a vertical position of a luminance sampling location relative to the display window, where the vertical position is represented as a signaled value of subtracting 128 from the specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0009] In one example, an apparatus includes means for signaling a phase indication information message corresponding to a coded video sequence; means for signaling within the phase indication information message a syntax element that specifies a horizontal position of a luminance sampling location relative to a display window, where the horizontal position is expressed as a signaled value represented as subtracting 128 from a specified horizontal position, dividing the difference by a phase denominator, and adding 0.5 to the quotient; and means for signaling within the phase indication information message one or more syntax elements that specify a vertical position of a luminance sampling location relative to the display window, where the vertical position is expressed as a signaled value represented as subtracting 128 from a specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0010] In one example, a method for decoding video data includes receiving a phase indication information message; determining a horizontal position of a luminance sampling location relative to a display window based on parsed values of one or more syntax elements included in the phase indication information message, where the horizontal position is represented as a signaled value of subtracting 128 from a specified horizontal position, dividing the difference by a phase denominator, and adding 0.5 to the quotient; and determining a vertical position of the luminance sampling location relative to the display window based on parsed values of one or more syntax elements included in the phase indication information message, where the vertical position is represented as a signaled value of subtracting 128 from a specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0011] In one example, a device includes one or more processors configured to receive a phase indication information message and determine a horizontal position of a luminance sampling location relative to a display window based on parsed values of one or more syntax elements included in the phase indication information message, where the horizontal position is represented as a signaled value of subtracting 128 from a specified horizontal position, dividing the difference by a phase denominator, and adding 0.5 to the quotient; and determine a vertical position of a luminance sampling location relative to the display window based on parsed values of one or more syntax elements included in the phase indication information message, where the vertical position is represented as a signaled value of subtracting 128 from a specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0012] In one example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause a phase indication information message to be received; and a horizontal position of a luminance sampling location relative to a display window based on parsed values of one or more syntax elements included in the phase indication information message, where the horizontal position is represented as a signaled value of subtracting 128 from a specified horizontal position, dividing the difference by a phase denominator, and adding 0.5 to the quotient; and a vertical position of a luminance sampling location relative to the display window based on parsed values of one or more syntax elements included in the phase indication information message, where the vertical position is represented as a signaled value of subtracting 128 from a specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0013] In one example, an apparatus includes means for receiving a phase indication information message; means for determining a horizontal position of a luminance sampling location relative to a display window based on parsed values of one or more syntax elements included in the phase indication information message, where the horizontal position is represented as a signaled value of subtracting 128 from a specified horizontal position, dividing the difference by a phase denominator, and adding 0.5 to the quotient; and means for determining a vertical position of a luminance sampling location relative to a display window based on parsed values of one or more syntax elements included in the phase indication information message, where the vertical position is represented as a signaled value of subtracting 128 from a specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0014] The details of one or more embodiments are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. [Brief description of the drawings]
[0015] [Figure 1]FIG. 1 is a block diagram illustrating an example of a system that may be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Diagram 2] 1 is a conceptual diagram illustrating encoded video data and corresponding data structures in accordance with one or more techniques of this disclosure. [Diagram 3] A conceptual diagram illustrating a data structure that encapsulates encoded video data and corresponding metadata in accordance with one or more techniques of this disclosure. [Figure 4] A conceptual diagram illustrating an example of a sampling format for video components that may be utilized by one or more techniques of this disclosure. [Figure 5A] A conceptual diagram illustrating examples of location types of video component sampling formats that may be utilized by one or more techniques of this disclosure. [Figure 5B] A conceptual diagram illustrating examples of location types of video component sampling formats that may be utilized by one or more techniques of this disclosure. [Figure 5C] A conceptual diagram illustrating examples of location types of video component sampling formats that may be utilized by one or more techniques of this disclosure. [Figure 5D] A conceptual diagram illustrating examples of location types of video component sampling formats that may be utilized by one or more techniques of this disclosure. [Figure 5E] A conceptual diagram illustrating examples of location types of video component sampling formats that may be utilized by one or more techniques of this disclosure. [Figure 5F] A conceptual diagram illustrating examples of location types of video component sampling formats that may be utilized by one or more techniques of this disclosure. [Figure 6A] A conceptual diagram illustrating an example of a pixel coordinate system that may be utilized by one or more techniques of this disclosure. [Figure 6B] A conceptual diagram illustrating an example of a pixel coordinate system that may be utilized by one or more techniques of this disclosure. [Figure 7]A conceptual diagram illustrating an example of downsampling of video components that may be utilized by one or more techniques of this disclosure. [Figure 8] FIG. 1 is a conceptual diagram illustrating an example of downsampling in accordance with one or more techniques of this disclosure. [Figure 9] A conceptual diagram illustrating an example of components that may be included in an implementation of a system that may be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Figure 10] FIG. 1 is a block diagram illustrating an example of a video encoding device that may be configured to encode video data in accordance with one or more techniques of this disclosure. [Figure 11] FIG. 1 is a block diagram illustrating an example of a video decoding device that may be configured to decode video data in accordance with one or more techniques of this disclosure.
[0016] Video content includes video sequences consisting of a series of frames (or pictures). A series of frames may also be referred to as a group of pictures (GOP). Each video frame or video picture may be divided into one or more regions. The regions may be defined according to a basic unit (e.g., a video block) and a set of rules that define the regions. For example, the rules that define the regions may be that a region must be an integer number of video blocks arranged in a rectangle. Furthermore, the video blocks may be ordered according to a scan pattern (e.g., a raster scan). As used herein, the term video block may refer generally to a region of a picture, or more specifically, to the largest array of sample values that can be predictively coded, a subdivision thereof, and / or a corresponding structure. Furthermore, the term current video block may refer to the region of a picture that is being coded or decoded. A video block may be defined as an array of sample values. It should be noted that in some cases, a pixel value may be described as including sample values of each component of video image data, which may also be referred to as color components (e.g., luma component (Y) and chroma components (Cb and Cr), or red, green, and blue color components). It should be noted that in some cases, the terms pixel value and sample value are used interchangeably. Furthermore, in some cases, a pixel or sample may be referred to as a pel. A video sampling format, which may also be referred to as a chroma format, may define the number of chroma samples included in a video block relative to the number of luma samples included in the video block. For example, in a 4:2:0 sampling format, the sampling rate of the luma component is twice the sampling rate of the chroma components in both the horizontal and vertical directions. It should be noted that in some cases, the terms pixel value and sample value are used interchangeably.
[0017] A video encoder can perform predictive coding on video blocks and their subdivisions. Video blocks and their subdivisions may be called nodes. ITU-T H.264 specifies macroblocks containing 16x16 luma samples. That is, in ITU-T H.264, a picture is segmented into multiple macroblocks. ITU-T H.265 specifies a similar Coding Tree Unit (CTU) structure, which may be referred to as a Largest Coding Unit (LCU). In ITU-T H.265, a picture is segmented into multiple CTUs. In ITU-T H.265, for one picture, the CTU size can be set to contain 16x16, 32x32, or 64x64 luma samples. In ITU-T H.265, a CTU is composed of a coding tree block (CTB) for each component of video image data, e.g., luma (Y) and chroma (Cb and Cr). It should be noted that a video having one luma component and two corresponding chroma components may be described as having two channels, i.e., a luma channel and a chroma channel. Furthermore, in ITU-T H.265, a CTU may be partitioned according to a quadtree (QT) partitioning structure, such that a CTB of a CTU is partitioned into multiple coding blocks (CBs). That is, in ITU-T H.265, a CTU may be partitioned into multiple leaf nodes of a quadtree. According to ITU-T H.265, one luma CB together with two corresponding chroma CBs and related syntax elements is referred to as a coding unit (CU). In ITU-T H.265, the minimum allowed size of a CB may be signaled. In H.265, the minimum allowed size of a luma CB is 8x8 luma samples. In ITU-T H.265, the decision to use intra or inter prediction to code a picture region is made at the CU level. In ITU-T H.265, a CU is associated with a prediction unit structure that has its root in the CU.In ITU-T H.265, the prediction unit structure allows the luma CB and the chroma CB to be divided for the purpose of generating corresponding reference samples. That is, in ITU-T H.265, the luma CB and the chroma CB can be divided into respective luma prediction blocks (PBs) and chroma prediction blocks (PBs), where a PB includes a block of sample values to which the same prediction is applied. In ITU-T H.265, the CB can be partitioned into 1, 2, or 4 PBs. ITU-T H.265 supports PB sizes from 64×64 samples to 4×4 samples. In ITU-T H.265, a square PB is supported for intra prediction, where one CB may form a PB, or the CB may be divided into four square PBs. In addition to the square PB, a rectangular PB is supported for inter prediction, where the CB may be bisected vertically or horizontally to form a PB. It should be further noted that ITU-T H.265 supports four asymmetric PB partitions for inter prediction, where the CB is partitioned into two PBs at one quarter of the height (top or bottom) or one quarter of the width (left or right) of the CB. The intra prediction data (e.g., intra prediction mode syntax element) or inter prediction data (e.g., motion data syntax element) corresponding to the PB is used to generate reference sample values and / or predicted sample values for the PB.
[0018] JEM specifies a CTU with a maximum size of 256x256 luma samples. JEM specifies a quadtree + binary tree (QTBT) block structure. In JEM, the QTBT structure allows the leaf nodes of the quadtree to be further partitioned by a binary tree structure (BT). That is, in JEM, the binary tree structure allows the leaf nodes of the quadtree to be recursively partitioned vertically or horizontally. In JVET-T2001, the CTU is partitioned according to a quadtree + multitype tree (QTMT or QT+MTT) structure. QTMT in JVET-T2001 is similar to QTBT in JEM. However, in JVET-T2001, the multitype tree, in addition to indicating a binary partition, can also indicate a so-called ternary (or triple tree (TT)) partition. A ternary partition divides a block into three blocks vertically or horizontally. In the case of vertical TT division, the block is divided at a quarter of the width from the left edge and at a quarter of the width from the right edge, and in the case of horizontal TT division, the block is divided at a quarter of the height from the top edge and at a quarter of the height from the bottom edge.
[0019] As mentioned above, each video frame or video picture may be divided into one or more regions. For example, according to ITU-T H.265, each video frame or video picture may be partitioned to include one or more slices, and further partitioned to include one or more tiles, where each slice includes a sequence of CTUs (e.g., in raster scan order), and a tile is a sequence of CTUs corresponding to a rectangular region of the picture. It should be noted that in ITU-T H.265, a slice is a sequence of one or more slice segments, starting with an independent slice segment and including all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any). A slice segment, like a slice, is a sequence of CTUs. Thus, in some cases, the terms slice and slice segment may be used interchangeably to indicate a sequence of CTUs arranged in raster scan order. It should also be noted that in ITU-T H.265, a tile may consist of CTUs contained in more than one slice, and a slice may consist of CTUs contained in more than one tile, but ITU-T H.265 specifies that one or both of the following conditions must be satisfied: (1) all CTUs in a slice belong to the same tile, and (2) all coding tree units in a tile belong to the same slice.
[0020] For JVET-T2001, a slice is not only required to consist of an integer number of CTUs, but also an integer number of complete tiles or an integer number of consecutive complete CTU rows. It should be noted that in JVET-T2001, the slice design does not include slice segments (i.e., there are no independent / dependent slice segments). Thus, in JVET-T2001, a picture may contain a single tile, in which case the single tile is contained in a single slice, or a picture may contain multiple tiles, in which case the multiple tiles (or their CTU rows) may be contained in one or more slices. In JVET-T2001, the partitioning of a picture into tiles is specified by specifying the height of each of the tile rows and the width of each of the tile columns. Thus, in JVET-T2001, a tile is a rectangular region of CTUs within a particular tile row and a particular tile column location. It should also be noted that JVET-T2001 specifies when a picture may be partitioned into sub-pictures, where a sub-picture is a rectangular region of CTUs within a picture. The top-left CTU of a sub-picture may be located at any CTU position within a picture, with the sub-picture being constrained to contain one or more slices. Thus, unlike tiles, sub-pictures are not necessarily limited to specific row and column positions. It should also be noted that sub-pictures may be useful for encapsulating regions of interest within a picture, and a sub-bitstream extraction process may be used to decode and display only the specific region of interest. That is, as described in more detail below, a bitstream of coded video data includes a sequence of network abstraction layer (NAL) units, where one NAL unit encapsulates coded video data (i.e., video data corresponding to a slice of a picture) or one NAL unit encapsulates metadata (e.g., a parameter set) used to decode the video data, and the sub-bitstream extraction process forms a new bitstream by removing one or more NAL units from the bitstream.
[0021] FIG. 2 is a conceptual diagram illustrating an example of a picture in a group of pictures partitioned by tiles, slices, and subpictures. It should be noted that the techniques described herein may be applicable to tiles, slices, subpictures, subdivisions thereof, and / or equivalent structures. That is, the techniques described herein may be generally applicable regardless of how a picture is partitioned into regions. For example, in some cases, the techniques described herein may be applicable when tiles may be partitioned into so-called bricks, where a brick is a rectangular region of CTU rows within a particular tile. Further, for example, in some cases, the techniques described herein may be applicable when one or more tiles may be included in a so-called tile group, where a tile group includes an integer number of adjacent tiles. In the example shown in FIG. 2, Pic3 is partitioned into 16 tiles (i.e., Tile 0 to Tile 1). 15 ) and three slices (i.e., Slice 0 through Slice 2). In the example shown in FIG. 2, Slice 0 contains four tiles (i.e., Tile 0 through Tile 3), Slice 1 contains eight tiles (i.e., Tile 4 through Tile 5), 11 ), Slice2 contains four tiles (i.e., Tile 12 From Tile 15), Pic3 includes two subpictures (i.e., Subpicture0 and Subpicture1), where Subpicture0 includes Slice0 and Slice1, and Subpicture1 includes Slice2. As discussed above, subpictures may be useful for encapsulating regions of interest within a picture, and a sub-bitstream extraction process may be used to selectively decode (and display) the regions of interest. For example, referring to FIG. 2, Subpicture0 may correspond to the action portion (e.g., a view of the field) of a sporting event representation, and Subpicture1 may correspond to a scrolling banner displayed within the sporting event representation. By organizing a picture into multiple subpictures in this manner, a viewer may disable the display of the scrolling banner. That is, the sub-bitstream extraction process may remove the NAL units of Slice2 from the bitstream (and thus not decoded and / or displayed), and the NAL units of Slice0 and Slice1 may be decoded and displayed. The encapsulation of slices of a picture into respective NAL unit data structures and sub-bitstream extraction are described in further detail below.
[0022] As mentioned above, a video sampling format, which may also be referred to as a chroma format, may define the number of chroma samples included in a CU relative to the number of luma samples included in the CU. For example, in a 4:2:0 format, the sampling rate for the luma component is twice the sampling rate for the chroma components in both the horizontal and vertical directions. As a result, in a CU formatted according to the 4:2:0 format, the width and height of the array of samples for the luma component is twice the width and height of each array of samples for the chroma component. FIG. 4 is a conceptual diagram illustrating an example of a coding unit formatted according to the 4:2:0 sample format. FIG. 4 illustrates the relative positions of the chroma samples with respect to the luma samples within the CU. As mentioned above, a CU is typically defined according to the number of luma samples horizontally and vertically. Thus, as shown in FIG. 4, a 16×16 CU formatted according to the 4:2:0 sample format includes 16×16 luma component samples and 8×8 samples for each chroma component. 4 illustrates the relative positions of chroma samples with respect to luma samples for video blocks adjacent to a 16×16 CU. For a CU formatted according to the 4:2:2 format, the width of the array of samples for the luma component is twice the width of the array of samples for each chroma component, but the height of the array of samples for the luma component is equal to the height of the array of samples for each chroma component. Furthermore, for a CU formatted according to the 4:4:4 format, the array of samples for the luma component has the same width and height as the array of samples for each chroma component.
[0023] Table 1 shows how in JVET-T2001, the chroma format is specified based on the value of the syntax element chroma_format_idc. Furthermore, Table 1 shows how the variables SubWidthC and SubHeightC are specified depending on the chroma format. SubWidthC and SubHeightC are used, for example, for deblocking. With respect to Table 1, JVET-T2001 provides the following: In monochrome sampling, there is only a single sample alignment, which is nominally considered the luma alignment. In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array. In 4:4:4 sampling, each of the two chroma arrays has the same height and width as the luma array.
[0024] [Table 1]
[0025] It should be noted that in a sampling format, such as the 4:2:0 sample format, a chroma location type may be specified, i.e., in the 4:2:0 sample format, horizontal and vertical offset values may be specified for the chroma samples that indicate their spatial positioning relative to the luma samples. Table 2 provides definitions of HorizontalOffsetC and VerticalOffsetC for the five chroma location types.
[0026] [Table 2]
[0027] 5A-5F show chroma location types for 4:2:0 sample format. HorizontalOffsetC and VerticalOffsetC may act as phase offsets for horizontal and vertical filtering operations, for example when using a downsampling filter. Furthermore, Rec. ITU-T H.274, Versatile supplemental enhancement information messages for coded bitstreams, August 2020, provides syntax for video usability information (VUI) parameters and semantics for indicating VUI parameters, which are applied to one or more coded layer video sequences (CLVS). VUI parameters may be included in a vui_parameters() syntax structure included in a vui_payload() included in a sequence parameter set (SPS). Sequence parameter sets are described in more detail below.
[0028] In intra-prediction coding, an intra-prediction mode may specify the location of a reference sample within a picture. In ITU-T H.265, the predefined possible intra-prediction modes include a planar (i.e., surface-fit) prediction mode, a DC (i.e., monotonic ensemble average) prediction mode, and 33 angular prediction modes (predMode: 2-34). In JEM, the predefined possible intra-prediction modes include a planar prediction mode, a DC prediction mode, and 65 angular prediction modes. It should be noted that the planar and DC prediction modes may be referred to as non-directional prediction modes, and the angular prediction modes may be referred to as directional prediction modes. It should be noted that the techniques described herein may be generally applicable regardless of the number of predefined possible prediction modes.
[0029] In inter-prediction coding, a reference picture is determined, and a motion vector (MV) identifies samples in the reference picture that are used to generate a prediction of a current video block. For example, a current video block may be predicted using reference sample values located in one or more previously encoded pictures, and a motion vector is used to indicate the location of the reference block relative to the current video block. The motion vector may, for example, describe a horizontal displacement component (i.e., MVx) of the motion vector, a vertical displacement component (i.e., MVy) of the motion vector, and a resolution (e.g., ¼ pixel precision, ½ pixel precision, 1 pixel precision, 2 pixel precision, 4 pixel precision) of the motion vector. Previously decoded pictures, which may include pictures output before or after the current picture, may be organized into one or more reference picture lists and identified using reference picture index values. Also, in inter-prediction coding, uni-prediction refers to generating a prediction using sample values from one reference picture, and bi-prediction refers to generating a prediction using respective sample values from two reference pictures. That is, in uni-prediction, a prediction for a current video block is generated using a single reference picture and corresponding motion vector, and in bi-prediction, a prediction for a current video block is generated using a first reference picture and corresponding first motion vector and a second reference picture and corresponding second motion vector. In bi-prediction, the respective sample values are combined (e.g., added, rounded, clipped, or averaged according to weights) to generate a prediction. Pictures and their regions may be classified based on which type of prediction mode may be used to encode the video block. That is, for a region having a B type (e.g., B slice), bi-prediction mode, uni-prediction mode, and intra-prediction mode may be used, for a region having a P type (e.g., P slice), uni-prediction mode and intra-prediction mode may be used, and for a region having an I type (e.g., I slice), only intra-prediction mode may be used. As described above, the reference picture is identified via a reference index.For example, there may be a single reference picture list RefPicList0 for a P slice, and a second independent reference picture list RefPicList1 for a B slice in addition to RefPicList0. It should be noted that in the case of uni-prediction in a B slice, one of RefPicList0 or RefPicList1 may be used to generate a prediction. It should also be noted that during the decoding process, at the start of decoding a picture, a reference picture list is generated from previously decoded pictures stored in a decoded picture buffer (DPB).
[0030] Furthermore, coding standards may support various modes of motion vector prediction. Motion vector prediction allows a value of a motion vector for a current video block to be derived based on another motion vector. For example, a set of candidate blocks with associated motion information may be derived from spatially and temporally neighboring blocks to the current video block. Furthermore, generated (or default) motion information may be used for motion vector prediction. Examples of motion vector prediction include advanced motion vector prediction (AMVP), temporal motion vector prediction (TMVP), the so-called "combined" mode, as well as "skip" and "direct" motion estimation. Furthermore, other examples of motion vector prediction include advanced temporal motion vector prediction (ATMVP) and spatial-temporal motion vector prediction (STMVP). In motion vector prediction, both the video coding device and the video decoding device perform the same process to derive a set of candidates. Thus, the same set of candidates is generated for the current video block during encoding and decoding.
[0031] As mentioned above, in inter-prediction coding, reference samples in previously coded pictures are used to code video blocks in a current picture. Previously coded pictures available for use as references when coding a current picture are called reference pictures. It should be noted that the decoding order does not necessarily correspond to the picture output order, i.e., the temporal order of pictures in a video sequence. In ITU-T H.265, pictures are stored in a decoded picture buffer (DPB) (which may be called a frame buffer, a reference buffer, a reference picture buffer, etc.) when they are decoded. In ITU-T H.265, pictures stored in the DPB are removed from the DPB when they are output, and are no longer needed for coding subsequent pictures. In ITU-T H.265, the decision of whether a picture should be removed from the DPB is called once for each picture after the decoding of the slice header, i.e., at the beginning of the decoding of the picture. For example, referring to FIG. 2, Pic2 is shown as referring to Pic1. Similarly, Pic3 is shown as referring to Pic0. With reference to Figure 2, assuming that the picture numbers correspond to the decoding order, the DPB is populated as follows: after the decoding of Pic0 the DPB contains {Pic0}, at the start of the decoding of Pic1 the DPB contains {Pic0}, after the decoding of Pic1 the DPB contains {Pic0,Pic1}, at the start of the decoding of Pic2 the DPB contains {Pic0,Pic1}. Then Pic2 is decoded with reference to Pic1, and after the decoding of Pic2 the DPB contains {Pic0,Pic1,Pic2}. At the start of the decoding of Pic3, pictures Pic0 and Pic1 are marked for removal from the DPB as they are not needed to decode Pic3 (or any subsequent pictures not shown), and assuming that Pic1 and Pic2 have already been output, the DPB is updated to contain {Pic0}. Then Pic3 is decoded with reference to Pic0. The process of marking pictures for removal from the DPB may be referred to as Reference Picture Set (RPS) management.
[0032] As mentioned above, intra prediction data or inter prediction data is used to generate reference sample values for a block of sample values. The difference between a sample value included in the current PB or another type of picture region structure and an associated reference sample (e.g., a reference sample generated using prediction) may be referred to as residual data. The residual data may include a respective array of difference values corresponding to each component of the video image data. The residual data may be in the pixel domain. A transform such as a discrete cosine transform, a discrete sine transform (DST), an integer transform, a wavelet transform, or a conceptually similar transform may be applied to the array of difference values to generate transform coefficients. It is noted that in ITU-T H.265 and JVET-T2001, a CU is associated with a transform tree structure rooted at the CU level. The transform tree is partitioned into one or more transform units (TUs). That is, the array of difference values may be partitioned (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) for the purpose of generating transform coefficients. For each component of video image data, such a subdivision of difference values may be referred to as a Transform Block (TB). It should be noted that in some cases, a core transform and a subsequent secondary transform may be applied (in a video encoder) to generate transform coefficients. In a video decoder, the order of the transforms is reversed.
[0033] A quantization process may be performed directly on the transform coefficients or residual sample values (e.g., in the case of palette coding quantization). Quantization approximates the transform coefficients with a magnitude limited to a set of specific values. Quantization essentially scales the transform coefficients to vary the amount of data required to represent a set of transform coefficients. Quantization may include dividing the transform coefficients (or the values resulting from adding an offset value to the transform coefficients) by a quantization scale factor and some associated rounding function (e.g., rounding to the nearest integer). Quantized transform coefficients are sometimes referred to as coefficient level values. Inverse quantization (or dequantization) may include multiplying the coefficient level values by the quantization scale factor and any inverse rounding or offset addition operations. It should be noted that, as used herein, the term quantization process may refer in some cases to division by the quantization scale factor to generate level values and in some cases to multiplication by the quantization scale factor to recover the transform coefficients. That is, the quantization process may refer, in some cases, to quantization, and, in some cases, to dequantization. Additionally, it should be noted that, in some examples, the quantization process is described with respect to arithmetic operations associated with decimal representation, but such description is for illustrative purposes and should not be construed as limiting. For example, the techniques described herein may be implemented in devices using binary arithmetic, and the like. For example, the multiplication and division operations described herein may be implemented using bit shifting operations, and the like.
[0034] The quantized transform coefficients and syntax elements (e.g., syntax elements indicating a coding structure of a video block) may be entropy coded according to an entropy coding technique. The entropy coding process includes coding the values of the syntax elements using a lossless data compression algorithm. Examples of entropy coding techniques include content adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), probability interval partitioning entropy coding (PIPE), etc. The entropy coded quantized transform coefficients and the corresponding entropy coded syntax elements may form an adapted bitstream that can be used to regenerate the video data at a video decoder. The entropy coding process, such as CABAC, may include performing binarization on the syntax elements. Binarization refers to the process of converting the value of a syntax element into a series of one or more bits. These bits may be referred to as "bins." The binarization may include one or a combination of the following coding techniques: fixed-length coding, unary coding, shortened unary coding, shortened Rice coding, Golomb coding, k-th exponential Golomb coding, and Golomb-Rice coding. For example, the binarization may include representing the integer value 5 of the syntax element as 00000101 using an 8-bit fixed-length binarization technique, or representing the integer value 5 as 11110 using a unary coding binarization technique. As used herein, each of the terms fixed-length coding, unary coding, shortened unary coding, shortened Rice coding, Golomb coding, k-th exponential Golomb coding, and Golomb-Rice coding may refer to general implementations of these techniques and / or more specific implementations of these coding techniques. For example, an implementation of Golomb-Rice coding may be specifically defined according to a video coding standard.In a CABAC example, a context provides, for a particular bin, a most likely state (MPS) value for that bin (i.e., the MPS for a bin is either 0 or 1) and a probability value that the bin is in the MPS or least likely state (LPS). For example, a context may indicate that a bin's MPS is 0 and that the probability of the bin being 1 is 0.3. It should be noted that the context may be determined based on values of previously coded bins, including bins in the current syntax element and previously coded syntax elements. For example, values of syntax elements associated with neighboring video blocks may be used to determine the context of the current bin.
[0035] As mentioned above, video content includes a video sequence consisting of a series of pictures, and each picture may be divided into one or more regions. In JVET-T2001, the coded representation of a picture includes video coding layer (VCL) NAL units of a particular layer in an access unit (AU) and encompasses all CTUs of that picture. For example, referring again to FIG. 2, the coded representation of Pic3 is encapsulated in three coded slice NAL units (i.e., Slice0 NAL unit, Slice1 NAL unit, and Slice2 NAL unit). It should be noted that the term video coding layer (VCL) NAL unit is used as a generic term for coded slice NAL units. That is, VCL NAL is a generic term that includes all types of slice NAL units. As explained above and in more detail below, NAL units may encapsulate metadata used to decode video data. NAL units encapsulating metadata used to decode video sequences are generally referred to as non-VCL NAL units. Thus, in JVET-T2001, a NAL unit can be a VCL NAL unit or a non-VCL NAL unit. It should be noted that a VCL NAL unit contains slice header data, which provides information used to decode a particular slice. Thus, in JVET-T2001, information used to decode video data, which in some cases may be referred to as metadata, is not necessarily included in a non-VCL NAL unit. JVET-T2001 specifies that a picture unit (PU) is a set of NAL units that are related to each other according to a specified classification rule, are consecutive in decoding order, and contain exactly one coded picture, and that an access unit (AU) is a set of PUs that belong to different layers and contain multiple coded pictures that are related at the same time for output from the DPB. JVET-T2001 further specifies that a layer is a set of VCL NAL units, all of which have a particular layer identifier value and associated non-VCL NAL units.Furthermore, in JVET-T2001, a PU consists of zero or one picture header (PH) NAL unit, one coded picture containing one or more VCL NAL units, and zero or more other non-VCL NAL units. Furthermore, in JVET-T2001, a coded video sequence (CVS) is a sequence of AUs, which consists, in decoding order, of a coded video sequence start (CVSS) AU followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to, but not including, any subsequent AUs that are CVSS AUs, where the CVSS AUs are AUs in which the PUs of each layer in the CVS reside, and the coded picture in each current picture unit is a coded layer video sequence start (CLVSS) picture. In JVET-T2001, a coding layer video sequence (CLVS) is a sequence of PUs in the same layer that consists, in decoding order, of a CLVSS PU followed by zero or more PUs that are not CLVSS PUs, including all subsequent PUs up to but not including a subsequent PU that is a CLVSS PU. In ITU-T JVET-T2001, a bitstream may be described as including a sequence of NAL units that form one or more CVSs.
[0036] Multi-layer video coding allows a video representation to be decoded / displayed as a representation corresponding to a base layer of the video data, and one or more additional representations decoded / displayed corresponding to enhancement layers of the video data. For example, a base layer may allow a video representation having a basic level of quality (e.g., high resolution rendering and / or 30 Hz frame rate) and an enhancement layer may allow a video representation having an enhanced quality level (e.g., ultra-high resolution rendering and / or 60 Hz frame rate). An enhancement layer may be coded by reference to the base layer. That is, for example, a picture in an enhancement layer may be coded (e.g., using inter-layer prediction techniques) by reference to one or more pictures (including scaled versions thereof) in the base layer. It should be noted that layers may also be coded independently of each other. In this case, there may be no inter-layer prediction between two layers. Each NAL unit may include an identifier indicating the layer of the video data to which the NAL unit is associated. As mentioned above, a sub-bitstream extraction process may be used to decode and display only a particular region of interest of a picture. Furthermore, a sub-bitstream extraction process may be used to decode and display only a particular layer of the video data. Sub-bitstream extraction may refer to a process in which a device receiving a compliant or conforming bitstream forms a new compliant or conforming bitstream by discarding and / or modifying data in the received bitstream. For example, sub-bitstream extraction may be used to form a new compliant or conforming bitstream that corresponds to a particular representation (e.g., a higher quality representation) of a video image.
[0037] In JVET-T2001, each of a video sequence, a GOP, a picture, a slice, and a CTU may be associated with metadata that describes video coding attributes, and some types of metadata are encapsulated within non-VCL NAL units. ITU-T JVET-T2001 defines parameter sets that may be used to describe video data attributes and / or video coding attributes. Specifically, JVET-T2001 includes four types of parameter sets: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), and adaptation parameter set (APS), where an SPS applies across zero or more CVSs, a PPS applies across zero or more coded pictures, an APS applies to zero or more slices, and a VPS may be optionally referenced by an SPS. A PPS applies to one or more individual coded pictures that reference it. In JVET-T2001, parameter sets may be encapsulated as non-VCL NAL units and / or signaled as messages. JVET-T2001 also includes a picture header (PH), which is either encapsulated as a non-VCL NAL unit if signaled in its own NAL unit, or as part of a VCL NAL unit if signaled in the slice header of a coded slice. In JVET-T2001, the picture header applies to all slices of a coded picture. JVET-T2001 also allows for the signaling of decoding capability information (DCI) and supplemental enhancement information (SEI) messages. In JVET-T2001, DCI and SEI messages aid in processes related to decoding, display, or other purposes, but may not be necessary for the decoding process to create luma or chroma samples. In JVET-T2001, DCI and SEI messages may be signaled in the bitstream using non-VCL NAL units.Furthermore, DCI and SEI messages may be conveyed by some mechanism other than being present in the bitstream (ie, signaled out-of-band).
[0038] FIG. 3 shows an example of a bitstream including multiple CVSs, where the CVSs include AUs, and the AUs include picture units. The example shown in FIG. 3 corresponds to the example of encapsulating multiple slice NAL units in a bitstream shown in the example of FIG. 2. In the example shown in FIG. 3, the picture unit corresponding to Pic3 includes three VCL NAL coded slice NAL units, namely, Slice0 NAL unit, Slice1 NAL unit, and Slice2 NAL unit, and two non-VCL NAL units, namely, PPS NAL unit and PH NAL unit. It should be noted that the header in FIG. 3 is a NAL unit header (i.e., should not be confused with a slice header). It should also be noted that in FIG. 3, other non-VCL NAL units, not shown, such as, for example, an SPS NAL unit, a VPS NAL unit, and an SEI message NAL unit, may be included in the CVS. It should also be noted that in other examples, the PPS NAL unit used to decode Pic3 may be included elsewhere in the bitstream, e.g., in the picture unit corresponding to Pic0, or may be provided by an external mechanism. As described in more detail below, in JVET-T2001, the PH syntax structure may be present in the slice header of a VCL NAL unit or in the PH NAL unit of the current PU.
[0039] For the formulas used herein, the following arithmetic operators may be used:
[0040] [Table 3]
[0041] In addition, the following mathematical functions can be used: Log2(x) The base 2 logarithm of x
[0042]
number
[0043] With respect to the example syntax used herein, the following definitions of logical operators may apply: x&&y The Boolean logic "product" of x and y x||y The Boolean "union" of x and y !Boolean logic "not" x?y:zIf x is true or not equal to 0, then evaluate the value of y, else evaluate the value of z.
[0044] In addition, the following relational operators may be applied:
[0045] [Table 4]
[0046] Furthermore, it should be noted that in the syntax descriptors used herein, the following descriptors may apply: -b(8): A byte (8 bits) with an arbitrary bit pattern. The parsing process of this descriptor is specified by the return value of the function read_bits(8). -f(n): a fixed pattern bit string using n bits written left bit first (left to right). The parsing process for this descriptor is specified by the return value of the function read_bits(n). -i(n): A signed integer using n bits. If n is "v" in the syntax table, the number of bits varies depending on the values of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits(n), which is interpreted as a two's complement integer representation written most significant bit first. - se(v): A signed integer zeroth order Exp-Golomb encoding syntax element with left bit first. -tb(v): A truncated binary using up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -tu(v): A shortened unary using up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -u(n): An unsigned integer using n bits. If n is "v" in the syntax table, the number of bits varies depending on the values of other syntax elements. The parsing process of this descriptor is specified by the return value of the function read_bits(n), which is interpreted as the binary representation of an unsigned integer written most significant bit first. -ue(v): An unsigned integer zero-order Exp-Golomb encoding syntax element, left bit first.
[0047] It should be noted that JVET-T2001 defines NAL unit header semantics that specify the type of raw byte sequence payload (RBSP) data structure contained in the NAL unit. Table 3 shows the syntax of the NAL unit header provided in JVET-T T2001. [Table 5]
[0048] JVET-T2001 specifies the following definitions for each syntax element shown in Table 3. forbidden_zero_bit shall be equal to 0. nuh_reserved_zero_bit SHALL be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified by ITU-T ISO / IEC in the future. Although this version of this specification requires that the value of nuh_reserved_zero_bit be equal to 0, decoders conforming to this version of this specification shall allow a value of nuh_reserved_zero_bit equal to 1 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1. nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs, or the identifier of the layer to which a non-VCL NAL unit applies. Values of nuh_layer_id shall be in the range 0 to 55, inclusive. Other values of nuh_layer_id are reserved for future use by ITU-T|ISO / IEC. Although values of nuh_layer_id are required to be in the range 0 to 55, inclusive, in this version of this specification, decoders conforming to this version of this specification shall allow values of nuh_layer_id greater than 55 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with a nuh_layer_id greater than 55. The value of nuh_layer_id shall be the same for all VCL NAL units of a coded picture. The value of nuh_layer_id for a coded picture or PU is the value of nuh_layer_id of the VCL NAL units of that coded picture or PU. If nal_unit_type is equal to PH_NUT or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit. If nal_unit_type is equal to EOS_NUT, then nuh_layer_id shall be equal to one of the nuh_layer_id values of a layer that exists in CVS. NOTE – The values of nuh_layer_id in the DCI, OPI, VPS, AUD, and EOB NAL units are not constrained. nuh_temporal_id_plus1-1 specifies the temporal identifier of that NAL unit. The value of nuh_temporal_id_plus1 shall not be equal to 0. The variable TemporalId is derived as follows: TemporalId=nuh_temporal_id_plus1-1 If nal_unit_type is in the range IDR_W_RADL to RSV_IRAP_11, inclusive, then TemporalId shall be equal to 0. If nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, then TemporalId shall be greater than 0. The value of TemporalId shall be the same for all VCL NAL units of an AU. The value of TemporalId for a coded picture, PU, or AU is the value of TemporalId of the VCL NAL units of that coded picture, PU, or AU. The value of TemporalId for a sublayer representation is the maximum value of TemporalId of all VCL NAL units in that sublayer representation. The values of TemporalId for non-VCL NAL units are constrained as follows: - If nal_unit_type is equal to DCI_NUT, OPI_NUT, VPS_NUT, or SPS_NUT, then TemporalId shall be equal to 0 and the TemporalId of the AU containing that NAL unit shall be equal to 0. Otherwise, if nal_unit_type is equal to PH_NUT, then TemporalId shall be equal to the TemporalId of the PU that contains the NAL unit. Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0. Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, then TemporalId shall be equal to the TemporalId of the AU containing the NAL unit. Otherwise, if nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, then TemporalId shall be equal to or greater than the TemporalId of the PU that contains the NAL unit. NOTE - If the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values of all AUs to which the non-VCL NAL unit belongs. If nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId may be equal to or greater than the TemporalId of the containing AU, since all PPS and APS may be included at the beginning of the bitstream (e.g., if they are carried out-of-band and the receiver places them at the beginning of the bitstream), and the first coded picture has TemporalId equal to 0. nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit, as specified in Table 4. NAL units with nal_unit_type in the range UNSPEC28 to UNSPEC31, inclusive, and with no semantics specified SHALL NOT affect the decoding process as defined in this specification. NOTE - NAL unit types in the range UNSPEC_28 to UNSPEC_31 may be used as determined by the application. The decoding process for these values of nal_unit_type is not specified by this specification. Because different applications may use these NAL unit types for different purposes, special care is expected to be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not prescribe any management of these values. The use of these nal_unit_type values may be appropriate only in situations where "collisions" of usage (i.e., multiple different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are insignificant, impossible, or are defined or governed, for example, by a controlling application or transport specification, or by controlling the environment in which the bitstream is delivered. For purposes other than determining the amount of data in a DU of the bitstream (as specified in Appendix C), a decoder SHALL ignore (remove from the bitstream and discard) the content of all NAL units that use reserved values of nal_unit_type. NOTE - This requirement allows for the definition of future compatibility extensions to this specification.
[0049] [Table 6] NOTE - A Clean Random Access (CRA) picture may have an associated RASL or RADL picture in the bitstream. NOTE - An immediate decoding refresh (IDR) picture with nal_unit_type equal to IDR_N_LP has no associated leading picture in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL has no associated RASL picture in the bitstream, but may have an associated RADL picture in the bitstream. The value of nal_unit_type shall be the same for all VCL NAL units of a subpicture. A subpicture is referred to as having the same NAL unit type as the VCL NAL units of that subpicture. For the VCL NAL units of any particular picture, the following applies: -If pps_mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and a picture or PU is referenced as having the same NAL unit type as the coded slice NAL units of that picture or PU. - Otherwise (pps_mixed_nalu_types_in_pic_flag is equal to 1), all of the following constraints apply. -A picture shall have at least two sub-pictures. - A VCL NAL unit of a picture shall have two or more distinct nal_unit_type values. There shall be no VCL NAL units with nal_unit_type equal to GDR_NUT in the picture. - If a VCL NAL unit of a picture has nal_unit_type equal to nalUnitTypeA, and nalUnitTypeA is equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT, then all other VCL NAL units of that picture shall have nal_unit_type equal to nalUnitTypeA or TRAIL_NUT. The value of nal_unit_type shall be the same for all pictures within an IRAP or GDR AU. If sps_video_parameter_set_id is greater than 0, and for j equal to GeneralLayerIdx[nuh_layer_id] and for any value of i in the range j+1 to vps_max_layers_minus1, inclusive, vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0, and pps_mixed_nalu_types_in_pic_flag is equal to 1, the value of nal_unit_type shall not be equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT. The requirements for bitstream conformance are that the following constraints apply: If a picture is the leading picture of an IRAP picture, it shall be a RADL or RASL picture. - If a subpicture is the leading subpicture of an IRAP subpicture, it shall be a RADL picture or a RASL subpicture. If a picture is not the leading picture of an IRAP picture, it shall not be a RADL or RASL picture. If a subpicture is not the leading subpicture of an IRAP subpicture, it shall not be a RADL or RASL subpicture. - No RASL pictures associated with an IDR picture shall be present in the bitstream. - No RASL sub-pictures associated with an IDR sub-picture shall be present in the bitstream. - There shall be no RADL pictures in the bitstream that are associated with an IDR picture with nal_unit_type equal to IIDR_N_LP. NOTE – Performing random access at the position of an IRAP AU (and correctly decoding the non-RASL pictures in the IRAP AU and all subsequent AUs in decoding order) is possible by discarding all PUs before the IRAP AU, provided that the respective parameter sets are available at the time of reference (either in the bitstream or by external means not specified in this specification). - There shall be no RADL sub-pictures associated with an IDR sub-picture with nal_unit_type equal to IIDR_N_LP in the bitstream. -Any picture with nuh_layer_id equal to layerId that precedes in decoding order an IRAP picture with nuh_layer_id equal to a particular value layerId shall precede that IRAP picture in output order and shall precede in output order the RADL picture associated with that IRAP picture. Any subpicture with nuh_layer_id equal to layerId and subpicture index equal to a particular value subpicIdx that precedes in decoding order an IRAP subpicture with nuh_layer_id equal to a particular value layerId and subpicture index equal to a particular value subpicIdx shall precede in output order that IRAP subpicture and all RADL subpictures associated with that IRAP subpicture. -A picture with nuh_layer_id equal to the specified value layerId that precedes in decoding order a recovery point picture with nuh_layer_id equal to layerId shall precede that recovery point picture in output order. -A subpicture with nuh_layer_id equal to a specific value layerId and a subpicture index equal to a specific value subpicIdx that precedes in decoding order a subpicture with nuh_layer_id equal to layerId and a subpicture index equal to subpicIdx in the recovery point picture shall precede that subpicture in output order in that recovery point picture. - The RASL picture associated with a CRA picture shall precede in output order the RADL picture associated with that CRA picture. - A RASL sub-picture associated with a CRA sub-picture shall precede in output order the RADL sub-picture associated with that CRA sub-picture. - A RASL picture with nuh_layer_id equal to a particular value layerId and associated with a CRA picture shall follow in output order the IRAP or GDR picture with nuh_layer_id equal to layerId that precedes the CRA picture in decoding order. - A RASL subpicture with nuh_layer_id equal to a particular value layerId and subpicture index equal to a particular value subpicIdx and associated with a CRA subpicture shall follow in output order the IRAP or GDR subpicture with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx that precedes the CRA subpicture in decoding order. -When sps_field_seq_flag is equal to 0, the following applies: if the current picture with nuh_layer_id equal to a particular value layerId is a leading picture associated with an IRAP picture, then it shall precede in decoding order all non-leading pictures associated with the same IRAP picture. Otherwise (sps_field_seq_flag is equal to 1), let picA and picB be the first and last leading pictures, respectively, in decoding order associated with an IRAP picture, then there shall be at most one non-leading picture with nuh_layer_id equal to layerId preceding picA in decoding order, and there shall be no non-leading picture with nuh_layer_id equal to layerId between picA and picB in decoding order. - if sps_field_seq_flag is equal to 0, the following applies: if the current subpicture with nuh_layer_id equal to a particular value layerId and subpicture index equal to a particular value subpicIdx is a leading subpicture associated with an IRAP subpicture, then it shall precede in decoding order all non-leading subpictures associated with the same IRAP subpicture. Otherwise (sps_field_seq_flag is equal to 1), let subpicA and subpicB be the first and last leading subpictures, respectively, in decoding order associated with an IRAP subpicture, then there shall be no more than one non-leading subpicture with nuh_layer_id equal to layerId and subpicIdx preceding subpicA in decoding order, and there shall be no non-leading picture with nuh_layer_id equal to layerId and subpicIdx between picA and picB in decoding order.
[0050] It should be noted that, in general, an Intra Random Access Point (IRAP) picture is a picture that does not reference any picture other than itself for prediction in the decoding process. In JVET-T2001, an IRAP picture can be a Clean Random Access (CRA) picture or an Immediate Decoding Refresh (IDR) picture. In JVET-T2001, the first picture in the bitstream in decoding order must be an IRAP or a Gradual Decoding Refresh (GDR) picture. JVET-T2001 describes the concept of a leading picture, which is a picture that precedes the associated IRAP picture in output order. JVET-T2001 further describes the concept of a trailing picture, which is a non-IRAP picture that follows the associated IRAP picture in output order. A trailing picture associated with an IRAP picture also follows that IRAP picture in decoding order. For an IDR picture, there are no trailing pictures that require references to pictures decoded before the IDR picture. JVET-T2001 specifies when a CRA picture may have leading pictures that follow the CRA picture in decoding order and contain inter-picture prediction references to pictures decoded before the CRA picture. Thus, when a CRA picture is used as a random access point, these leading pictures may not be decodable and are identified as random access skipped leading (RASL) pictures. Another type of picture that can follow an IRAP picture in decoding order and precede it in output order is a random access decodable leading (RADL) picture, which cannot contain references to pictures that precede the IRAP picture in decoding order. A GDR picture is a picture whose VCL NAL unit has a nal_unit_type equal to GDR_NUT.If the current picture is a GDR picture associated with a picture header that signals the syntax element recovery_poc_cnt, and there is a picture picA in the CLVS that follows the current GDR picture in decoding order and has a PicOrderCntVal equal to the PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt, then picture picA is called a recovery point picture.
[0051] As mentioned above, HorizontalOffsetC and VerticalOffsetC may act as phase offsets for horizontal and vertical filtering operations, for example when using a downsampling filter. To some extent similarly, for downsampling of a single component (e.g., luma), the location of the filtered sample may be relative to the number of filter support samples. It should also be noted that different graphics rendering systems may utilize different pixel space coordinate systems. For example, FIG. 6A shows an example of a pixel space coordinate system in which the rendering object origin (0,0) is located in the upper left corner and the pixel center is located at (x+0.5,y+0.5). The example of FIG. 6A may correspond to Microsoft Direct3D 10 graphics. FIG. 6B shows an example of a pixel space coordinate system in which the rendering object origin (0,0) is located offset from the upper left corner by (0.5,0.5) and the pixel center is located at (x,y). The example of FIG. 6B may correspond to Microsoft Direct3D 9 graphics. It should be noted that in some cases, the pixel space coordinate system may be aligned with the display window coordinate system (which may include, for example, texel space). However, in some cases, the pixel coordinate system and the display window may not be aligned. In these cases, when rendering an image to a display window, equations are needed to transform the pixel space coordinate system to the display window coordinate system to ensure alignment. It should be noted that in some cases, the terms display window and rendering window may be used interchangeably.
[0052] When downsampling components of video image data, there may be several ways to make the downsampled result values correspond to a pixel coordinate system. For example, referring to FIG. 7, in the example of FIG. 7, the input samples are downsampled by a factor of 4 in each direction. However, in the upper example of FIG. 7, the downsampling filter generates the filtered samples at a center location of the 4×4 region, and in the lower example of FIG. 7, the downsampling filter generates the filtered samples at a co-located location at the top left of the 4×4 region. As further shown in FIG. 7, the pixel center values of the filtered samples may be expressed based on a pixel space coordinate system (e.g., Microsoft Direct3D 10). In most cases, the device receiving the filtered sample values simply receives a 2×2 array of four filtered sample values. Thus, when the device performs upsampling, the device assumes a pixel center location for the received samples (i.e., according to a default process, for example). As shown in Figure 7, depending on the downsampling filter employed, the assumed set of filtered pixel center locations (e.g., [(2,2),(6,2),(2,6),(6,6)]) may be quite different from the actual set of filtered pixel center locations (e.g., [(0.5,0.5),(4.5,0.5),(0.5,4.5),(4.5,4.5)]). When the assumed set of filtered pixel center locations differs from the actual set of filtered pixel center locations, the result is a so-called phase shift that may be visible during rendering of the image.
[0053] There can be many ways to position the filtered samples relative to the input samples. That is, many filtering designs can be used for downsampling. “Down-sample phase indication (SEI message)”, 24th Meeting of ISO / IEC JTC1 / SC29 / WG11 6-15 October 2020, Teleconference, document JVET-X0092-v1, referred to as JVET-X0092, describes several examples of downsampling filters, such as Lanczons1 phase 0, Linear, etc. It should be noted that downsampling filters can be classified as either odd or even based on the number of filter coefficients. JVET-X0092 further provides an example where a phase shift is perceptible in a rendered image, specifying that the phase shift can generally be calculated as follows: Phase shift = (downsampling scale factor - 1) / 2
[0054] It should be noted that the downsampling / upsampling scale factors may be non-integer values. For example, a 2×2 array of filtered sample values may be upsampled to a 9×9 array. In such cases, the assumed position of the filtered sample pixel center location may affect the interpolated sample values.
[0055] JVET-X0092 describes the addition of a phase indication SEI message to a versatile supplemental enhancement information message for coded video bitstreams, where the phase indication SEI message contains a single syntax element indicating a proposed phase indication for the video. Table 5 shows the syntax of the phase indication SEI message provided in JVET-X0092.
[0056] [Table 7]
[0057] With respect to Table 5, JVET-X0092 provides the following semantics: The phase instruction SEI message provides information about the proposed phase position of the video image in the rescaling operation. The phase indication SEI message lasts from the current picture to the end of the CLVS for the current layer in decoding order. All phase indication SEI messages that apply to the same CLVS shall have the same content. phase_idc indicates the phase of the video as specified in Table 6. Values of phase_idc shall be in the range 0 to 2, inclusive. Values of phase_idc specified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091 2 shall not be present in bitstreams conforming to this version of this specification.
[0058] [Table 8] NOTE - The phase indication SEI message can also be used to provide information about the phase of the filter used to create the video from the higher spatial resolution video through a downsampling process. As an example, a video generated by decimating the number of samples of the higher resolution video by 50% both horizontally and vertically (i.e., scale factor 0.5) using a symmetric filter with an odd number of filter taps (e.g., [A,B,C,B,A]) may signal a phase_idc equal to 1. A video generated by decimating the number of samples of the higher resolution video by 50% both horizontally and vertically (i.e., scale factor 0.5) using a symmetric filter with an even number of filter taps (e.g., [D,E,F,F,E,D]) may signal a phase_idc equal to 2.
[0059] It should be noted that because JVET-X0092's phase_idc only supports indicating one of top-left corner alignment or center alignment, it cannot clearly and / or adequately support at least the following cases: (1) downsampling that occurs in only one dimension (e.g., from 1920x1080 to 960x1080); (2) downsampling is cascaded and different filters are used for each downsampling step (e.g., a 7680x4320 picture is downsampled to 3840x2160 using an odd filter, then downsampled to 1920x1080 using an even filter); and (3) downsampling is cascaded and different ratios are used for each downsampling step (e.g., a 3840x2160 picture is subsampled to 1920x1080 using an odd filter, then subsampled to 640x360). With regard to (2) and (3), FIG. 8 shows a conceptual example illustrating various filtering steps for downsampling a 3840×2160 video to a 640×360 video. Thus, as shown in FIG. 8, there are multiple ways to downsample a video to a particular scaling factor (i.e., 6, as shown in FIG. 8), and each filtering step can be implemented with different relative positions of the filtered samples. It should be noted that in a dynamic streaming infrastructure, cascaded filters may be used to generate videos of different resolutions. Thus, it may be useful to allow signaling of more information regarding the phase of the filters used to create a video at a particular size, according to the techniques herein. Thus, the signaling of phase_idc in JVET-X0092 may be less than ideal.
[0060] FIG. 1 is a block diagram illustrating an example of a system that may be configured to encode (encode and / or decode) video data in accordance with one or more techniques of this disclosure. System 100 represents an example of a system that may encapsulate video data in accordance with one or more techniques of this disclosure. As shown in FIG. 1, system 100 includes a source device 102, a communication medium 110, and a destination device 120. In the example shown in FIG. 1, source device 102 may include any device configured to encode video data and transmit the encoded video data to communication medium 110. Destination device 120 may include any device configured to receive the encoded video data via communication medium 110 and decode the encoded video data. Source device 102 and / or destination device 120 may include computing devices equipped for wired and / or wireless communication, and may include, for example, set-top boxes, digital video recorders, televisions, desktop, laptop, or tablet computers, gaming consoles, medical imaging devices, and mobile devices including, for example, smartphones, cellular phones, personal gaming devices.
[0061] The communication medium 110 may include any combination of wireless and wired communication media and / or storage devices. The communication medium 110 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The communication medium 110 may include one or more networks. For example, the communication medium 110 may include a network configured to enable access to the World Wide Web, e.g., the Internet. The network may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, Global System Mobile Communications (GSM) standards, code division multiple access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, Internet Protocol (IP) standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards.
[0062] A storage device may include any type of device or storage medium capable of storing data. A storage medium may include a tangible or non-transitory computer-readable medium. A computer-readable medium may include an optical disk, a flash memory, a magnetic memory, or any other suitable digital storage medium. In some examples, a memory device or a portion thereof may be described as a non-volatile memory, and in other examples, a portion of a memory device may be described as a volatile memory. Examples of volatile memory may include random access memory (RAM), dynamic random access memory (DRAM), and static random access memory (SRAM). Examples of non-volatile memory may include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or a form of electrically programmable memory (EPROM) or electrically erasable and programmable memory (EEPROM). The storage device(s) may include memory cards (e.g., Secure Digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid state drives. Data may be stored on the storage devices according to a defined file format. FIG. 9 is a conceptual diagram illustrating an example of components that may be included in an implementation of system 100. In the exemplary implementation shown in FIG. 9, system 100 includes one or more computing devices 402A-402N, a television service network 404, a television service provider site 406, a wide area network 408, a local area network 410, and one or more content provider sites 412A-412N.The implementation shown in FIG. 9 represents an example of a system that may be configured to enable digital media content, such as movies, live sporting events, and data and applications and their associated media presentations, to be distributed to and accessed by a plurality of computing devices, such as computing devices 402A-402N. In the example shown in FIG. 9, computing devices 402A-402N may include any device configured to receive data from one or more of a television service network 404, a wide area network 408, and / or a local area network 410. For example, computing devices 402A-402N may include televisions that may be equipped for wired and / or wireless communication and configured to receive services over one or more data channels, including so-called smart televisions, set-top boxes, and digital video recorders. Additionally, computing devices 402A-402N may include desktop, laptop, or tablet computers, gaming consoles, mobile devices including, for example, "smart" phones, cellular phones, and personal gaming devices.
[0063] Television service network 404 is an example of a network configured to enable delivery of digital media content, which may include television services. For example, television service network 404 may include a public terrestrial television network, a public or subscription-based satellite television service provider network, and a public or subscription-based cable television provider network and / or an over the top service provider or an Internet service provider. It should be noted that, in some embodiments, television service network 404 may be primarily used to enable the provision of television services, but television service network 404 may also enable the provision of other types of data and services based on any combination of telecommunication protocols described herein. It should also be noted that, in some embodiments, television service network 404 may enable bidirectional communication between television service provider site 406 and one or more of computing devices 402A-402N. Television service network 404 may include any combination of wireless communication media and / or wired communication media. The television service network 404 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The television service network 404 may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include the DVB standard, the ATSC standard, the ISDB standard, the DTMB standard, the DMB standard, the Data Over Cable Service Interface Specification (DOCSIS) standard, the HbbTV standard, the W3C standard, and the UPnP standard.
[0064] Referring again to FIG. 9, the television service provider site 406 may be configured to distribute television services over the television service network 404. For example, the television service provider site 406 may include one or more broadcast stations, cable television providers, or satellite television providers, or Internet-based television providers. For example, the television service provider site 406 may be configured to receive transmissions including television programming over satellite uplinks / downlinks. Additionally, as shown in FIG. 9, the television service provider site 406 may be in communication with a wide area network 408 and configured to receive data from content provider sites 412A-412N. It should be noted that in some embodiments, the television service provider site 406 may include a television studio from which content may originate.
[0065] The wide area network 408 may include a packet-based network and may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include Global System Mobile Communications (GSM) standards, code division multiple access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, European Standards (EN), IP standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards, such as, for example, one or more of the IEEE 802 standards (e.g., Wi-Fi). The wide area network 408 may include any combination of wireless communication media and / or wired communication media. The wide area network 408 may include coaxial cables, fiber optic cables, twisted pair cables, Ethernet cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. In one embodiment, the wide area network 408 may include the Internet. The local area network 410 may include a packet-based network and operate according to a combination of one or more telecommunications protocols. The local area network 410 may be distinguished from the wide area network 408 based on the level of access and / or physical infrastructure. For example, the local area network 410 may include a secure home network.
[0066] Referring again to FIG. 9, the content provider sites 412A-412N represent examples of sites that may provide multimedia content to the television service provider site 406 and / or the computing devices 402A-402N. For example, the content provider sites may include a studio having one or more studio content servers configured to provide multimedia files and / or streams to the television service provider site 406. In one embodiment, the content provider sites 412A-412N may be configured to provide multimedia content using an IP suite. For example, the content provider sites may be configured to provide multimedia content to receiving devices according to the Real Time Streaming Protocol (RTSP), HTTP, or the like. Additionally, the content provider sites 412A-412N may be configured to provide data, including hypertext-based content, or the like, over the wide area network 408 to one or more of the receiving devices, the computing devices 402A-402N, and / or the television service provider site 406. The content provider sites 412A-412N may include one or more web servers. The data provided by the data provider sites 412A-412N may be defined according to a data format.
[0067] Referring again to FIG. 1, the source device 102 includes a video source 104, a video encoder 106, a data encapsulator 107, and an interface 108. The video source 104 may include any device configured to capture and / or store video data. For example, the video source 104 may include a video camera and a storage device operatively coupled thereto. The video encoder 106 may include any device configured to receive video data and generate a compliant bitstream representing the video data. A compliant bitstream may refer to a bitstream that a video decoder may receive and from which the video data can be reproduced. The aspects of a compliant bitstream may be defined according to a video coding standard. When generating a compliant bitstream, the video encoder 106 may compress the video data. The compression may be lossy (perceptible or imperceptible to a viewer) or lossless. FIG. 10 is a block diagram illustrating an example of a video encoder 500 capable of implementing techniques for encoding video data described herein. It should be noted that, although the example video encoder 500 is shown as having separate functional blocks, such illustration is for illustrative purposes and is not intended to limit the video encoder 500 and / or its subcomponents to any particular hardware or software architecture. The functionality of the video encoder 500 may be realized using any combination of hardware, firmware, and / or software implementations.
[0068] The video encoder 500 may perform intra-predictive and inter-predictive coding of picture regions and may therefore be referred to as a hybrid video encoder. In the example shown in FIG. 10, the video encoder 500 receives a source video block. In some examples, the source video block may include a region of a picture partitioned according to a coding structure. For example, the source video data may include a macroblock, a CTU, a CB, a subdivision thereof, and / or another equivalent coding unit. In some examples, the video encoder 500 may be configured to perform additional subdivision of the source video block. It should be noted that the techniques described herein are generally applicable to video coding, regardless of how the source video data is partitioned before and / or during encoding. In the example shown in Fig. 10, the video encoding device 500 includes an adder 502, a transform coefficient generating device 504, a coefficient quantization unit 506, an inverse quantization and transform coefficient processing unit 508, an adder 510, an intra prediction processing unit 512, an inter prediction processing unit 514, a filter unit 516, and an entropy encoding unit 518. As shown in Fig. 10, the video encoding device 500 receives source video blocks and outputs a bitstream.
[0069] In the example shown in FIG. 10, the video coding device 500 may generate residual data by subtracting a prediction video block from a source video block. Selection of the prediction video block is described in more detail below. The adder 502 represents a component configured to perform this subtraction operation. In one example, the subtraction of the video blocks is performed in the pixel domain. The transform coefficient generator 504 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block or a subdivision thereof (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) to generate a set of residual transform coefficients. The transform coefficient generator 504 may be configured to perform any and all combinations of transforms included in the family of discrete triangular transforms, including approximations of the discrete triangular transform. The transform coefficient generator 504 may output the transform coefficients to a coefficient quantizer 506. The coefficient quantizer 506 may be configured to perform quantization of the transform coefficients. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may change the rate distortion (i.e., video bit rate vs. quality) of the coded video data. The degree of quantization may be modified by adjusting a quantization parameter (QP). The quantization parameter may be determined based on a slice level value and / or a CU level value (e.g., a CU delta QP value). The QP data may include any data used to determine a QP for quantizing a particular set of transform coefficients. As shown in FIG. 10, the quantized transform coefficients (which may also be referred to as level values) are output to an inverse quantization and transform coefficient processor 508. The inverse quantization and transform coefficient processor 508 may be configured to apply inverse quantization and inverse transform to generate reconstructed residual data. As shown in FIG. 10, the reconstructed residual data may be added to a predictive video block at an adder 510. In this manner, the coded video block may be reconstructed, and the resulting reconstructed video block may be used to evaluate the coding quality for a given prediction, transform, and / or quantization.The video encoding device 500 may be configured to perform multiple encoding passes (e.g., performing encoding while varying one or more of prediction, transformation parameters, and quantization parameters). The rate-distortion or other system parameters of the bitstream may be optimized based on evaluation of the reconstructed video blocks. Furthermore, the reconstructed video blocks may be stored and used as references for predicting subsequent blocks.
[0070] Referring again to FIG. 10, the intra-predictor 512 may be configured to select an intra-prediction mode for a video block to be coded. The intra-predictor 512 may be configured to evaluate a frame and determine an intra-prediction mode to use to code a current block. As described above, possible intra-prediction modes may include planar, DC, and angular prediction modes. It should also be noted that in some examples, the prediction mode for a chroma component may be inferred from the prediction mode for a luma prediction mode. The intra-predictor 512 may select the intra-prediction mode after performing one or more coding passes. Furthermore, in one embodiment, the intra-predictor 512 may select the prediction mode based on a rate-distortion analysis. As shown in FIG. 10, the intra-predictor 512 outputs intra-prediction data (e.g., syntax elements) to the entropy encoder 518 and the transform coefficient generator 504. As described above, the transform performed on the residual data may be mode-dependent (e.g., a secondary transform matrix may be determined based on the prediction mode).
[0071] Referring again to FIG. 10, the inter prediction processor 514 may be configured to perform inter prediction coding on the current video block. The inter prediction processor 514 may be configured to receive a source video block and calculate a motion vector for a PU of the video block. The motion vector may indicate the displacement of a prediction block of the video block in the current video frame relative to a prediction block in a reference frame. The inter prediction coding may use one or more reference pictures. Furthermore, the motion prediction may be uni-predictive (using one motion vector) or bi-predictive (using two motion vectors). The inter prediction processor 514 may be configured to select a prediction block by calculating pixel differences determined by, for example, sum of absolute difference (SAD), sum of square difference (SSD), or other difference measures. As described above, a motion vector may be determined and determined according to the motion vector prediction. The inter prediction processor 514 may be configured to perform motion vector prediction as described above. The inter prediction processor 514 may be configured to generate a prediction block using the motion prediction data. For example, the inter prediction processor 514 may place the prediction video block in a frame buffer (not shown in FIG. 10). It should be noted that the inter prediction processor 514 may be further configured to apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for use in motion estimation. The inter prediction processor 514 may output motion prediction data for the calculated motion vectors to the entropy encoder 518.
[0072] As shown in FIG. 10, the filter unit 516 receives the reconstructed video blocks and the coding parameters and outputs a modified version of the reconstructed video data. The filter unit 516 may be configured to perform deblocking and / or Sample Adaptive Offset (SAO) filtering. SAO filtering is a non-linear amplitude mapping that can be used to improve the reconstruction by adding an offset to the reconstructed video data. It should be noted that, as shown in FIG. 10, the intra-prediction unit 512 and the inter-prediction unit 514 may receive the modified version of the reconstructed video blocks via the filter unit 516. The entropy coder 518 receives the quantized transform coefficients and prediction syntax data (i.e., intra-prediction data, motion prediction data). It should be noted that in some examples, the coefficient quantizer 506 may perform a scan of a matrix including the quantized transform coefficients before the coefficients are output to the entropy coder 518. In other examples, the entropy coder 518 may perform the scan. The entropy encoder 518 may be configured to perform entropy encoding according to one or more of the techniques described herein. Thus, the video encoder 500 represents one example of a device configured to generate video data encoded according to one or more techniques of this disclosure.
[0073] Referring again to FIG. 1, data encapsulator 107 may receive the encoded video data and generate a compliant bitstream, such as a sequence of NAL units, according to a defined data structure. A device receiving the compliant bitstream may regenerate the video data therefrom. Additionally, as described above, sub-bitstream extraction may refer to a process in which a device receiving a compliant bitstream forms a new compliant bitstream by discarding and / or modifying data in the received bitstream. It should be noted that the term conforming bitstream may be used instead of the term compliant bitstream. In one example, data encapsulator 107 may be configured to generate syntax according to one or more techniques described herein. It should be noted that data encapsulator 107 need not be located in the same physical device as video encoder 106. For example, the functions described as being performed by video encoder 106 and data encapsulator 107 may be distributed among multiple devices shown in FIG. 9.
[0074] As mentioned above, it may be useful for the techniques herein to allow for signaling of more information regarding the phase of the filter used to create a motion image at a particular size. In one example, in accordance with the techniques herein, an SEI message may include syntax elements indicating fractional offset values for each of the horizontal and vertical dimensions. That is, for example, the location of the top-left filtered sample relative to the top-left corner of the display window may be given by dividing a designated numerator by a designated denominator. Table 7A shows the syntax of an exemplary phase indication SEI message in accordance with the techniques herein.
[0075] [Table 9]
[0076] With respect to Table 7A, the semantics may be based on the following: hor_phase_num and hor_phase_den specify the horizontal position of the luminance sampling locations relative to the display window. If hor_phase_den is equal to 0, the horizontal position is undefined. Otherwise, if hor_phase_den is greater than 0, the horizontal position hor_phase_num÷hor_phase_den is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. hor_phase_num shall be greater than or equal to 0 and less than or equal to hor_phase_den. ver_phase_num and ver_phase_den specify the vertical position of the luminance sampling locations relative to the display window. If ver_phase_den is equal to 0, then the vertical position is undefined. Otherwise, if ver_phase_den is greater than 0, then the vertical position ver_phase_num÷ver_phase_den is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. ver_phase_num shall be greater than or equal to 0 and less than or equal to ver_phase_den. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0077] Table 7B illustrates the syntax of an example phase indication SEI message in accordance with the techniques herein.
[0078] [Table 10]
[0079] With respect to Table 7B, the semantics may be based on the following: hor_phase_num and hor_phase_den specify the horizontal position of the luminance sampling locations relative to the display window. If hor_phase_den is equal to 0, then the horizontal position is undefined. Otherwise, if hor_phase_den is greater than 0, then the horizontal position hor_phase_num÷hor_phase_den is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. If present, hor_phase_num shall be greater than or equal to 0 and less than or equal to hor_phase_den. ver_phase_num and ver_phase_den specify the vertical position of the luminance sampling locations relative to the display window. If ver_phase_den is equal to 0, then the vertical position is undefined. Otherwise, if ver_phase_den is greater than 0, then the vertical position ver_phase_num÷ver_phase_den is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. If present, ver_phase_num shall be greater than or equal to 0 and less than or equal to ver_phase_den. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0080] Tables 7C and 7D show the syntax of an example phase indication SEI message in accordance with the techniques herein.
[0081] [Table 11]
[0082] [Table 12]
[0083] With respect to Tables 7C and 7D, the semantics may be based on the following: hor_phase_num and hor_phase_den specify the horizontal position of the luminance sampling locations relative to the display window. If hor_phase_den is equal to 0, the horizontal position is undefined. Otherwise, if hor_phase_den is greater than 0, the horizontal position hor_phase_num÷hor_phase_den is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. If present, the hor_phase_num syntax element shall be Ceil(Log2(hor_phase_den+1)) bits in length and its value shall be greater than or equal to 0 and less than or equal to hor_phase_den. ver_phase_num and ver_phase_den specify the vertical position of the luminance sampling locations with respect to the display window. If ver_phase_den is equal to 0, then the vertical position is undefined. Otherwise, if ver_phase_den is greater than 0, then the vertical position ver_phase_num÷ver_phase_den is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. If present, the ver_phase_num syntax element shall be Ceil(Log2(ver_phase_den+1)) bits in length and its value shall be greater than or equal to 0 and less than or equal to ver_phase_den. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled. With reference to FIG. 7, for the example shown in FIG. 7, the following values of hor_phase_num, hor_phase_den, ver_phase_num, and ver_phase_den may be signaled: For [(2,2),(6,2),(2,6),(6,6)]: hor_phase_num=1,hor_phase_den=2,ver_phase_num=1,ver_phase_den=2 For [(0.5,0.5),(4.5,0.5),(0.5,4.5),(4.5,4.5)]: hor_phase_num=1, hor_phase_den=8, ver_phase_num=1, ver_phase_den=8
[0084] In one example, a fixed denominator may be used for either or both the horizontal and vertical directions in accordance with the techniques herein. In one example, the denominator may be a power of two (e.g., 128). It should be noted that using a fixed denominator requires fewer syntax elements compared to the syntax in Table 7A, does not imply division by non-powers of two, and may introduce minor rounding errors (e.g., 1 / 12=0.08333...is expressed as 11 / 128=0.0859375). Table 8 shows the syntax of an exemplary phase indication SEI message in accordance with the techniques herein.
[0085] [Table 13]
[0086] Regarding Table 8, the semantics may be based on the following: hor_phase_num specifies the horizontal position of the luminance sampling locations relative to the display window. If hor_phase_num is equal to 255, the horizontal position is undefined. Otherwise, if hor_phase_num is less than or equal to 128, the horizontal position, hor_phase_num÷128, is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. hor_phase_num values 129 through 254 are reserved for future use. ver_phase_num specifies the vertical position of the luminance sampling locations relative to the display window. If ver_phase_num is equal to 255, the vertical position is undefined. Otherwise, if ver_phase_num is less than or equal to 128, the vertical position, ver_phase_num÷128, is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. ver_phase_num values 129 through 254 are reserved for future use. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0087] In one example, in accordance with the techniques herein, a common denominator may be signaled for both the horizontal and vertical directions. Table 9 shows the syntax of an example phase indication SEI message in accordance with the techniques herein.
[0088] [Table 14]
[0089] With respect to Table 9, the semantics may be based on the following: hor_phase_num and phase_den specify the horizontal position of the luminance sampling locations relative to the display window. If phase_den is equal to 0, the horizontal position is undefined. Otherwise, if phase_den is greater than 0, the horizontal position hor_phase_num÷phase_den is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. hor_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. ver_phase_num and phase_den specify the vertical position of the luminance sampling locations with respect to the display window. If phase_den is equal to 0, the vertical position is undefined. Otherwise, if phase_den is greater than 0, the vertical position ver_phase_num÷phase_den is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. ver_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0090] With respect to Table 9, in one example, the semantics may be based on the following: hor_phase_num and phase_den specify the horizontal position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then the horizontal position is undefined. Otherwise, if phase_den is greater than 0, then the horizontal position hor_phase_num / (2 * phase_den) is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. hor_phase_num must be greater than or equal to 0 and less than or equal to (2 * phase_den) or less. ver_phase_num and phase_den specify the vertical position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then the vertical position is undefined. Otherwise, if phase_den is greater than 0, then the vertical position is ver_phase_num÷(2 * phase_den) is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. ver_phase_num must be greater than or equal to 0 and less than or equal to (2 * phase_den) or less. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0091] Note that for the above example, multiplying the signaled denominator by 2 allows for a wider range of denominator values without signaling additional bits. * phase_den) and ver_phase_num÷(2 * The denominator phase_den is typically less than or equal to 1 / 2. It is unlikely that the values hor_phase_num and ver_phase_num are greater than phase_den. Note that the denominator phase_den is often an even number.
[0092] With respect to Table 9, in one example, the semantics may be based on the following: hor_phase_num and phase_den specify the horizontal position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then the horizontal position is undefined. Otherwise, if phase_den is greater than 0, then the horizontal position is (hor_phase_num-128+phase_den)÷(2 * phase_den) is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. ver_phase_num and phase_den specify the vertical position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then the vertical position is undefined. Otherwise, if phase_den is greater than 0, then the vertical position is (ver_phase_num-128+phase_den)÷(2 * phase_den) is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0093] Note that for the above example, the midpoint of the u(8) range used to code phase_den is 128. When upsampling a picture with any offset in the range 0 to 1, the use of reference picture resampling (RPR) may result in offsets outside the range 0 to 1.
[0094] In one example, in accordance with the techniques herein, a flag may be signaled to indicate if the phase is the same for the horizontal and vertical directions. Table 10 shows the syntax of an example phase indication SEI message in accordance with the techniques herein.
[0095] [Table 15]
[0096] With respect to Table 10, the semantics may be based on the following: same_hor_ver_phase equal to 0 specifies that the vertical position of the luminance sampling location relative to the display window may not be the same as the horizontal position of the luminance sampling location relative to the display window. hor_phase_num and hor_phase_den specify the horizontal position of the luminance sampling locations relative to the display window. If hor_phase_den is equal to 0, the horizontal position is undefined. Otherwise, if hor_phase_den is greater than 0, the horizontal position hor_phase_num÷hor_phase_den is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. hor_phase_num shall be greater than or equal to 0 and less than or equal to hor_phase_den. ver_phase_num and ver_phase_den specify the vertical position of the luminance sampling locations with respect to the display window. If ver_phase_den is equal to 0, then the vertical position is undefined. Otherwise, if ver_phase_den is greater than 0, then the vertical position ver_phase_num÷ver_phase_den is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. ver_phase_num shall be greater than or equal to 0 and less than or equal to ver_phase_den. If not present, ver_phase_num and ver_phase_den shall be inferred to be equal to hor_phase_num and hor_phase_den, respectively. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0097] It should be noted that in one example, the precision of hor_phase_num, hor_phase_den, phase_den, ver_phase_num, and / or ver_phase_den may be less than 8 bits. For example, hor_phase_num, hor_phase_den, ver_phase_num, and ver_phase_den may each be signaled with 4 bits. Signaling hor_phase_num, hor_phase_den, ver_phase_num, and ver_phase_den with 4 bits has the advantage of reducing signaling overhead while keeping the data bytes aligned.
[0098] As provided in the semantics above, when using a shader, such as when used with OpenGL, the texture coordinates may be offset by an amount proportional to the signaled horizontal and vertical phase indicators. In one example, in accordance with the techniques herein, the signaling techniques described above may be used in conjunction with phase_idc. That is, in one example, the signaled value may be used to derive the value of tex_offset. For example, Table 11 shows the syntax of an example phase indicator SEI message in accordance with the techniques herein.
[0099] [Table 16]
[0100] With respect to Table 11, semantics may be based on the following: phase_idc indicates the tex_offset information as provided in Table 12. Values of phase_idc shall be in the range 0 to 3, inclusive. Values of phase_idc reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091 2 shall not be present in bitstreams conforming to this version of this specification.
[0101] [Table 17] hor_phase_num and phase_den specify the horizontal position of the luminance sampling locations relative to the display window. If phase_den is equal to 0, the horizontal position is undefined. Otherwise, if phase_den is greater than 0, the horizontal position hor_phase_num÷phase_den is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. hor_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. ver_phase_num and phase_den specify the vertical position of the luminance sampling locations with respect to the display window. If phase_den is equal to 0, the vertical position is undefined. Otherwise, if phase_den is greater than 0, the vertical position ver_phase_num÷phase_den is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. ver_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. NOTE: Phase indicators may be used during the rendering process. For example, when using a shader, such as when used in OpenGL, texture coordinates may be offset by an amount proportional to the horizontal and vertical phase indicators signaled.
[0102] That is, assuming the picture is rendered into a rectangle with the lower left corner at (-1,-1) and the upper right corner at (1,1), the value of tex_offset may be derived using phase_idc, hor_phase_num, ver_phase_num, and phase_den, where the vertex and fragment shaders are specified as follows: Vertex Shader
[0103] [Table 18] Fragment Shader
[0104] [Table 19] Here, vec2 may point to a vector of two floating point values.
[0105] In one example, if the texture coordinate system is defined with (0,0) as the bottom left corner, tex_offset can be derived from hor_phase_num, ver_phase_num, and phase_den as follows: tex_offset=vec2(0.5f-hor_phase_num / phase_den,ver_phase_num / phase_den-0.5f));
[0106] In one example, if the texture coordinate system is defined with the upper left corner as (0,0), then tex_offset may be derived from hor_phase_num, ver_phase_num, and phase_den as follows: tex_offset=vec2(0.5f-hor_phase_num / phase_den,0.5f-ver_phase_num / phase_den));
[0107] Table 13A illustrates the syntax of an example phase indication SEI message in accordance with the techniques herein.
[0108] [Table 20]
[0109] With respect to Table 13A, in one example, the semantics may be based on the following: hor_phase_num_plus128 and phase_den specify the horizontal position of the luminance sampling location with respect to the display window. If phase_den is equal to 0, then no horizontal position is defined. Otherwise, if phase_den is greater than 0, then the horizontal position 0.5 + ((hor_phase_num_plus128-128)÷phase_den) defines the distance between the display window and the first horizontal luminance sample location, expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. ver_phase_num_plus128 and phase_den specify the vertical position of the luminance sampling location with respect to the display window. If phase_den is equal to 0, then no vertical position is defined. Otherwise, if phase_den is greater than 0, then the vertical position 0.5 + ((ver_phase_num_plus128-128)÷phase_den) defines the distance between the display window and the first vertical luminance sample location, expressed in units of the vertical distance between two vertically adjacent luminance sampling locations.
[0110] If the texture coordinate system is defined with (0,0) as the bottom left corner, then tex_offset can be derived from hor_phase_num_plus128, ver_phase_num_plus128, and phase_den as follows: tex_offset=vec2((hor_phase_num_plus128-128)÷phase_den,(128-ver_phase_num_plus128)÷phase_den)
[0111] If the texture coordinate system is defined with the upper left corner as (0,0), then tex_offset can be derived from hor_phase_num_plus128, ver_phase_num_plus128, and phase_den as follows: tex_offset=vec2((hor_phase_num_plus128-128)÷phase_den,(ver_phase_num_plus128-128)÷phase_den)
[0112] Table 13B illustrates an example phase indication SEI message syntax in accordance with the techniques herein, where separate horizontal and vertical denominator values are signaled.
[0113] [Table 21]
[0114] With respect to Table 13B, in one example, the semantics may be based on the following: hor_phase_num_plus128 and hor_phase_den specify the horizontal position of the luminance sampling location relative to the display window. If hor_phase_den is equal to 0, then no horizontal position is defined. Otherwise, if hor_phase_den is greater than 0, then the horizontal position 0.5 + ((hor_phase_num_plus128-128)÷hor_phase_den) defines the distance between the display window and the first horizontal luminance sample location, expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. ver_phase_num_plus128 and ver_phase_den specify the vertical position of the luminance sampling location with respect to the display window. If ver_phase_den is equal to 0, then no vertical position is defined. Otherwise, if ver_phase_den is greater than 0, then the vertical position 0.5+((ver_phase_num_plus128-128)÷ver_phase_den) defines the distance between the display window and the first vertical luminance sample location, expressed in units of the vertical distance between two vertically adjacent luminance sampling locations.
[0115] If the texture coordinate system is defined with (0,0) as the bottom left corner, then tex_offset can be derived from hor_phase_num_plus128, ver_phase_num_plus128, hor_phase_den, and ver_phase_den as follows: tex_offset=vec2((hor_phase_num_plus128-128)÷hor_phase_den,(128-ver_phase_num_plus128)÷ver_phase_den)
[0116] If the texture coordinate system is defined with the upper left corner as (0,0), then tex_offset can be derived from hor_phase_num_plus128, ver_phase_num_plus128, hor_phase_den, and ver_phase_den as follows: tex_offset=vec2((hor_phase_num_plus128-128)÷hor_phase_den,(ver_phase_num_plus128-128)÷ver_phase_den)
[0117] Table 14A illustrates the syntax of an example phase indication SEI message in accordance with the techniques herein.
[0118] [Table 22]
[0119] With respect to Table 14A, in one example, the semantics may be based on the following: hor_phase_sign, abs_hor_phase_num, and phase_den specify the horizontal position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then the horizontal position is undefined. Otherwise, if phase_den is greater than 0, then the horizontal position 0.5+((1-2 * hor_phase_sign) *abs_hor_phase_num)÷phase_den defines the distance between the display window and the first horizontal luminance sample location and is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. ver_phase_sign, abs_ver_phase_num, and phase_den specify the vertical position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then the vertical position is undefined. Otherwise, if phase_den is greater than 0, then the vertical position is 0.5+((1-2 * ver_phase_sign) * abs_ver_phase_num)÷phase_den defines the distance between the display window and the first vertical luminance sample location and is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations.
[0120] In one example, if the texture coordinate system is defined with (0,0) as the bottom left corner, tex_offset can be derived from hor_phase_sign, abs_hor_phase_num, ver_phase_sign, abs_ver_phase_num, and phase_den as follows: tex_offset=vec2(((1-2 * hor_phase_sign) * abs_hor_phase_num)÷phase_den,((2 * ver_phase_sign-1) * abs_ver_phase_num)÷phase_den)
[0121] If the texture coordinate system is defined with (0,0) as the bottom left corner, then tex_offset can be derived from hor_phase_sign, abs_hor_phasenum, ver_phase_sign, abs_ver_phase_num, and phase_den as follows: tex_offset=vec2(((1-2 * hor_phase_sign) * abs_hor_phase_num)÷phase_den,((1-2 * ver_phase_sign) * abs_ver_phase_num)÷phase_den)
[0122] Table 14B illustrates an example phase indication SEI message syntax in accordance with the techniques herein, where separate horizontal and vertical denominator values are signaled.
[0123] [Table 23]
[0124] With respect to Table 14B, in one example, the semantics may be based on the following: hor_phase_sign, abs_hor_phase_num, and hor_phase_den specify the horizontal position of the luminance sampling location relative to the display window. If hor_phase_den is equal to 0, the horizontal position is undefined. Otherwise, if hor_phase_den is greater than 0, the horizontal position is 0.5+((1-2 * hor_phase_sign) * abs_hor_phase_num)÷hor_phase_den defines the distance between the display window and the first horizontal luminance sample location, expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. ver_phase_sign, abs_ver_phase_num, and ver_phase_den specify the vertical position of the luminance sampling location relative to the display window. If ver_phase_den is equal to 0, then the vertical position is undefined. Otherwise, if ver_phase_den is greater than 0, then the vertical position is 0.5+((1-2 *ver_phase_sign) * abs_ver_phase_num)÷ver_phase_den defines the distance between the display window and the first vertical luminance sample location and is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations.
[0125] In one example, if the texture coordinate system is defined with (0,0) as the bottom left corner, tex_offset can be derived from hor_phase_sign, abs_hor_phase_num, ver_phase_sign, abs_ver_phase_num, hor_phase_den, and ver_phase_den as follows: tex_offset=vec2(((1-2 * hor_phase_sign) * abs_hor_phase_num)÷hor_phase_den,((2 * ver_phase_sign-1) * abs_ver_phase_num)÷ver_phase_den)
[0126] If the texture coordinate system is defined with (0,0) as the bottom left corner, then tex_offset can be derived from hor_phase_sign, abs_hor_phasenum, ver_phase_sign, abs_ver_phase_num, hor_phase_den, and ver_phase_den as follows: tex_offset=vec2(((1-2 * hor_phase_sign) * abs_hor_phase_num)÷hor_phase_den,((1-2 * ver_phase_sign) * abs_ver_phase_num)÷ver_phase_den)
[0127] Table 15A illustrates the syntax of an example phase indication SEI message in accordance with the techniques herein.
[0128] [Table 24]
[0129] With respect to Table 15A, in one example, the semantics may be based on the following: hor_phase_num and phase_den specify the horizontal position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then no horizontal position is defined. Otherwise, if phase_den is greater than 0, then the horizontal position 0.5+hor_phase_num÷phase_den defines the distance between the display window and the first horizontal luminance sample location, expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. ver_phase_num and phase_den specify the vertical position of the luminance sampling location relative to the display window. If phase_den is equal to 0, then no vertical position is defined. Otherwise, if phase_den is greater than 0, then the vertical position 0.5+ver_phase_num÷phase_den defines the distance between the display window and the first vertical luminance sample location, expressed in units of the vertical distance between two vertically adjacent luminance sampling locations.
[0130] If the texture coordinate system is defined with (0,0) as the bottom left corner, then tex_offset can be derived from hor_phase_num, ver_phase_num, and phase_den as follows: tex_offset=vec2(hor_phase_num÷phase_den,-ver_phase_num÷phase_den)
[0131] If the texture coordinate system is defined with the upper left corner as (0,0), then tex_offset can be derived from hor_phase_num, ver_phase_num, and phase_den as follows: tex_offset=vec2(hor_phase_num÷phase_den,ver_phase_num÷phase_den)
[0132] Table 15B illustrates an example phase indication SEI message syntax in accordance with the techniques herein, where separate horizontal and vertical denominator values are signaled.
[0133] [Table 25]
[0134] With respect to Table 15B, in one example, the semantics may be based on the following: hor_phase_num and hor_phase_den specify the horizontal position of the luminance sampling location relative to the display window. If hor_phase_den is equal to 0, then no horizontal position is defined. Otherwise, if hor_phase_den is greater than 0, then the horizontal position 0.5+hor_phase_num÷hor_phase_den defines the distance between the display window and the first horizontal luminance sample location, expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. ver_phase_num and ver_phase_den specify the vertical position of the luminance sampling location relative to the display window. If ver_phase_den is equal to 0, then no vertical position is defined. Otherwise, if ver_phase_den is greater than 0, then the vertical position 0.5+ver_phase_num÷ver_phase_den defines the distance between the display window and the first vertical luminance sample location, expressed in units of the vertical distance between two vertically adjacent luminance sampling locations.
[0135] If the texture coordinate system is defined with (0,0) as the bottom left corner, then tex_offset can be derived from hor_phase_num, ver_phase_num, hor_phase_den, and ver_phase_den as follows: tex_offset=vec2(hor_phase_num÷hor_phase_den,-ver_phase_num÷ver_phase_den)
[0136] If the texture coordinate system is defined with the upper left corner as (0,0), then tex_offset can be derived from hor_phase_num, ver_phase_num, hor_phase_den, and ver_phase_den as follows: tex_offset=vec2(hor_phase_num÷hor_phase_den,ver_phase_num÷ver_phase_den)
[0137] As mentioned above, a chroma location type may be specified. In some cases, downsampling luma samples and chroma samples with the same particular filter according to a scaling factor may result in a chroma sample position relative to the luma sample position that does not match any of the entries defined for the chroma location type ChromaLocType. For example, in chroma location type 2, downsampling a luma sample with an even filter according to a scaling factor of 2 may result in a chroma location that cannot be represented using the entries defined for ChromaLocType when the same filter is applied to the chroma sample. In some cases, for example, in chroma location type 0, if the same filter is used for luma and chroma, it is necessary to use an odd filter horizontally and an even filter vertically to keep the relative positions of luma and chroma consistent with chroma location type 0.
[0138] With respect to examples corresponding to Tables 11 and 12 above, in one example, in accordance with the techniques herein, HorizontalOffsetC and VerticalOffsetC can be used to derive tex_offset_chroma as follows: tex_offset_chroma=(0.5-HorizontalOffetC, 0.5-VerticalOffsetC), where the vertex and fragment shaders are specified as follows: Vertex Shader
[0139] [Table 26] Fragment Shader
[0140] [Table 27]
[0141] In one example, in accordance with the techniques herein, information regarding chroma location alignment may be signaled. Further, in one example, in accordance with the techniques herein, the signaling techniques described above may be used in conjunction with phase_idc. For example, in one example, the semantics of phase_idc may be based on the following: phase_idc indicates the phase of the video as specified in Table 16. Values of phase_idc shall be in the range 0 to 6, inclusive. Values of phase_idc specified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091 2 shall not be present in bitstreams conforming to this version of this specification.
[0142] [Table 28]
[0143] In one example, in accordance with the techniques herein, Table 17 illustrates the syntax of an exemplary phase indication SEI message in accordance with the techniques herein.
[0144] [Table 29]
[0145] With respect to Table 17, the semantics may be based on the following: same_hor_ver_phase equal to 0 specifies that the vertical position of the luminance sampling location relative to the display window may not be the same as the horizontal position of the luminance sampling location relative to the display window. chroma_phase_idc indicates the presence of chroma-specific syntax elements. When chroma_phase_idc is equal to 0, no chroma-specific syntax elements are present. When chroma_phase_idc is equal to 1, syntax elements specific to the first chroma component are present. When chroma_phase_idc is equal to 2, syntax elements specific to the first and second chroma components are present. The value 3 for chroma_phase_idc is reserved for future use. hor_phase_num and phase_den specify the horizontal position of the luminance sampling locations relative to the display window. If phase_den is equal to 0, the horizontal position is undefined. Otherwise, if phase_den is greater than 0, the horizontal position hor_phase_num÷phase_den is expressed in units of the horizontal distance between two horizontally adjacent luminance sampling locations. hor_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. ver_phase_num and phase_den specify the vertical position of the luminance sampling locations with respect to the display window. If phase_den is equal to 0, the vertical position is undefined. Otherwise, if phase_den is greater than 0, the vertical position ver_phase_num÷phase_den is expressed in units of the vertical distance between two vertically adjacent luminance sampling locations. ver_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. cb_hor_phase_num and phase_den specify the horizontal position of the Cb chrominance sampling locations relative to the display window. If phase_den is equal to 0, then the horizontal position is undefined. Otherwise, if phase_den is greater than 0, then the horizontal position cb_hor_phase_num÷phase_den is expressed in units of the horizontal distance between two horizontally adjacent Cb chrominance sampling locations. cb_hor_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. If not present, cb_hor_phase_num is set equal to hor_phase_num. cb_ver_phase_num and phase_den specify the vertical position of the Cb chrominance sampling locations relative to the display window. If phase_den is equal to 0, then the vertical position is undefined. Otherwise, if phase_den is greater than 0, then the vertical position cb_ver_phase_num÷phase_den is expressed in units of the vertical distance between two vertically adjacent Cb chrominance sampling locations. ver_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. If not present, cb_ver_phase_num is set equal to ver_phase_num. cr_hor_phase_num and phase_den specify the horizontal position of the Cr chrominance sampling locations relative to the display window. If phase_den is equal to 0, then the horizontal position is undefined. Otherwise, if phase_den is greater than 0, then the horizontal position hor_phase_num÷phase_den is expressed in units of the horizontal distance between two horizontally adjacent Cr chrominance sampling locations. cr_hor_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. If not present, cr_hor_phase_num is set equal to hor_phase_num. cr_ver_phase_num and phase_den specify the vertical position of the Cr chrominance sampling locations relative to the display window. If phase_den is equal to 0, then the vertical position is undefined. Otherwise, if phase_den is greater than 0, then the vertical position cb_ver_phase_num÷phase_den is expressed in units of the vertical distance between two vertically adjacent Cr chrominance sampling locations. ver_phase_num shall be greater than or equal to 0 and less than or equal to phase_den. If not present, cb_ver_phase_num is set equal to ver_phase_num.
[0146] With respect to examples corresponding to Tables 16 and 17 above, in one example, in accordance with the techniques herein, the value of tex_offset_chroma may be derived using phase_chroma_idc, cb_hor_phase_num, cb_ver_phase_num, cr_hor_phase_num, cr_ver_phase_num, and phase_den, where the vertex shader and fragment shader are specified as described above.
[0147] It should be noted that in some examples, phase_indication() may be included within a vui_parameters() syntax structure.
[0148] It should be noted that the phase indication proposal is not only used as guidance for rendering video. It can also be used when performing other operations on the decoded video, such as rescaling, color conversion, format conversion, transcoding, transrating, mixing, editing, or other forms of post-production. In particular, rendering is one important aspect / use of the SEI messages of this disclosure. Another important aspect is to use the phase indication to perform accurate filtering (e.g., upsampling) of the decoded video. Upsampling with accurate phase shifting can also be useful for cases that do not involve rendering, such as when the decoded video is transcoded to a different format.
[0149] With respect to the example SEI message described above, in one example, the duration of the message may be specified by (1) the picture unit (PU) that contains the SEI message, or (2) the syntax of the SEI message. In one example, the phase indication SEI message applies to the currently decoded picture and persists for all pictures of the current layer that follow in output order until one or more of the following conditions are true: a new CLVS for the current layer starts; the bitstream ends; or a picture in the current layer that has an associated phase indication SEI message is output and follows the current picture in output order. In one example, the phase indication SEI message is applied to the current decoded picture and persists for all pictures of the current layer that follow in output order with the same value of ph_pic_parameter_set_id as the current picture until one or more of the following conditions become true: a new CLVS for the current layer starts; the bitstream ends; or a picture in the current layer with an associated phase indication SEI message and the same value of ph_pic_parameter_set_id as the current picture is output and follows the current picture in output order.
[0150] In one example, the phase indication SEI message is applied to the current cropped decoded picture having a cropped picture width and cropped picture height expressed in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively, and persists for all pictures of the current layer that follow in output order having the same CroppedWidth value as the current picture and the same CroppedHeight value as the current picture until one or more of the following conditions become true: a new CLVS for the current layer starts; the bitstream ends; or a picture in the current layer that has an associated phase indication SEI message and the same CroppedWidth value as the current picture and the same CroppedHeight value as the current picture is output and follows the current picture in output order.
[0151] In one example, the phase indication SEI message is applied to the current cropped decoded picture having a cropped picture width and cropped picture height expressed in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively, and persists for all pictures of the current layer that follow in output order having the same CroppedWidth value as the current picture and the same CroppedHeight value as the current picture until one or more of the following conditions become true: a new CLVS for the current layer starts; the bitstream ends; or a picture in the current layer having an associated phase indication SEI message and the same CroppedWidth value as the current picture and the same CroppedHeight value as the current picture is output and follows the current picture in output order.
[0152] In one example, the phase indication SEI message is applied to the current cropped decoded picture having a cropped picture width and cropped picture height expressed in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively, and persists for all pictures of the current layer that follow in output order having the same value of CroppedWidth as the current picture and the same value of CroppedHeight as the current picture until one or more of the following conditions become true: a new CLVS of the current layer starts; the bitstream ends; a picture in the current layer with an associated phase indication SEI is output and follows the current picture in output order, or a picture in the current layer that has a different value of CroppedWidth compared to the current picture or has a different value of CroppedHeight compared to the current picture is output and follows the current picture in output order.
[0153] In one example, an additional condition may be added to the list of conditions listed above: in one example, the condition "same value of CroppedWidth as the current picture and same value of CroppedHeight as the current picture" may be further augmented by "and same value of pps_pic_width_in_luma_samples as the current picture and same value of pps_pic_height_in_luma_samples as the current picture".
[0154] In this manner, source device 102 represents one example of a device that is configured to signal a phase indication information message corresponding to a coded video sequence, signal within the phase indication information message a syntax element that specifies a horizontal position of a luminance sampling location relative to a display window, where the horizontal position is represented as a signaled value of subtracting 128 from the specified horizontal position, dividing the difference by the phase denominator, and adding 0.5 to the quotient, and signal one or more syntax elements in the phase indication information message that specify a vertical position of a luminance sampling location relative to the display window, where the vertical position is represented as a signaled value of subtracting 128 from the specified vertical position, dividing the difference by the phase denominator, and adding 0.5 to the quotient.
[0155] 1, interface 108 may include any device configured to receive data generated by data encapsulator 107 and transmit and / or store the data on a communications medium. Interface 108 may include a network interface card, such as an Ethernet card, and may include an optical transceiver, a radio frequency transceiver, or any other type of device capable of transmitting and / or receiving information. Additionally, interface 108 may include a computer system interface that may allow files to be stored on a storage device. For example, interface 108 may include a Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, a proprietary bus protocol, a Universal Serial Bus (USB) protocol, an I / O protocol, or any other type of device that may be configured to transmit and / or receive information. 2 C, or any other logical and physical structure that can be used to interconnect peer devices.
[0156] 1, destination device 120 includes an interface 122, a data decapsulator 123, a video decoder 124, and a display 126. Interface 122 may include any device configured to receive data from a communication medium. Interface 122 may include a network interface card, such as an Ethernet card, and may include an optical transceiver, a radio frequency transceiver, or any other type of device capable of receiving and / or transmitting information. Additionally, interface 122 may include an interface for a computer system that allows a compliant video bitstream to be obtained from a storage device. For example, interface 122 may include interfaces for PCI and PCIe bus protocols, proprietary bus protocols, USB protocols, I / O protocols, and the like. 2 C, or any other logical and physical structures that may be used to interconnect peer devices. The data decapsulator 123 may be configured to receive and parse any of the example syntax structures described herein.
[0157] Video decoder 124 may include any device configured to receive a bitstream (e.g., a sub-bitstream extract) and / or an acceptable variant thereof and regenerate video data therefrom. Display 126 may include any device configured to display video data. Display 126 may include one of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display. Display 126 may include a high-resolution display or an ultra-high-resolution display. It should be noted that, although in the example shown in FIG. 1, video decoder 124 is described as outputting data to display 126, video decoder 124 may be configured to output video data to various types of devices and / or subcomponents thereof. For example, video decoder 124 may be configured to output video data to any communication medium as described herein.
[0158] FIG. 11 is a block diagram illustrating an example of a video decoding device that may be configured to decode video data according to one or more techniques of this disclosure (e.g., the decoding process for constructing a reference picture list described above). In one example, video decoding device 600 may be configured to decode transform data and recover residual data from transform coefficients based on the decoded transform data. Video decoding device 600 may be configured to perform intra-prediction decoding and inter-prediction decoding, and may therefore be referred to as a hybrid decoding device. Video decoding device 600 may be configured to parse any combination of syntax elements described above in Tables 1-17. Video decoding device 600 may render a picture based on or in accordance with the above-described process and further based on the parsed values in Tables 1-17.
[0159] In the example shown in FIG. 11, the video decoding device 600 includes an entropy decoding unit 602, an inverse quantization unit 604, an inverse transform processing unit 606, an intra prediction processing unit 608, an inter prediction processing unit 610, an adder 612, a post filter unit 614, and a reference buffer 616. The video decoding device 600 may be configured to decode video data in a manner consistent with a video encoding system. It should be noted that, although the example video decoding device 600 is shown having separate functional blocks, such illustration is for illustrative purposes and does not limit the video decoding device 600 and / or its subcomponents to a particular hardware or software architecture. The functionality of the video decoding device 600 may be realized using any combination of hardware, firmware, and / or software implementations.
[0160] As shown in FIG. 11, the entropy decoder 602 receives an entropy coded bitstream. The entropy decoder 602 may be configured to decode syntax elements and quantized coefficients from the bitstream according to a reciprocal process of the entropy coding process. The entropy decoder 602 may be configured to perform entropy decoding according to any of the entropy coding techniques described above. The entropy decoder 602 may determine values of syntax elements in the coded bitstream in a manner conforming to a video coding standard. As shown in FIG. 11, the entropy decoder 602 may determine quantization parameters, quantized coefficient values, transform data, and prediction data from the bitstream. In the example shown in FIG. 11, the inverse quantizer 604 and the inverse transform processor 606 receive the quantized coefficient values from the entropy decoder 602 and output reconstructed residual data.
[0161] Referring again to FIG. 11 , the reconstructed residual data may be provided to the adder 612. The adder 612 may add the reconstructed residual data to a prediction video block to generate reconstructed video data. The prediction video block may be determined according to a prediction video technique (i.e., intra prediction and inter-frame prediction). The intra prediction processor 608 may be configured to receive the intra prediction syntax element and obtain the prediction video block from the reference buffer 616. The reference buffer 616 may include a memory device configured to store one or more frames of video data. The intra prediction syntax element may identify an intra prediction mode, such as the intra prediction modes described above. The inter prediction processor 610 may receive the inter prediction syntax element and generate a motion vector that identifies a prediction block in one or more reference frames stored in the reference buffer 616. The inter prediction processor 610 may perform interpolation, possibly based on an interpolation filter, to generate a motion-interpolated block. The syntax element may include an identifier of an interpolation filter to be used for motion estimation with sub-pixel accuracy. The inter-prediction processor 610 may use an interpolation filter to calculate interpolated values for sub-integer pixels of the reference block. The post-filter unit 614 may be configured to perform filtering on the reconstructed video data. For example, the post-filter unit 614 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering, for example, based on parameters specified in the bitstream. It should also be noted that in some examples, the post-filter unit 614 may be configured to perform proprietary discretionary filtering (e.g., visual enhancement such as mosquito noise reduction). For example, the post-filter unit 614 may be configured to perform a rendering operation based on the received phase_indication() message described above. As shown in FIG. 11, the reconstructed video block may be output by the video decoding device 600.In this manner, the video decoding apparatus 600 represents an example of a device that is configured to receive a phase indication information message, determine a horizontal position of a luminance sampling location relative to a display window based on the parsed value of one or more syntax elements included in the phase indication information message, determine a vertical position of the luminance sampling location relative to the display window based on the parsed value of one or more syntax elements included in the phase indication information message, and render video based on the determined horizontal and vertical positions.
[0162] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processor. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium or a communication medium, including any medium that facilitates the transfer of a computer program from one place to another, for example according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0163] By way of example, and without limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, other magnetic storage, flash memory, or any other medium, i.e., any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer readable media.
[0164] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured to encode and decode, or incorporated into a composite codec. Also, the techniques may be implemented entirely in one or more circuit or logic elements.
[0165] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are illustrated in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but are not necessarily realized by different hardware units. Rather, as previously described, the various units may be combined into a codec hardware unit or may be provided by a collection of interoperating hardware units, including one or more processors as previously described, in conjunction with suitable software and / or firmware.
[0166] Furthermore, each functional block and various functions of the base station device and terminal device used in each of the above-mentioned implementations can be realized or performed by an electric circuit, which is generally an integrated circuit or a plurality of integrated circuits. The circuit designed to perform the functions described herein may comprise a general-purpose processor, a digital signal processor (DSP), an application specific or general-purpose application integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or individual hardware components, or a combination thereof. The general-purpose processor may be a microprocessor, or the processor may be a conventional processor, controller, microcontroller, or state machine. The general-purpose processor or each circuit described above may be composed of digital circuits or analog circuits. Furthermore, if an integrated circuit technology that replaces the current integrated circuits appears due to the progress of semiconductor technology, the integrated circuit using this technology can also be used.
[0167] Various embodiments have been described. These and other embodiments are within the scope of the following claims.
Claims
1. A video encoding device for encoding video data, comprising one or more processors, wherein the one or more processors: signaling a phase indication information message corresponding to the coded video sequence; signaling in the phase indication information message a first syntax element for indicating a horizontal phase numerator value; signaling in the phase indication information message a second syntax element for indicating a horizontal phase denominator value, wherein a horizontal position of a luma sampling location relative to a rendering window is equal to the quotient of the horizontal phase numerator value indicated in the first syntax element and the horizontal phase denominator value indicated in the second syntax element, expressed in units of the horizontal distance between two horizontally adjacent luma sampling locations; signaling in the phase indication information message a third syntax element for indicating a vertical phase numerator value; signaling in the phase indication information message a fourth syntax element for indicating a vertical phase denominator value, wherein a vertical position of a luma sampling location relative to a rendering window is equal to a quotient of the vertical phase numerator value indicated in the third syntax element and the vertical phase denominator value indicated in the fourth syntax element, expressed in units of a vertical distance between two vertically adjacent luma sampling locations; the phase indication information message is applied to the current cropped decoded picture and persists for all pictures of the current layer that follow in output order and have the same cropped picture width and cropped picture height as the current picture until a picture in the current layer that has an associated phase indication information message and the same cropped picture width and cropped picture height as the current picture has been output and follows the current picture in output order; Video encoding device.
2. A video decoding device for decoding video data, comprising one or more processors, the one or more processors: receiving a phase indication information message; Parsing a first syntax element in the phase indication information message to indicate a horizontal phase numerator value; Parsing a second syntax element in the phase indication information message to indicate a horizontal phase denominator value; the first syntax element and the second syntax element specify a horizontal position of a luma sampling location relative to a rendering window, the horizontal position being equal to the quotient of the horizontal phase numerator value and the horizontal phase denominator value, and expressed in units of the horizontal distance between two horizontally adjacent luma sampling locations; Parsing a third syntax element in the phase indication information message to indicate a vertical phase numerator value; Parsing a fourth syntax element in the phase indication information message to indicate a vertical phase denominator value; the third syntax element and the fourth syntax element are configured to specify a vertical position of a luma sampling location relative to a rendering window, the vertical position being equal to a quotient of the vertical phase numerator value and the vertical phase denominator value, expressed in units of a vertical distance between two vertically adjacent luma sampling locations; the phase indication information message is applied to the current cropped decoded picture and persists for all pictures of the current layer that follow in output order and have the same cropped picture width and cropped picture height as the current picture until a picture in the current layer that has an associated phase indication information message and the same cropped picture width and cropped picture height as the current picture has been output and follows the current picture in output order; Video decoding device.
3. A computer-readable storage medium storing a program for causing a computer to decode video data, the program being receiving a phase indication information message; Parsing a first syntax element in the phase indication information message to indicate a horizontal phase numerator value; Parsing a second syntax element in the phase indication information message to indicate a horizontal phase denominator value; the first syntax element and the second syntax element specify a horizontal position of a luma sampling location relative to a rendering window, the horizontal position being equal to the quotient of the horizontal phase numerator value and the horizontal phase denominator value, and expressed in units of the horizontal distance between two horizontally adjacent luma sampling locations; Parsing a third syntax element in the phase indication information message to indicate a vertical phase numerator value; Parsing a fourth syntax element in the phase indication information message to indicate a vertical phase denominator value; the third syntax element and the fourth syntax element specify a vertical position of a luma sampling location relative to a rendering window, the vertical position being equal to a quotient of the vertical phase numerator value and the vertical phase denominator value, expressed in units of a vertical distance between two vertically adjacent luma sampling locations; the phase indication information message is applied to the current cropped decoded picture and persists for all pictures of the current layer that follow in output order and have the same cropped picture width and cropped picture height as the current picture until a picture in the current layer that has an associated phase indication information message and the same cropped picture width and cropped picture height as the current picture has been output and follows the current picture in output order; A computer-readable storage medium.