Encoder, decoder and corresponding method

By treating sub-pictures as independent entities with clipping functions for motion vectors and interpolation filters, the method addresses errors in sub-picture decoding, ensuring accurate and efficient extraction in video coding systems.

JP7704491B2Active Publication Date: 2025-07-08HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021555037
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-29
Filing Date
2020-03-11
Publication Date
2025-07-08
Estimated Expiration
2040-03-11

AI Technical Summary

Technical Problem

Existing video coding systems face challenges in efficiently encoding and decoding sub-pictures due to errors caused by motion vectors pointing outside the sub-picture boundaries, leading to incomplete data availability and incorrect interpolation, which affects the separation and independent extraction of sub-pictures.

Method used

A method is introduced where a flag indicates that a sub-picture should be treated as a picture, applying a clipping function to motion vectors and interpolation filters to ensure they do not depend on data from adjacent sub-pictures, maintaining separation and enabling separate extraction.

Benefits of technology

This approach prevents errors during sub-picture extraction by ensuring that interpolation filters operate within sub-picture boundaries, allowing for accurate and independent decoding of sub-pictures without relying on data from other sub-pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704491000059
    Figure 0007704491000059
  • Figure 0007704491000060
    Figure 0007704491000060
  • Figure 0007704491000061
    Figure 0007704491000061
Patent Text Reader

Abstract

A video coding mechanism is disclosed. The mechanism includes receiving a bitstream including a current picture including a sub-picture coded according to inter prediction. Motion vectors for blocks of the sub-picture are determined. A clipping function is applied to sample positions in a reference block to support application of an interpolation filter when the motion vector points outside the sub-picture and when a flag is set to indicate that the sub-picture is treated as a picture. The interpolation filter is applied to the result of the clipping function to obtain predicted sample values. The block is decoded based on the predicted sample values. The block is transmitted for display as part of a decoded video sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [ Technical Field The present disclosure generally relates to video coding, and more particularly to coding sub-pictures of a picture in video coding.

Background Art

[0002] The amount of video data required to depict even relatively short videos can be substantial, which can cause difficulties when the data is streamed or communicated across a communication network having a limited bandwidth capacity. Thus, video data is generally compressed before being communicated across modern telecommunications networks. The size of the video can also be a problem when the video is stored on a storage device because memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Due to limited network resources and the ever-increasing demands for higher video quality, improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the image quality are desirable.

Summary of the Invention

[0003] In an embodiment, the present disclosure includes a method implemented in a decoder, the method comprising: receiving, by a receiver of the decoder, a bitstream including a current picture including a sub-picture coded according to inter prediction; determining, by a processor of the decoder, a motion vector for a block of the sub-picture; applying, by the processor, a clipping function to a sample position in a reference block to support application of an interpolation filter when the motion vector points outside the sub-picture and a flag is set to indicate that the sub-picture is to be treated as a picture; applying, by the processor, an interpolation filter to the result of the clipping function to obtain a predicted sample value; and decoding, by the processor, a block based on the predicted sample value. Inter prediction may be performed according to one of several inter prediction modes. A particular inter prediction mode generates a candidate list of motion vector predictors at both the encoder and the decoder. This enables signaling of the motion vector by signaling an index from the candidate list instead of the encoder signaling the entire motion vector. Further, some systems encode sub-pictures for independent extraction. This enables the current sub-picture to be decoded and displayed without decoding information from other sub-pictures. This may cause an error when a motion vector pointing outside the sub-picture is used. The reason is that the data pointed to by the motion vector may not be decoded and thus may not be available. The present disclosure includes a flag indicating that the sub-picture should be treated as a picture. When the current sub-picture is treated like a picture, the current sub-picture should be extracted without referring to other sub-pictures. Specifically, this example uses a clipping function applied when applying an interpolation filter. This clipping function ensures that the interpolation filter does not depend on data from adjacent sub-pictures in order to maintain separation between sub-pictures to support separate extraction.Therefore, the clipping function is applied when the flag is set and the motion vector points outside the current subpicture. The interpolation filter is then applied to the result of the clipping function. Thus, this example provides additional functionality to the video codec by preventing errors when performing subpicture extraction.

[0004] Optionally, in any of the above aspects, other implementations of the aspect provide that the interpolation filter includes a luma sample bilinear interpolation process, the block includes a block of luma samples, and the predicted sample value includes a predicted luma sample value.

[0005] Optionally, in any of the above aspects, other implementations of the aspect provide that the luma sample bilinear interpolation process receives an input including the luma position at all sample units (xIntL, yIntL), the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), and the clipping function is as follows, i.e., when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that a subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, Clip3 is a clipping function that follows the following,

Number

[0006] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the interpolation filter includes a luma sample 8-tap interpolation filtering process, the block includes a block of luma samples, and the predicted sample value includes a predicted luma sample value.

[0007] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the luma sample 8-tap interpolation filtering process receives an input including the luma position at all sample units (xIntL, yIntL), the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), and the clipping function follows the following, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i - 3), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i - 3) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, Clip3 is a clipping function that follows the following, [Number] Here, according to x, y, and z being numerical input values, it is provided to be applied to the sample position.

[0008] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the interpolation filter includes a chroma sample interpolation process, the block includes a block of chroma samples, and the predicted sample value includes a predicted chroma sample value.

[0009] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the chroma sample interpolation process receives an input including the chroma position at all sample units (xIntC, yIntC), the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC), and the clipping function follows the following, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xIntC + i), and yInti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yIntC + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, SubWidthC and SubHeightC indicate the horizontal and vertical sampling rate ratios between luma samples and chroma samples, Clip3 is a clipping function that follows the following,

Number

[0010] In an embodiment, the present disclosure includes a method implemented in an encoder, the method comprising: dividing a current picture into sub-pictures and further dividing the sub-pictures into blocks by a processor of the encoder; determining by the processor to encode a block according to inter prediction; selecting by the processor a motion vector for encoding the block; when the motion vector points outside the sub-picture and a flag is set to indicate that the sub-picture is to be treated as a picture, applying a clipping function to a sample position in a reference block to support the application of an interpolation filter by the processor; applying by the processor an interpolation filter to the result of the clipping function to obtain a predicted sample value; encoding by the processor the block into a bitstream based on the predicted sample value and the motion vector; and storing by a memory coupled to the processor the bitstream for communication to a decoder. Inter prediction may be performed according to one of several inter prediction modes. A particular inter prediction mode generates a candidate list of motion vector predictors at both the encoder and the decoder. This enables signaling of the motion vector by signaling an index from the candidate list instead of the encoder signaling the entire motion vector. Further, some systems encode sub-pictures for independent extraction. This enables the current sub-picture to be decoded and displayed without decoding information from other sub-pictures. This may cause an error when a motion vector pointing outside the sub-picture is used. The reason is that the data pointed to by the motion vector may not be decoded and thus may not be available. The present disclosure includes a flag indicating that the sub-picture should be treated as a picture. When the current sub-picture is treated as a picture, the current sub-picture should be extracted without referring to other sub-pictures. Specifically, this example uses a clipping function applied when applying an interpolation filter.This clipping function ensures that the interpolation filter does not depend on data from adjacent sub-pictures in order to maintain the separation between sub-pictures that support separate extractions. Thus, the clipping function is applied when the flag is set and the motion vector points outside the current sub-picture. The interpolation filter is then applied to the result of the clipping function. Thus, this example provides additional functionality to the video codec by preventing errors when performing sub-picture extraction.

[0011] Optionally, in any of the above aspects, other implementations of the aspect provide that the interpolation filter includes a luma sample bilinear interpolation process, the block includes a block of luma samples, and the predicted sample value includes a predicted luma sample value.

[0012] Optionally, in any of the above aspects, other implementations of the aspect provide that the luma sample bilinear interpolation process receives an input including the luma position at all sample units (xIntL, yIntL), the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), and the clipping function is as follows, i.e., when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that a subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, Clip3 is a clipping function that follows the following,

Number

[0013] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the interpolation filter includes a luma sample 8-tap interpolation filtering process, the block includes a block of luma samples, and the predicted sample value includes a predicted luma sample value.

[0014] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the luma sample 8-tap interpolation filtering process receives an input including the luma position at all sample units (xIntL, yIntL), the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), and the clipping function follows the following, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i - 3), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i - 3) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, Clip3 is a clipping function that follows the following,

Number

[0015] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the interpolation filter includes a chroma sample interpolation process, the block includes a block of chroma samples, and the predicted sample value includes a predicted chroma sample value.

[0016] Optionally, in any of the above aspects, other implementation manners of the aspect provide that the chroma sample interpolation process receives an input including the chroma position at all sample units (xIntC, yIntC), the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC), and the clipping function follows the following, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xIntC + i), and yInti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yIntC + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that the sub - picture is treated as a picture, SubPicIdx is the index of the sub - picture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the sub - picture, SubPicLeftBoundaryPos is the position of the left boundary of the sub - picture, SubPicTopBoundaryPos is the position of the upper boundary of the sub - picture, SubPicBotBoundaryPos is the position of the lower boundary of the sub - picture, SubWidthC and SubHeightC indicate the horizontal and vertical sampling rate ratios between luma samples and chroma samples, Clip3 is a clipping function that follows:

Number

[0017] In an embodiment, the present disclosure includes a video coding device including a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, and the processor, receiver, memory, and transmitter are configured to execute the method of any of the above aspects.

[0018] In an embodiment, the present disclosure includes a non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to execute any of the above-described methods.

[0019] In an embodiment, the present disclosure includes receiving means for receiving a bitstream including a current picture including a sub-picture coded according to inter prediction, determining means for determining a motion vector for a block of the sub-picture, and when the motion vector points outside the sub-picture and a flag is set to indicate that the sub-picture is to be treated as a picture, applying a clipping function to a sample position in a reference block to support the application of an interpolation filter, and applying means for applying an interpolation filter to the result of the clipping function to obtain a predicted sample value, decoding means for decoding the block based on the predicted sample value, and transfer means for transferring the block for display as part of a decoded video sequence.

[0020] Optionally, in any of the above aspects, another implementation of the aspect provides that the decoder is further configured to execute any of the above-described methods.

[0021] In an embodiment, the present disclosure includes partitioning means for partitioning a current picture into sub-pictures and partitioning the sub-pictures into blocks, determination means for determining to encode a block according to inter prediction, selection means for selecting a motion vector for encoding a block, when the motion vector points outside the sub-picture and a flag is set to indicate that the sub-picture is to be treated as a picture, applying a clipping function to a sample position in a reference block to support the application of an interpolation filter, applying means for applying an interpolation filter to the result of the clipping function to obtain a predicted sample value, encoding means for encoding a block into a bitstream based on the predicted sample value and the motion vector, and storage means for storing the bitstream for communication to a decoder.

[0022] Optionally, in any of the above aspects, another implementation of the aspect provides that the encoder is further configured to execute any of the methods of the above aspects.

[0023] Optionally, in any of the above aspects, another implementation of the aspect provides that the encoder is further configured to execute any of the methods of the above aspects.

[0024] For the purpose of clarity, any one of the above embodiments may be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present disclosure.

[0025] These and other features will be more clearly understood from the following detailed description considered in connection with the accompanying drawings and the claims.

Brief Description of the Drawings

[0026] For a more complete understanding of the present disclosure, reference is now made to the following brief description considered in connection with the accompanying drawings and detailed description, in which like reference numerals represent like parts.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Best Mode for Carrying Out the Invention

[0027] First, exemplary implementation manners of one or more embodiments are provided below, but it should be understood that the disclosed system and / or method may be implemented using any number of technologies, whether currently known or existing. The present disclosure should in no way be limited to the exemplary implementation manners, drawings, and technologies shown below, including the exemplary design and implementation manners exemplified and described herein, but may be modified within the scope of the appended claims, together with the full scope of these equivalents.

[0028] The following abbreviations, namely, Adaptive Loop Filter (ALF), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion-Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sample Adaptive Offset (SAO), Sequence Parameter Set (SPS), Temporal Motion Vector Prediction (TMVP), Versatile Video Coding (VVC), and Working Draft (WD) are used herein.

[0029] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques can include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a part of a video picture) may be divided into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTB), coding tree units (CTU), coding units (CU), and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks within an inter-coded unidirectional prediction (P) or bidirectional prediction (B) slice of a picture may be coded by using spatial prediction with respect to reference samples in adjacent blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or an image, and a reference picture may be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a prediction block representing the image block. Residual data represents the pixel difference between the original image block and the prediction block. Thus, an inter-coded block is coded according to a motion vector pointing to a block of reference samples forming the prediction block and residual data indicating the difference between the coded block and the prediction block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to the transform domain. These result in residual transform coefficients, which may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients.Entropy coding may be applied to achieve even more compression. Such video compression techniques will be described in more detail below.

[0030] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded in accordance with the corresponding video coding standard. Video coding standards include H.261 of the International Telecommunication Union (ITU) Standardization Sector (ITU-T), Part 2 of the Moving Picture Experts Group (MPEG)-1 of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC), H.262 of ITU-T or Part 2 of MPEG-2 of ISO / IEC, H.263 of ITU-T, Part 2 of MPEG-4 of ISO / IEC, Advanced Video Coding (AVC), also known as H.264 of ITU-T or Part 10 of MPEG-4 of ISO / IEC, and High Efficiency Video Coding (HEVC), also known as H.265 of ITU-T or Part 2 of MPEG-H. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).The ITU-T and ISO / IEC's Joint Video Experts Team (JVET) has started developing a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD), which includes JVET-M1001-v6 that provides algorithm descriptions, encoder-side descriptions of the VVC WD, and reference software.

[0031] To code a video image, the image is first segmented, and the segments are coded into a bitstream. Various picture segmentation methods are available. For example, an image can be segmented into normal slices, dependent slices, tiles, and / or according to Wavefront Parallel Processing (WPP). For simplicity, HEVC restricts the encoder so that only normal slices, dependent slices, tiles, WPP, and combinations thereof can be used when segmenting slices into groups of CTBs for video coding. Such segmentation can be applied to support matching of the Maximum Transfer Unit (MTU) size, parallel processing, and reduced end-to-end delay. The MTU indicates the maximum amount of data that can be transmitted in a single packet. If the packet payload exceeds the MTU, the payload is split into two packets through a process called fragmentation.

[0032] A normal slice, also simply called a slice, is a segmented part of an image that can be reproduced independently of other normal slices within the same picture despite some dependencies, for loop filtering operations. Each normal slice is encapsulated into its own Network Abstraction Layer (NAL) unit for transmission. Further, the in-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries may be disabled to support independent reproduction. Such independent reproduction supports parallelization. For example, normal slice-based parallelization uses minimal inter-processor or inter-core communication. However, since each normal slice is independent, each slice is associated with an individual slice header. The use of normal slices may incur substantial coding overhead due to the bit cost of the slice header for each slice and the lack of prediction across slice boundaries. Further, normal slices may be used to support matching for MTU size requirements. Specifically, since normal slices can be encapsulated into individual NAL units and coded independently, each normal slice should be smaller than the MTU in the MTU scheme to avoid splitting the slice into multiple packets. Therefore, the goals of parallelization and MTU size matching may impose conflicting requirements on the slice layout within a picture.

[0033] A dependent slice is similar to a normal slice but has a shortened slice header and enables segmentation at picture tree block boundaries without breaking in-picture prediction. Thus, a dependent slice enables a normal slice to be fragmented into multiple NAL units, which provides reduced end-to-end latency by allowing a part of the normal slice to be sent before the coding of the entire normal slice is complete.

[0034] A tile is a segmented part of an image created by horizontal and vertical boundaries that create columns and rows of tiles. Tiles may be coded in raster scan order (right to left and top to bottom). The scan order of CTBs is local within a tile. Thus, CTBs within the first tile are coded in raster scan order before proceeding to CTBs in the next tile. Similar to normal slices, tiles break picture-in-prediction dependencies and entropy decoding dependencies. However, tiles may not be contained in individual NAL units and thus may not be used for MTU size matching. Each tile may be processed by one processor / core, and inter-processor / inter-core communication used for intra-picture prediction between processing units decoding adjacent tiles may be limited to communicating the shared slice header and performing sharing related to loop filtering of the reconstructed samples and metadata (when adjacent tiles are within the same slice). When more than one tile is included in a slice, the entry point byte offset for each tile other than the first entry point offset within the slice may be signaled in the slice header. For each slice and tile, at least one of the following conditions, namely, 1) all coded tree blocks within the slice belong to the same tile, and 2) all coded tree blocks within the tile belong to the same slice, should be satisfied.

[0035] In WPP, the picture is partitioned into single rows of CTBs. The entropy decoding and prediction mechanisms may use data from CTBs in other rows. Parallel processing is enabled through parallel decoding of CTB rows. For example, the current row may be decoded in parallel with the previous row. However, the decoding of the current row is delayed from the decoding process of the row only two CTBs before. This delay ensures that data related to the CTB above and the top-right CTB of the current CTB in the current row is available before the current CTB is coded. This approach appears as a wavefront when represented graphically. This staggered start enables parallelization with up to the same number of processors / cores as the picture has CTB rows. Since intra-picture prediction between adjacent tree-block rows within a picture is allowed, the inter-processor / inter-core communication for enabling intra-picture prediction can be substantial. The WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support matching of MTU sizes. However, normal slices can be used with respect to WPP with a specific coding overhead to achieve matching of MTU sizes as desired.

[0036] The tile may also include a motion constrained tile set. A motion constrained tile set (MCTS) is a tile set designed such that the associated motion vectors are constrained to all sample positions within the MCTS and fractional sample positions that require only all sample positions within the MCTS for interpolation. Further, the use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is prohibited. Thus, each MCTS may be decoded independently without the presence of tiles not included in the MCTS. A supplemental enhancement information (SEI) message for the temporal MCTS may indicate the presence of the MCTS in the bitstream and may be used to signal the MCTS. The SEI message for the MCTS provides supplemental information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a compliant bitstream for the MCTS set. The information includes a number of extraction information sets, each defining a number of MCTS sets and including the raw bytes sequence payload (RBSP) bytes of the video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) that should be used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced, and the slice header may be updated since one or all of the syntax elements related to slice addresses (including first_slice_segment_in_pic_flag and slice_segment_address) may use different values in the extracted sub-bitstream.

[0037] The picture may also be divided into one or more sub - pictures. Dividing the picture into sub - pictures can enable different parts of the picture to be treated differently from the perspective of coding. For example, a sub - picture can be extracted and displayed without extracting other sub - pictures. As another example, different sub - pictures can be displayed at different resolutions, (e.g., in a video conferencing application) rearranged relative to each other, or coded as separate pictures even if the sub - pictures together contain data from a common picture.

[0038] Exemplary implementation manners of sub - pictures are as follows. A picture can be divided into one or more sub - pictures. A sub - picture is a rectangular or square set of slice / tile groups starting with a slice / tile group having an address equal to 0. Each sub - picture may refer to a different PPS, and thus each sub - picture may use a different partitioning mechanism. A sub - picture may be treated like a picture in the decoding process. The current reference picture used to decode the current sub - picture may be generated by extracting an area at the same position as the current sub - picture from the reference pictures in the decoded picture buffer. The extracted area may be the decoded sub - picture, and thus, inter - prediction may be performed between sub - pictures of the same size and at the same position within the picture. A tile group may be a sequence of tiles in the tile raster scan of a sub - picture. The following may be derived to determine the position of sub - pictures within a picture. Each sub - picture may be included at the next non - occupied position in the CTU raster scan order within the picture that is large enough to fit the sub - picture within the picture boundaries.

[0039] The sub-picture methods used by various video coding systems include various problems that reduce coding efficiency and / or functionality. This disclosure includes various solutions to such problems. In a first exemplary problem, inter prediction may be performed according to one of several inter prediction modes. A particular inter prediction mode generates a candidate list of motion vector predictors at both the encoder and the decoder. This enables signaling of the motion vector by signaling an index from the candidate list instead of the encoder signaling the overall motion vector. Further, some systems encode sub-pictures for independent extraction. This enables the current sub-picture to be decoded and displayed without decoding information from other sub-pictures. This can cause an error when motion vectors pointing outside the sub-picture are used. The reason for this is that the data pointed to by the motion vector may not be decoded and thus may not be available.

[0040] Accordingly, in a first example, a flag is disclosed here that indicates that the sub-picture should be treated as a picture. This flag is set to support separate extraction of the sub-picture. When the flag is set, the motion vector predictors obtained from blocks at the same position include only motion vectors pointing within the sub-picture. Any motion vector predictor pointing outside the sub-picture is excluded. This ensures that motion vectors pointing outside the sub-picture are not selected and related errors are avoided. Blocks at the same position are blocks from a picture different from the current picture. Motion vector predictors from blocks within the current picture (blocks not at the same position) may point outside the sub-picture since other processes such as an interpolation filter can prevent errors for such motion vector predictors. Thus, this example provides additional functionality to a video encoder / decoder (codec) by preventing errors when performing sub-picture extraction.

[0041] In a second example, a flag is disclosed here indicating that a subpicture should be treated as a picture. When the current subpicture is treated as a picture, the current subpicture should be extracted without referring to other subpictures. Specifically, this example uses a clipping function that is applied when applying an interpolation filter. This clipping function ensures that the interpolation filter does not depend on data from adjacent subpictures in order to maintain the separation between subpictures to support separate extractions. Thus, the clipping function is applied when the flag is set and the motion vector points outside the current subpicture. The interpolation filter is then applied to the result of the clipping function. Thus, this example provides additional functionality to the video codec by preventing errors when performing subpicture extraction. Thus, the first example and the second example address the first exemplary problem.

[0042] In a second exemplary problem, a video coding system divides a picture into sub-pictures, slices, tiles and / or coding tree units, which are then further divided into blocks. These blocks are then encoded for transmission to a decoder. Decoding such blocks can result in a decoded picture that includes various types of noise. To correct such problems, the video coding system may apply various filters across block boundaries. These filters can remove blocking, quantization noise and other unwanted coding artifacts. As described above, some systems encode sub-pictures for independent extraction. This allows the current sub-picture to be decoded and displayed without decoding information from other sub-pictures. In such a system, the sub-picture may be divided into blocks for encoding. Thus, block boundaries along a sub-picture edge may be aligned with the sub-picture boundary. In some cases, the block boundaries may also be aligned with the tile boundary. Filters may be applied across such block boundaries and thus across sub-picture boundaries and / or tile boundaries. This can cause errors when the current sub-picture is independently extracted because the filtering process may operate in an unexpected manner when data from adjacent sub-pictures is unavailable.

[0043] In a third example, a flag for controlling filtering at the sub-picture level is disclosed herein. When the flag is set for a sub-picture, the filter can be applied across the sub-picture boundary. When the flag is not set, the filter is not applied across the sub-picture boundary. Thus, the filter can be turned off for sub-pictures encoded for separate extraction or turned on for sub-pictures encoded for display as a group. Accordingly, this example provides additional functionality to a video codec by preventing filter-related errors when performing sub-picture extraction.

[0044] In a fourth example, a flag is disclosed here that can be set to control filtering at the tile level. When the flag is set for a tile, the filter can be applied across the tile boundary. When the flag is not set, the filter is not applied across the tile boundary. Thus, the filter can be turned off or on for use at the tile boundary (while, for example, continuing to filter the interior of the tile). Accordingly, this example provides an additional function to the video codec by supporting selective filtering across tile boundaries. Thus, the third and fourth examples address the second exemplary problem.

[0045] In a third exemplary problem, a video coding system may divide a picture into sub-pictures. This enables different sub-pictures to be treated differently when coding video. For example, sub-pictures can be separately extracted and displayed, and can be resized independently based on application-level changes, etc. In some cases, sub-pictures may be created by dividing a picture into tiles and assigning tiles to sub-pictures. Some video coding systems describe sub-picture boundaries with respect to the tiles included in the sub-picture. However, the tile approach may not be used for some pictures. Thus, such boundary descriptions may limit the use of sub-pictures to pictures that use tiles.

[0046] In a fifth example, a mechanism for signaling sub-picture boundaries with respect to CTBs and / or CTUs is disclosed herein. Specifically, the width and height of a sub-picture can be signaled in units of CTBs. Also, the position of the top-left CTU of the sub-picture can be signaled as an offset from the top-left CTU of the picture as measured in CTBs. The CTU and CTB sizes may be set to predetermined values. Thus, signaling the sub-picture dimensions and position with respect to CTBs and CTUs provides sufficient information for a decoder to position the sub-picture for display. This enables the use of sub-pictures even when tiles are not used. Also, this signaling mechanism can be coded using a relatively small number of bits while avoiding complexity. Thus, this example provides additional functionality to a video codec by enabling sub-pictures to be used independently of tiles. Further, this example increases coding efficiency and thus reduces the use of processor, memory, and / or network resources in an encoder and / or decoder. Thus, the fifth example addresses the third exemplary problem.

[0047] In a fourth exemplary problem, a picture can be partitioned into a plurality of slices for encoding. In some video coding systems, slices are addressed based on these positions relative to the picture. Still other video coding systems use the concept of sub-pictures. As noted above, sub-pictures can be treated differently from other sub-pictures from a coding perspective. For example, a sub-picture can be extracted and displayed independently of other sub-pictures. In such cases, slice addresses generated based on picture position may not operate properly because a significant number of assumed slice addresses are omitted. Some video coding systems address this problem by dynamically rewriting the slice header in response to a request to change the slice address to support sub-picture extraction. Such a process can be resource intensive because this process can occur each time a user requests to view a sub-picture.

[0048] In a sixth example, a slice addressed to a sub-picture including the slice is disclosed herein. For example, the slice header may include a sub-picture identifier (ID) and the address of each slice included in the sub-picture. Additionally, a sequence parameter set (SPS) may include the dimensions of the sub-picture that can be referenced by the sub-picture ID. Thus, when separate extraction of sub-pictures is requested, the slice header need not be rewritten. The slice header and SPS contain sufficient information to support positioning the slices within the sub-picture for display. Thus, this example increases coding efficiency and / or avoids redundant rewriting of the slice header, and thus reduces the use of processor, memory and / or network resources in the encoder and / or decoder. Thus, the sixth example addresses the fourth exemplary problem.

[0049] FIG. 1 is a flowchart of an exemplary operating method 100 for coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by using various mechanisms to reduce the video file size. A smaller file size enables the compressed video file to be transmitted to the user, while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reproduce the original video signal for display to the end user. The decoding process generally mirrors the encoding process to enable the decoder to consistently reproduce the video signal.

[0050] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that give a visual impression of movement when viewed in sequence. The frames include pixels that are represented here in terms of light, called luma components (or luma samples), and colors, called chroma components (or color samples). In some examples, the frames may also include depth values to support three-dimensional displays.

[0051] In step 103, the video is segmented into blocks. The segmentation includes subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC), also known as H.265 and MPEG-H Part 2, a frame can first be divided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels × 64 pixels). A CTU contains both luma samples and chroma samples. A coding tree may be used to divide the CTU into blocks and then recursively subdivide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of a frame may be subdivided until the individual blocks contain relatively uniform illumination values. Further, the chroma component of a frame may be subdivided until the individual blocks contain relatively uniform color values. Thus, the segmentation mechanism varies depending on the content of the video frame.

[0052] In step 105, various compression mechanisms are used to compress the image blocks segmented in step 103. For example, inter prediction and / or intra prediction may be used. Inter prediction is designed to utilize the fact that objects within a common scene tend to appear in consecutive frames. Therefore, blocks representing objects in the reference frame do not need to be repeatedly described in adjacent frames. Specifically, an object such as a table can remain in a fixed position over multiple frames. Thus, the table is described once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism may be used to match objects across multiple frames. Additionally, a moving object may be represented across multiple frames, for example, due to the movement of the object or the camera. As a specific example, a video may show a car moving across the screen over multiple frames. To describe such movement, motion vectors can be used. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in the reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame.

[0053] Intra prediction encodes blocks within a common frame. Intra prediction utilizes the fact that the luma and chroma components tend to concentrate within the frame. For example, green fragments in a part of a tree tend to be located adjacent to similar green fragments. Intra prediction uses multiple direction prediction modes (e.g., 33 in HEVC), the planar mode, and the direct current (DC) mode. The direction mode indicates that the current block is similar / same as the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the end of the row. The planar mode indicates a smooth transition of light / color across the row / column by using a relatively constant slope when substantially changing values. The DC mode is used for boundary smoothing and indicates that the block is similar / same as the average value related to the samples of all adjacent blocks related to the angular direction of the direction prediction mode. Therefore, the intra prediction block can represent the image block as various relationship prediction mode values instead of the actual values. Furthermore, the inter prediction block can represent the image block as motion vector values instead of the actual values. In either case, the prediction block may not accurately represent the image block in some cases. Any difference is stored in the residual block. To further compress the file, a transform may be applied to the residual block.

[0054] In step 107, various filtering techniques may be applied. In HEVC, the filters are applied according to the in-loop filtering method. The above block-based prediction may result in the generation of a block-shaped image in the decoder. Further, the block-based prediction method may encode a block and then reproduce the encoded block for later use as a reference block. The in-loop filtering method repeatedly applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reproduced. Further, these filters reduce the artifacts in the reproduced reference block so that there is a low possibility of creating further artifacts in subsequent blocks encoded based on the reference block in which the artifacts are reproduced.

[0055] Once the video signal is segmented, compressed, and filtered, in step 109, the resulting data is encoded into a bitstream. The bitstream includes the above data and any signaling data desirable to support the proper reproduction of the video signal in the decoder. For example, such data may include segmentation data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. The creation of the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may occur continuously and / or simultaneously over many frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of explanation and is not intended to limit the video coding process to a specific order.

[0056] The decoder receives a bitstream and starts the decoding process at step 111. Specifically, the decoder uses an entropy decoding method to convert the bitstream into corresponding syntax and video data. At step 111, the decoder uses syntax data from the bitstream to determine the frame segmentation. The segmentation should match the result of the block segmentation at step 103. Here, the entropy coding / decoding used at step 111 will be described. The encoder makes many choices during the compression process, such as selecting a block segmentation method from several possible options based on the spatial positioning of the values within the input image. Signaling the exact option may require using a large number of bins. When used here, a bin is a binary value treated as a variable (e.g., a bit value that can vary depending on the context). Entropy coding enables the encoder to discard any option that is clearly infeasible in a particular case and leave a set of acceptable options. Then, a codeword is assigned to each acceptable option. The length of the codeword is based on the number of acceptable options (e.g., 1 bin for 2 options, 2 bins for 3 to 4 options, etc.). Then, the encoder encodes the codeword for the selected option. This method reduces the size of the codeword because, in contrast to uniquely indicating a selection from a potentially large set of all possible options, it is desirable for the codeword to uniquely indicate a selection from a small subset of acceptable options. Then, the decoder decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.

[0057] In step 113, the decoder performs block decoding. Specifically, the decoder uses inverse transformation to generate residual blocks. Then, the decoder uses the residual blocks and the corresponding prediction blocks to reproduce the image blocks according to the partition. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks as generated by the encoder in step 105. Then, the reproduced image blocks are positioned in the frame of the reproduced video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding as described above.

[0058] In step 115, filtering is performed on the frame of the reproduced video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal can be output to the display in step 117 for viewing by the end user.

[0059] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functions to support the implementation of the operation method 100. The codec system 200 is generalized to show components used in both the encoder and the decoder. As described with respect to steps 101 and 103 in the operation method 100, the codec system 200 receives and partitions a video signal, which results in a partitioned video signal 201. Next, as described with respect to steps 105, 107, and 109 in method 100, when functioning as an encoder, the codec system 200 compresses the partitioned video signal 201 into a coded bitstream. When functioning as a decoder, the codec system 200 generates an output video signal from the bitstream, as described with respect to steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes an overall encoder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header format and context adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, the solid lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may be present in the encoder. The decoder may include a subset of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. Here, these components will be described.

[0060] The segmented video signal 201 is a captured video sequence that is segmented into blocks of pixels by a coding tree. The coding tree uses various splitting modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into even smaller blocks. The blocks may be referred to as nodes on the coding tree. Larger parent nodes are split into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. In some cases, the segmented blocks can be included in a coding unit (CU). For example, a CU can be a sub-part of a CTU that includes a luma block, a chroma red difference (Cr) block, and a chroma blue difference (Cb) block, along with the corresponding syntax instruction for the CU. The splitting modes may include a binary tree (BT), a triple tree (TT), and a quad tree (QT) that are used to divide a node into two, three, or four child nodes of varying shapes depending on the splitting mode used. The segmented video signal 201 is transferred for compression to an overall coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221.

[0061] The overall coder control component 211 is configured to make determinations related to the coding of images of a video sequence into a bitstream in accordance with the constraints of the application. For example, the overall coder control component 211 manages the optimization of the bitrate / bitstream size with respect to the reproduction quality. Such determinations may be made based on the availability of memory space / bandwidth and the image resolution requirements. The overall coder control component 211 also manages the utilization of the buffer in consideration of the transmission speed in order to mitigate the problems of buffer underrun and overrun. To manage these problems, the overall coder control component 211 manages the partitioning, prediction, and filtering by other components. For example, the overall coder control component 211 may dynamically increase the complexity of compression to increase the resolution and thereby increase the use of bandwidth, or may decrease the complexity of compression to decrease the resolution and the use of bandwidth. Thus, the overall coder control component 211 controls the other components of the codec system 200 in order to balance the concerns of bitrate and the video signal reproduction quality. The overall coder control component 211 creates control data for controlling the operation of the other components. The control data is also transferred to the header formatter and the CABAC component 231 for being encoded into a bitstream for signaling the parameters for decoding at the decoder.

[0062] The partitioned video signal 201 is also transmitted to the motion estimation component 221 and the motion compensation component 219 for inter prediction. A frame or slice of the partitioned video signal 201 may be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video blocks with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of the video data.

[0063] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating motion vectors that estimate the motion for video blocks. The motion vectors may indicate, for example, the displacement of an object coded for a prediction block. The prediction block is a block that has been found to closely match the block to be coded in terms of pixel differences. The prediction block may also be referred to as a reference block. Such pixel differences may be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. HEVC uses several coded objects including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into CTBs, which can then be divided into CUs for inclusion in CBs. A CU can be coded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing transform residual data for the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and may select the reference block, motion vector, etc. having the best rate-distortion characteristics. The best rate-distortion characteristics balance both the quality of video reproduction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final coding).

[0064] In some examples, the codec system 200 may calculate values for sub-integer pixel positions of reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values for 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation component 221 may perform motion searches for both full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates motion vectors for PUs of video blocks in an inter-coded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header format and the CABAC component 231 for encoding, and outputs the motion to the motion compensation component 219.

[0065] Motion compensation performed by the motion compensation component 219 may involve fetching or generating a predicted block based on the motion vector determined by the motion estimation component 221. Similarly, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving a motion vector for the current video block's PU, the motion compensation component 219 may identify the predicted block pointed to by the motion vector. Then, by subtracting the pixel values of the predicted block from the pixel values of the currently coded video block, a residual video block is formed, forming pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The predicted block and the residual block are transferred to the transform scaling and quantization component 213.

[0066] The separated video signal 201 is also sent to the intra picture estimation component 215 and the intra picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217, as described above, perform intra prediction of the current block for blocks within the current frame as an alternative to the inter prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames. In particular, the intra picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra prediction mode for encoding the current block from a plurality of tested intra prediction modes. The selected intra prediction mode is then transferred to the header format and the CABAC component 231 for encoding.

[0067] For example, the intra picture estimation component 215 calculates rate-distortion values using rate-distortion analysis for various tested intra prediction modes and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the original unencoded block and the encoded block used to generate the encoded block, and the bit rate (e.g., the number of bits) used to generate the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. Further, the intra picture estimation component 215 may be configured to code depth blocks of the depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).

[0068] When implemented in the encoder, the intra picture prediction component 217 may generate a residual block from a prediction block based on the selected intra picture prediction mode determined by the intra picture estimation component 215, or when implemented in the decoder, may read the residual block from the bitstream. The residual block includes the difference in values between the prediction block and the original block, represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra picture estimation component 215 and the intra picture prediction component 217 may be executed for both the luma component and the chroma component.

[0069] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information such that different frequency information is quantized at different granularities, which may affect the final visual quality of the reproduced video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header format and the CABAC component 231 for encoding into the bitstream.

[0070] The scaling and inverse transform component 229 applies the inverse operation of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization, for example, to reproduce the residual block in the pixel domain for later use as a reference block that can be a predicted block for other current blocks. The motion estimation component 221 and / or the motion compensation component 219 may calculate a reference block by returning the residual block to the corresponding predicted block and adding them for use in motion estimation of subsequent blocks / frames. A filter is applied to the reproduced reference block to reduce artifacts created during scaling, quantization, and transformation. Otherwise, such artifacts may cause inaccurate predictions (and create further artifacts) when subsequent blocks are predicted.

[0071] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reproduced image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reproduce the original image block. Then, the filter may be applied to the reproduced image block. In some examples, the filter may be applied to the residual block instead. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reproduced reference block is applied to a specific spatial region and includes a plurality of parameters to adjust how such a filter is applied. The filter control analysis component 227 analyzes the reproduced reference block and sets the corresponding parameters to determine where such a filter should be applied. Such data is transferred as filter control data for encoding to the header format and the CABAC component 231. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied in the spatial / pixel domain (e.g., the reproduced pixel block) or the frequency domain depending on the example.

[0072] When operating as an encoder, the filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks and transfers them towards the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0073] The header format and the CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. Specifically, the header format and the CABAC component 231 generate various headers for encoding control data such as overall control data and filter control data. Further, prediction data including intra prediction and motion data and residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder for reproducing the original segmented video signal 201. Such information may also include an intra prediction mode index table (also called a codeword mapping table), definitions of coding contexts for various blocks, an indication of the most accurate intra prediction mode, an indication of segmentation information, etc. Such data may be encoded by using entropy coding. For example, the information may be encoded by using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. Following the entropy coding, the coded bitstream may be transmitted to other devices (e.g., a video decoder), or may be archived for later transmission or acquisition.

[0074] FIG. 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 divides the input video signal, resulting in a divided video signal 301, which is substantially the same as the divided video signal 201. The divided video signal 301 is then compressed by components of the encoder 300 and encoded into a bitstream.

[0075] Specifically, the divided video signal 301 is transferred to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. The divided video signal 301 is also transferred to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to a transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially the same as the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (along with associated control data) are transferred to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially the same as the header format and CABAC component 231.

[0076] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 in order to reproduce them as reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. The in-loop filter within the in-loop filter component 325 is also applied to the residual blocks and / or the reproduced reference blocks, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters, as described for the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.

[0077] FIG. 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reproduced based on the bitstream for display to an end user.

[0078] The bitstream is received by an entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide a context for interpreting further data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding the video signal, such as overall control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to an inverse transform and quantization component 429 for reproduction of the residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0079] The reproduced residual block and / or prediction block is transferred to the intra-picture prediction component 417 to reproduce the image block based on the intra prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses a prediction mode to identify a reference block within the frame and applies the residual block to the result to reproduce the intra-predicted image block. The reproduced intra-predicted image block and / or residual block and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, and these components may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reproduced image block, residual block and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reproduced image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses a motion vector from a reference block to generate a prediction block and applies the residual block to the result to reproduce the image block. The resulting reproduced block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store further reproduced image blocks, which can be reproduced into a frame via the segmentation information. Such a frame may also be arranged in a sequence. The sequence is output to a display as a reproduced output video signal.

[0080] FIG. 5A is a schematic diagram showing an exemplary picture 500 divided into sub-pictures 510. For example, picture 500 can be divided for encoding by codec system 200 and / or encoder 300, and can be divided for decoding by codec system 200 and / or decoder 400. As another example, picture 500 may be divided by the encoder in step 103 of method 100 for use by the decoder in step 111.

[0081] Picture 500 is an image showing the complete visual part of the video sequence at a specified time position. Picture 500 may also be referred to as an image and / or a frame. Picture 500 may be specified by a picture order count (POC). POC is an index indicating the output / display order of picture 500 in the video sequence. Picture 500 can be divided into sub-pictures 510. Sub-picture 510 is a rectangular or square area of one or more slices / tile groups within picture 500. Sub-picture 510 is optional, and thus some video sequences include sub-picture 510 while others do not. Four sub-pictures 510 are shown, but picture 500 can be divided into any number of sub-pictures 510. The division of sub-pictures 510 may be consistent across the entire coded video sequence.

[0082] Sub-picture 510 may be used to enable different regions of picture 500 to be treated differently. For example, a specified sub-picture 510 may be extracted independently and transmitted to a decoder. As a specific example, a user using a virtual reality (VR) headset may view a subset of picture 500, which may provide the user with the impression of physically existing in the space as shown in picture 500. In such a case, streaming only the sub-pictures 510 that may be displayed to the user can increase coding efficiency. As another example, different sub-pictures 510 may be treated differently in a specific application. As a specific example, a video conferencing application may display an active speaker at a more prominent position and at a higher resolution than users who are not currently speaking. Positioning different users within different sub-pictures 510 supports real-time reconstruction of the displayed image to support this function.

[0083] Each sub-picture 510 can be identified by a unique sub-picture ID, which may be consistent for the entire CVS. For example, the upper-left sub-picture 510 of the picture 500 may have a sub-picture ID of 0. In such a case, the upper-left sub-picture 510 of any picture 500 in the sequence can be referenced by the sub-picture ID of 0. Further, each sub-picture 510 may include a defined configuration, which may be consistent for the entire CVS. For example, the sub-picture 510 may include a height, a width, and / or an offset. The height and width describe the size of the sub-picture 510, and the offset describes the position of the sub-picture 510. For example, the sum of all the widths of the sub-pictures 510 in a row is the width of the picture 500. Further, the sum of all the heights of the sub-pictures 510 in a column is the height of the picture 500. Further, the offset indicates the position of the upper-left corner of the sub-picture 510 relative to the upper-left corner of the picture 500. The height, width, and offset of the sub-picture 510 provide sufficient information to position the corresponding sub-picture 510 within the picture 500. Since the division of the sub-pictures 510 may be consistent across the entire CVS, the parameters related to the sub-pictures may be included in the sequence parameter set (SPS).

[0084] FIG. 5B is a schematic diagram showing an exemplary sub-picture 510 divided into slices 515. As shown, the sub-picture 510 of the picture 500 may include one or more slices 515. A slice 515 is an integral number of complete tiles or an integral number of consecutive complete CTU rows within a tile of a picture that is exclusively included in a single network abstraction layer (NAL) unit. Four slices 515 are shown, but the sub-picture 510 may include any number of slices 515. A slice 515 includes visual data specific to the picture 500 of a specified POC. Accordingly, the parameters related to the slice 515 may be included in a picture parameter set (PPS) and / or a slice header.

[0085] FIG. 5C is a schematic diagram showing an exemplary slice 515 divided into tiles 517. As shown, the slice 515 of the picture 500 may include one or more tiles 517. A tile 517 may be created by dividing the picture 500 into rectangular rows and columns. Accordingly, a tile 517 is a rectangular or square region of CTUs within a specific tile column and a specific tile row within the picture. Tiling is optional, and thus some video sequences include tiles 517 while others do not. Four tiles 517 are shown, but the slice 515 may include any number of tiles 517. A tile 517 may include visual data specific to the slice 515 of the picture 500 of a specified POC. In some cases, a slice 515 may also be included in a tile 517. Accordingly, the parameters related to the tile 517 may be included in the PPS and / or the slice header.

[0086] FIG. 5D is a schematic diagram showing an exemplary slice 515 partitioned into CTUs 519. As shown, a slice 515 of picture 500 (or a tile 517 of slice 515) may include one or more CTUs 519. A CTU 519 is a region of picture 500 subdivided by a coding tree to create coding blocks to be encoded / decoded. A CTU 519 may include luma samples for a monochrome picture 500, or a combination of luma samples and chroma samples for a color picture 500. The grouping of luma samples or chroma samples that can be partitioned by the coding tree is referred to as a coding tree block (CTB) 518. Thus, a CTU 519 includes a CTB 518 of luma samples and two corresponding CTBs 518 of chroma samples of picture 500 having three sample arrays, or a CTB 518 of samples of a picture coded using a syntax structure used to code a monochrome picture, or three separate color planes and samples.

[0087] As shown above, picture 500 may be divided into sub - pictures 510, slices 515, tiles 517, CTUs 519 and / or CTBs 518, and then these are further divided into blocks. Then, such blocks are encoded for transmission to a decoder. Decoding such blocks may result in a decoded image containing various types of noise. To correct such problems, a video coding system may apply various filters across block boundaries. These filters can remove blocking, quantization noise, and other unwanted coding artifacts. As described above, sub - picture 510 may be used when performing independent extraction. In this case, the current sub - picture 510 may be decoded and displayed without decoding information from other sub - pictures 510. Therefore, block boundaries along the sub - picture 510 edges may align with sub - picture boundaries. In some cases, block boundaries may also align with tile boundaries. Filters may be applied across such block boundaries and thus may be applied across sub - picture boundaries and / or tile boundaries. This may cause an error when the current sub - picture 510 is independently extracted because the filtering process may operate in an unexpected manner when data from adjacent sub - pictures 510 is unavailable.

[0088] To address these issues, a flag may be used to control filtering at the sub-picture 510 level. For example, the flag may be indicated as loop_filter_across_subpic_enabled_flag. When the flag is set for a sub-picture 510, the filter can be applied across the corresponding sub-picture boundary. When the flag is not set, the filter is not applied across the corresponding sub-picture boundary. Thus, the filter can be turned off for sub-pictures 510 that are encoded for separate extraction, or can be turned on for sub-pictures 510 that are encoded for display as a group. Other flags can be set to control filtering at the tile 517 level. The flag may be indicated as loop_filter_across_tiles_enabled_flag. When the flag is set for a tile 517, the filter can be applied across the tile boundary. When the flag is not set, the filter is not applied across the tile boundary. Thus, the filter can be turned off or on for use at the tile boundary (while, for example, continuing to filter within the tile). As used herein, the filter is applied across the sub-picture 510 or tile 517 boundary when the filter is applied to samples on both sides of the boundary.

[0089] Also, as described above, tiling is optional. However, some video coding systems describe sub-picture boundaries with respect to tiles 517 included in sub-picture 510. In such systems, the sub-picture boundary description with respect to tile 517 restricts the use of sub-picture 510 to picture 500 that uses tile 517. To expand the applicability of sub-picture 510, sub-picture 510 may be described with respect to CTB 518 and / or CTU 519 regarding the boundaries. Specifically, the width and height of sub-picture 510 can be signaled in units of CTB 518. Further, the position of the top-left CTU 519 of sub-picture 510 can be signaled as an offset from the top-left CTU 519 of picture 500 as measured by CTB 518. The CTU 519 and CTB 518 sizes may be set to predetermined values. Thus, signaling sub-picture dimensions and positions with respect to CTB 518 and CTU 519 provides sufficient information for the decoder to position sub-picture 510 for display. This enables sub-picture 510 to be used even when tile 517 is not used.

[0090] Furthermore, some video coding systems address slice 515 based on these positions relative to picture 500. This causes problems when sub-picture 510 is coded for independent extraction and display. In such cases, the slice 515 and corresponding addresses associated with the omitted sub-picture 510 are also omitted. The omission of the slice 515 addresses may prevent the decoder from properly positioning slice 515. Some video coding systems address this problem by dynamically rewriting the addresses in the slice header associated with slice 515. Since the user may request any sub-picture, such rewriting occurs every time the user requests a video, which is extremely resource-intensive. To overcome this problem, slice 515 is addressed relative to sub-picture 510 that includes slice 515 when sub-picture 510 is used. For example, slice 515 can be identified by an index or other value specific to sub-picture 510 that includes slice 515. The slice address can be coded in the slice header associated with slice 515. The sub-picture ID of sub-picture 510 that includes slice 515 can also be coded in the slice header. Further, the dimensions / configuration of sub-picture 510 can be coded in the SPS along with the sub-picture ID. Thus, the decoder can obtain the sub-picture 510 configuration from the SPS based on the sub-picture ID and position slice 515 relative to sub-picture 510 without referring to the complete picture 500. Thus, when sub-picture 510 is extracted, rewriting of the slice header can be omitted, which significantly reduces resource usage in the encoder, decoder, and / or corresponding slicer.

[0091] Once the picture 500 is partitioned into CTB518 and / or CTU519, CTB518 and / or CTU519 can be further divided into coding blocks. Then, the coding blocks can be coded according to intra prediction and / or inter prediction. The present disclosure also includes improvements related to the inter prediction mechanism. The inter prediction can be performed in several different modes that can operate according to uni - directional inter prediction and / or bi - directional inter prediction.

[0092] FIG. 6 is a schematic diagram showing an example of uni - directional inter prediction 600 that is executed to determine a motion vector in, for example, block compression step 105, block decoding step 113, motion estimation component 221, motion compensation component 219, motion compensation component 321, and / or motion compensation component 421. For example, the uni - directional inter prediction 600 can be used to determine the motion vector for an encoded and / or decoded block created when partitioning a picture such as picture 500.

[0093] The uni - directional inter prediction 600 uses a reference frame 630 having a reference block 631 to predict a current block 611 within the current frame 610. The reference frame 630 may be temporally located after the current frame 610 (e.g., as a subsequent reference frame) as shown in the figure, but in some examples, it may be temporally located before the current frame 610 (e.g., as a previous reference frame). The current frame 610 is an exemplary frame / picture encoded / decoded at a specific time. The current frame 610 includes an object within the current block 611 that matches an object within the reference block 631 of the reference frame 630. The reference frame 630 is a frame used as a reference for encoding the current frame 610, and the reference block 631 is a block within the reference frame 630 that includes an object also included in the current block 611 of the current frame 610.

[0094] The current block 611 is any coding unit that is encoded / decoded at a specified point in the coding process. The current block 611 may be the entire segmented block, or may be a sub-block when using the affine inter prediction mode. The current frame 610 is separated from the reference frame 630 by some temporal distance (TD) 633. The TD 633 indicates the amount of time between the current frame 610 and the reference frame 630 in the video sequence, and may be measured in units of frames. The prediction information for the current block 611 may refer to the reference frame 630 and / or the reference block 631 by a reference index indicating the direction and temporal distance between the frames. Over the period represented by the TD 633, the object within the current block 611 moves from its position within the current frame 610 to another position within the reference frame 630 (e.g., the position of the reference block 631). For example, the object may move along a motion trajectory 613 which is the direction of the movement of the object over time. The motion vector 635 describes the direction and magnitude of the movement of the object along the motion trajectory 613 over the TD 633. Thus, the encoded motion vector 635, the reference block 631, and the residual including the difference between the current block 611 and the reference block 631 provide sufficient information to reproduce the current block 611 and position the current block 611 within the current frame 610.

[0095] FIG. 7 is a schematic diagram showing an example of a bidirectional inter prediction 700 that is executed to determine an MV, for example, in a block compression step 105, a block decoding step 113, a motion estimation component 221, a motion compensation component 219, a motion compensation component 321, and / or a motion compensation component 421. For example, the bidirectional inter prediction 700 can be used to determine motion vectors for the encoded and / or decoded blocks created when segmenting a picture such as picture 500.

[0096] The bidirectional inter prediction 700 is similar to the unidirectional inter prediction 600, but uses a pair of reference frames to predict the current block 711 within the current frame 710. Thus, the current frame 710 and the current block 711 are substantially the same as the current frame 610 and the current block 611, respectively. The current frame 710 is temporally positioned between a previous reference frame 720 that occurs before the current frame 710 in the video sequence and a subsequent reference frame 730 that occurs after the current frame 710 in the video sequence. The previous reference frame 720 and the subsequent reference frame 730 are substantially the same as the reference frame 630 in other respects.

[0097] The current block 711 is matched with a preceding reference block 721 in a preceding reference frame 720 and a succeeding reference block 731 in a succeeding reference frame 730. Such matching indicates that, in the process of a video sequence, an object moves from the position of the preceding reference block 721 along a motion trajectory 713 through the current block 711 to the position of the succeeding reference block 731. The current frame 710 is separated from the preceding reference frame 720 by some preceding time distance (TD0) 723 and from the succeeding reference frame 730 by some succeeding time distance (TD1) 733. TD0 723 indicates, in terms of frames, the amount of time between the preceding reference frame 720 and the current frame 710 in the video sequence. TD1 733 indicates, in terms of frames, the amount of time between the current frame 710 and the succeeding reference frame 730 in the video sequence. Thus, the object moves from the preceding reference block 721 to the current block 711 along the motion trajectory 713 over the period indicated by TD0 723. The object also moves from the current block 711 to the succeeding reference block 731 along the motion trajectory 713 over the period indicated by TD1 733. Prediction information about the current block 711 may refer to the preceding reference frame 720 and / or the preceding reference block 721 and the succeeding reference frame 730 and / or the succeeding reference block 731 by a pair of reference indices indicating the direction and time distance between the frames.

[0098] The previous motion vector (MV0) 725 describes the direction and magnitude of the movement of an object along a motion trajectory 713 spanning TD0 723 (e.g., between the previous reference frame 720 and the current frame 710). The subsequent motion vector (MV1) 735 describes the direction and magnitude of the movement of an object along a motion trajectory 713 spanning TD1 733 (e.g., between the current frame 710 and the subsequent reference frame 730). Thus, in bidirectional inter prediction 700, the current block 711 can be coded and reproduced by using the previous reference block 721 and / or the subsequent reference block 731, MV0 725, and MV1 735.

[0099] In both the merge mode and the advanced motion vector prediction (AMVP) mode, the candidate list is generated by adding candidate motion vectors to the candidate list in the order defined by the candidate list determination pattern. Such candidate motion vectors may include motion vectors from uni-directional inter prediction 600, bidirectional inter prediction 700, or a combination thereof. Specifically, for an adjacent block, the motion vector is generated when such a block is coded. Such a motion vector is added to the candidate list for the current block, and the motion vector for the current block is selected from the candidate list. The motion vector can then be signaled as the index of the selected motion vector within the candidate list. The decoder can construct the candidate list using the same process as the encoder and determine the motion vector selected from the candidate list based on the signaled index. Thus, the candidate motion vectors include motion vectors generated according to uni-directional inter prediction 600 and / or bidirectional inter prediction 700, depending on which method is used when such adjacent blocks are coded.

[0100] FIG. 8 is a schematic diagram showing an example 800 of coding a current block 801 based on candidate motion vectors from adjacent coded blocks 802. The operation method 100 of the encoder 300 and / or decoder 400, and / or the use of the functions of the codec system 200 can use adjacent blocks 802 to generate a candidate list. Such a candidate list can be used in inter prediction by unidirectional inter prediction 600 and / or bidirectional inter prediction 700. The candidate list can then be used to encode / decrypt the current block 801, which may be generated by partitioning a picture such as picture 500.

[0101] The current block 801 is, depending on the example, a block that is being encoded by the encoder or decoded by the decoder at a specified time. The coded block 802 is a block that has already been encoded at a specified time. Therefore, the coded block 802 is potentially available for use in generating a candidate list. The current block 801 and the coded block 802 may be included in a common frame and / or may be included in temporally adjacent frames. When the coded block 802 is included in a common frame with the current block 801, the coded block 802 includes a boundary that is immediately adjacent to (e.g., touches) the boundary of the current block 801. When the coded block 802 is included in temporally adjacent frames, the coded block 802 is located at the same position as the position of the current block 801 in the current frame within the temporally adjacent frames. The candidate list can be generated by adding the motion vectors from the coded block 802 as candidate motion vectors. The current block 801 can then be coded by selecting a candidate motion vector from the candidate list and signaling the index of the selected candidate motion vector.

[0102] FIG. 9 is a schematic diagram showing an exemplary pattern 900 for determining a candidate list of motion vectors. Specifically, the operation method 100 of the encoder 300 and / or decoder 400, and / or the use of the functions of the codec system 200 can use the candidate list determination pattern 900 used when generating a candidate list 911 to encode the current block 801 segmented from the picture 500. The resulting candidate list 911 may be a merge candidate list or an AMVP candidate list, which can be used in inter prediction by one-way inter prediction 600 and / or bi-directional inter prediction 700.

[0103] When encoding the current block 901, the candidate list determination pattern 900 obtains valid candidate motion vectors and searches for positions 905 shown as A0, A1, B0, B1 and / or B2 within the same picture / frame as the current block 901. The candidate list determination pattern 900 may also obtain valid candidate motion vectors and search for a block 909 at the same position. The block 909 at the same position is a block at the same position as the current block 901 but is included in temporally adjacent pictures / frames. The candidate motion vectors can then be placed in the candidate list 911 in a predetermined inspection order. Thus, the candidate list 911 is a procedurally generated list of indexed candidate motion vectors.

[0104] The candidate list 911 can be used to select a motion vector for performing inter prediction for the current block 901. For example, the encoder can obtain samples of a reference block pointed to by a candidate motion vector from the candidate list 911. The encoder can then select a candidate motion vector that points to the reference block that most closely matches the current block 901. The index of the selected candidate motion vector can then be encoded to represent the current block 901. In some cases, the candidate motion vector points to a reference block that includes partial reference samples 915. In this case, the interpolation filter 913 can be used to reproduce a complete set of reference samples 915 to support motion vector selection. The interpolation filter 913 is a filter that can upsample a signal. Specifically, the interpolation filter 913 is a filter that can accept a partial / lower quality signal as input and determine an approximation of a more complete / higher quality signal. Thus, the interpolation filter 913 can be used to obtain a complete set of reference samples 915 for use in selecting a reference block for the current block 901 and thus in selecting a motion vector for encoding the current block 901 in certain cases.

[0105] The above mechanism for coding a block based on inter prediction by using a candidate list may cause certain errors when a subpicture such as subpicture 510 is used. Specifically, problems may occur when the current block 901 is included in the current subpicture but the motion vector points to a reference block that is at least partially located in an adjacent subpicture. In such a case, the current subpicture may be extracted for presentation without the adjacent subpicture. When this occurs, portions of the reference block within the adjacent subpicture may not be transmitted to the decoder and thus the reference block may not be available for use in decoding the current block 901. When this occurs, the decoder does not have access to sufficient data to decode the current block 901.

[0106] The present disclosure provides a mechanism for addressing this problem. In one example, a flag indicating that the current sub-picture should be treated as a picture is used. This flag can be set to support separate extraction of sub-pictures. Specifically, when the flag is set, the current sub-picture should be encoded without referring to data within other sub-pictures. In this case, the current sub-picture is treated like a picture in that it is coded separately from other sub-pictures and can be displayed as a separate picture. Thus, this flag may be denoted as subpic_treated_as_pic_flag[i], where i is the index of the current sub-picture. When the flag is set, the motion vector candidates (also known as motion vector predictors) obtained from block 909 at the same location include only motion vectors pointing within the current sub-picture. Any motion vector predictor pointing outside the current sub-picture is excluded from candidate list 911. This ensures that a motion vector pointing outside the current sub-picture is not selected and associated errors are avoided. This example particularly applies to motion vectors from block 909 at the same location. Motion vectors from search position 905 within the same picture / frame may be corrected by a separate mechanism as described below.

[0107] Another example may be used to address the search position 905 when the current subpicture is treated as a picture (e.g., when subpic_treated_as_pic_flag[i] is set). When the current subpicture is treated as a picture, the current subpicture should be extracted without referring to other subpictures. An exemplary mechanism relates to the interpolation filter 913. The interpolation filter 913 can be applied to samples at one position to interpolate (e.g., predict) related samples at other positions. In this example, the motion vector from the coded block at the search position 905 may point to such a reference sample 915 as long as the interpolation filter 913 can interpolate reference samples 915 outside the current subpicture based only on the reference sample 915 from the current subpicture. Thus, this example uses a clipping function that is applied when applying the interpolation filter 913 to motion vector candidates from the search position 905 from the same picture. This clipping function clips data from adjacent subpictures and thus removes such data as input to the interpolation filter 913 when determining the reference sample 915 pointed to by the motion vector candidate. This approach maintains the separation between subpictures during encoding to support separate extraction and decoding when the subpicture is treated as a picture. The clipping function may be applied to the luma sample bilinear interpolation process, the luma sample 8-tap interpolation filtering process, and / or the chroma sample interpolation process.

[0108] FIG. 10 is a block diagram showing an exemplary in-loop filter 1000. The in-loop filter 1000 may be used to implement in-loop filters 225, 325, and / or 425. Further, the in-loop filter 1000 may be applied in an encoder and a decoder when method 100 is executed. Further, the in-loop filter 1000 can be applied to filter a current block 801 partitioned from picture 500, which may be coded according to uni-directional inter prediction 600 and / or bi-directional inter prediction 700 based on a candidate list generated according to pattern 900. The in-loop filter 1000 includes a deblocking filter 1043, a SAO filter 1045, and an adaptive loop filter (ALF) 1047. The filters of the in-loop filter 1000 are applied to the reproduced image blocks in order in the encoder (e.g., before being used as a reference block) and in the decoder before display.

[0109] The deblocking filter 1043 is configured to remove block-shaped edges created by block-based inter and intra prediction. The deblocking filter 1043 scans an image portion (e.g., an image slice) to find discontinuities in chroma and / or luma values occurring at partition boundaries. The deblocking filter 1043 then applies a smoothing function to the block boundaries to remove such discontinuities. The strength of the deblocking filter 1043 may be changed depending on the spatial activity (e.g., the variance of the luma / chroma components) occurring in the regions adjacent to the block boundaries.

[0110] The SAO filter 1045 is configured to remove artifacts related to sample distortion caused by the encoding process. The SAO filter 1045 in the encoder classifies the deblocked samples of the reconstructed image into several categories based on the relative deblocking edge shape and / or direction. Then, an offset is determined based on the category and added to the sample. The offset is then encoded into the bitstream and used by the SAO filter 1045 in the decoder. The SAO filter 1045 removes banding artifacts (bands of values instead of smooth transitions) and ringing artifacts (spurious signals near sharp edges).

[0111] The ALF 1047 is configured to compare the reconstructed image with the original image in the encoder. The ALF 1047 determines, for example, coefficients that describe the difference between the reconstructed image and the original image via a Wiener-based adaptive filter. Such coefficients are encoded into the bitstream and used by the ALF 1047 in the decoder to remove the difference between the reconstructed image and the original image.

[0112] The image data filtered by the in-loop filter 1000 is output to a picture buffer 1023 that is substantially similar to the decoded picture buffers 223, 323, and / or 423. As described above, the deblocking filter 1043, the SAO filter 1045, and / or the ALF 1047 can be turned off at sub-picture boundaries and / or tile boundaries by flags such as the loop_filter_across_subpic_enabled flag and / or the loop_filter_across_tiles_enabled_flag, respectively.

[0113] FIG. 11 is a schematic diagram showing an exemplary bitstream 1100 including coding tool parameters for supporting decoding of sub-pictures of a picture. For example, the bitstream 1100 can be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400. As another example, the bitstream 1100 may be generated by the encoder in step 109 of method 100 for use by the decoder in step 111. Further, the bitstream 1100 may include coded blocks such as the coded picture 500, the corresponding sub-picture 510, and / or the current blocks 801 and / or 901, which may be coded according to the one-way inter prediction 600 and / or the bidirectional inter prediction 700 based on a candidate list generated according to the pattern 900. The bitstream 1100 may also include parameters for configuring the in-loop filter 1000.

[0114] The bitstream 1100 includes a sequence parameter set (SPS) 1110, a plurality of picture parameter sets (PPS) 1111, a plurality of slice headers 1115, and picture data 1120. The SPS 1110 includes sequence data common to all pictures within the video sequence included in the bitstream 1100. Such data can include picture size, bit depth, coding tool parameters, bitrate limits, and the like. The PPS 1111 includes parameters applicable to an entire picture. Thus, each picture within the video sequence may refer to the PPS 1111. Although each picture refers to the PPS 1111, it should be noted that in some examples, a single PPS 1111 can include data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such a case, a single PPS 1111 may include data for such similar pictures. The PPS 1111 can indicate coding tools, quantization parameters, offsets, etc. available for slices within the corresponding picture. The slice header 1115 includes parameters specific to each slice within a picture. Thus, there may be one slice header 1115 per slice in the video sequence. The slice header 1115 may include slice type information, picture order count (POC), reference picture list, prediction weights, tile entry points, deblocking parameters, and the like. It should also be noted that the slice header 1115 may also be referred to as a tile group header in some situations.

[0115] The picture data 1120 includes video data encoded according to inter prediction and / or intra prediction, and corresponding transformed and quantized residual data. For example, a video sequence includes a plurality of pictures coded as picture data. A picture is a single frame of the video sequence and thus is generally displayed as a single unit when the video sequence is displayed. However, sub-pictures may be displayed to implement certain technologies such as virtual reality, picture-in-picture, etc. The pictures each reference a PPS 1111. The pictures are divided into sub-pictures, tiles and / or slices as described above. In some systems, a slice is called a tile group that includes tiles. The slice and / or tile group of tiles references a slice header 1115. The slice is further divided into CTUs and / or CTBs. The CTU / CTB is further divided into coding blocks based on a coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.

[0116] The parameter set within bitstream 1100 includes various data that can be used to implement the examples described herein. To support the implementation manner of the first example, the SPS 1110 of bitstream 1100 includes a flag 1131 of a subpicture treated as a picture for a specified subpicture. In some examples, the flag 1131 of the subpicture treated as a picture is denoted as subpic_treated_as_pic_flag[i], where i is the index of the subpicture related to the flag. For example, the flag 1131 of the subpicture treated as a picture may be set equal to 1 to specify that the i-th subpicture of each coded picture within the coded video sequence (within image data 1120) is treated as a picture in the decoding process excluding the loop filtering operation. The flag 1131 of the subpicture treated as a picture may be used when the current subpicture within the current picture is coded according to inter prediction. When the flag 1131 of the subpicture treated as a picture is set to indicate that the current subpicture is treated as a picture, the candidate list of candidate motion vectors for the current block can be determined by excluding from the candidate list the motion vectors at the same position that are included in the block at the same position and point outside the current subpicture. This ensures that when the current subpicture is separately extracted from other subpictures, motion vectors pointing outside the current subpicture are not selected and related errors are avoided.

[0117] In some examples, the candidate list of motion vectors for the current block is determined according to temporal luma motion vector prediction. For example, when the current block is a luma block of luma samples, the selected current motion vector for the current block is a temporal luma motion vector pointing to a reference luma sample within the reference block, and the current block is coded based on the reference luma sample, temporal luma motion vector prediction may be used. In such a case, the temporal luma motion vector prediction is performed as follows. xColBr = xCb + cbWidth; yColBr = yCb + cbHeight; rightBoundaryPos = subpic_treated_as_pic_flag[SubPicIdx]? SubPicRightBoundaryPos : pic_width_in_luma_samples - 1; And botBoundaryPos = subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos : pic_height_in_luma_samples - 1 Here, xColBr and yColBR specify the position of the block at the same location, xCb and yCb specify the top - left sample of the current block relative to the top - left sample of the current picture, cbWidth is the width of the current block, cbHeight is the height of the current block, SubPicRightBoundaryPos is the position of the right boundary of the sub - picture, SubPicBotBoundaryPos is the position of the bottom boundary of the sub - picture, pic_width_in_luma_samples is the width of the current picture measured in luma samples, pic_height_in_luma_samples is the height of the current picture measured in luma samples, botBoundaryPos is the calculated position of the bottom boundary of the sub - picture, rightBoundaryPos is the calculated position of the right boundary of the sub - picture, SubPicIdx is the index of the sub - picture. When yCb >> CtbLog2SizeY is not equal to yColBr >> CtbLog2SizeY, the motion vector at the same location is excluded, and CtbLog2SizeY indicates the size of the coding tree block.

[0118] The flag 1131 of the sub-picture treated as a picture may also be used for the implementation of the second example. As in the first example, the flag 1131 of the sub-picture treated as a picture may be used when the current sub-picture within the current picture is coded according to inter-prediction. In this example, a motion vector can be determined for the current block of the sub-picture (e.g., from the candidate list). When the flag 1131 of the sub-picture treated as a picture is set, the clipping function can be applied to the sample position within the reference block. The sample position is a position within the picture that can include a single sample including a luma value and / or a pair of chroma values. Then, when the motion vector points outside the current sub-picture, an interpolation filter can be applied. This clipping function ensures that the interpolation filter does not depend on data from adjacent sub-pictures in order to maintain the separation between sub-pictures to support separate extraction.

[0119] The clipping function can be applied in the luma sample bilinear interpolation process. The luma sample bilinear interpolation process may receive an input including the luma position at all sample units (xIntL, yIntL). The luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). The clipping function is applied to the sample position as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, and Clip3 is a clipping function that follows the following. [Number] Here, x, y, and z are numerical input values.

[0120] The clipping function can also be applied in the luma sample 8-tap interpolation filtering process. The luma sample 8-tap interpolation filtering process receives an input including the luma position at all sample units (xIntL, yIntL). The luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). The clipping function is applied to the sample positions as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i - 3), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i - 3) Here, subpic_treated_as_pic_flag is a flag set to indicate that a subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, and Clip3 is as described above.

[0121] The clipping function can also be applied in the chroma sample interpolation process. The chroma sample interpolation process receives an input including the chroma positions at all sample units (xIntC, yIntC). The chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC). The clipping function is applied to the sample positions as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xIntC + i), and yInti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yIntC + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that a subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, SubWidthC and SubHeightC indicate the horizontal and vertical sampling rate ratios between luma samples and chroma samples, and Clip3 is as described above.

[0122] The loop filter activation flag 1132 across sub - pictures within SPS1110 may be used for the implementation method of the third example. The loop filter activation flag 1132 across sub - pictures may be set to control whether filtering is used across the boundaries of a specified sub - picture. For example, the loop filter activation flag 1132 across sub - pictures may be indicated as loop_filter_across_subpic_enabled_flag. The loop filter activation flag 1132 across sub - pictures may be set to 1 when specifying that the in - loop filtering operation can be performed across the boundary of the sub - picture, or may be set to 0 when specifying that the in - loop filtering operation is not performed across the boundary of the sub - picture. Therefore, the filtering operation may or may not be performed across the sub - picture boundary based on the value of the loop filter activation flag 1132 across sub - pictures. The filtering operation may include applying the de - blocking filter 1043, ALF1047, and / or SAO filter 1045. Thus, the filter can be turned off for sub - pictures encoded for separate extraction, or can be turned on for sub - pictures encoded for display as a group.

[0123] The loop filter activation flag 1134 across tiles within PPS1111 may be used for the implementation method of the fourth example. The loop filter activation flag 1134 across tiles may be set to control whether filtering is used across the boundaries of the specified tiles. For example, the loop filter activation flag 1134 across tiles may be denoted as loop_filter_across_tiles_enabled_flag. The loop filter activation flag 1134 across tiles may be set to 1 when specifying that the in-loop filtering operation can be executed across the boundaries of the tiles, or may be set to 0 when specifying that the in-loop filtering operation is not executed across the boundaries of the tiles. Therefore, the filtering operation may or may not be executed across the specified tile boundaries based on the value of the loop filter activation flag 1134 across tiles. The filtering operation may include applying the deblocking filter 1043, ALF 1047, and / or SAO filter 1045.

[0124] The sub-picture data 1133 in the SPS 1110 may be used for the implementation method of the fifth example. The sub-picture data 1133 may include the width, height, and offset for each sub-picture in the image data 1120. For example, the width and height of each sub-picture can be described in the sub-picture data 1133 in units of CTB. In some examples, the width and height of the sub-picture are stored in the sub-picture data 1133 as subpic_width_minus1 and subpic_height_minus1, respectively. Further, the offset of each sub-picture can be described in the sub-picture data 1133 in units of CTU. For example, the offset of each sub-picture can be specified as the vertical position and horizontal position of the top-left CTU of the sub-picture. Specifically, the offset of the sub-picture can be specified as the difference between the top-left CTU of the picture and the top-left CTU of the sub-picture. In some examples, the vertical position and horizontal position of the top-left CTU of the sub-picture are stored in the sub-picture data 1133 as subpic_ctu_top_left_y and subpic_ctu_top_left_x, respectively. The implementation method of this example describes the sub-pictures in the sub-picture data 1133 with respect to CTB / CTU, rather than with respect to tiles. This enables sub-pictures to be used even when tiles are not used in the corresponding picture / sub-picture.

[0125] The sub-picture data 1133 within SPS 1110, the slice address 1136 within the slice header 1115, and the slice sub-picture ID 1135 within the slice header 1115 may be used for the implementation of the sixth example. The sub-picture data 1133 may be implemented as described in the implementation of the fifth example. The slice address 1136 may include a sub-picture level slice index of a slice related to the slice header 1115 (e.g., the image data 1120). For example, the slice may be indexed based on the position of the slice within the sub-picture rather than based on the position of the slice within the picture. The slice address 1136 may be stored in the slice_address variable. The slice sub-picture ID 1135 includes the ID of the sub-picture including the slice related to the slice header 1115. Specifically, the slice sub-picture ID 1135 may refer to the description (e.g., width, height, and offset) of the corresponding sub-picture within the sub-picture data 1133. The slice sub-picture ID 1135 may be stored in the slice_subpic_id variable. Therefore, the slice address 1136 is signaled as an index based on the position of the slice within the sub-picture indicated by the sub-picture ID 1135, as described in the sub-picture data 1133. In this way, even when sub-pictures are separately extracted and other sub-pictures are omitted from the bitstream 1100, the position of the slice within the sub-picture can be determined. This is because this addressing method separates the addresses of each sub-picture from other sub-pictures. Therefore, the slice header 1115 does not need to be rewritten when the sub-picture is extracted, as required in an addressing method where slices are addressed based on the position of the slices within the picture. It should be noted that this technique may be used when the slice is a rectangular / square slice (as opposed to a raster scan slice). For example, the rect_slice_flag within PPS 1111 can be set equal to 1 to indicate that the slice is a rectangular slice.

[0126] Exemplary implementation methods of sub-pictures used in some video coding systems are as follows. Information related to sub-pictures that may exist in the CVS may be signaled in the SPS. Such signaling may include the following information. The number of sub-pictures present in each picture of the CVS may be included in the SPS. In the context of the SPS or CVS, sub-pictures at the same position for all access units (AUs) may be collectively referred to as a sub-picture sequence. A loop for further specifying information related to the characteristics of each sub-picture may also be included in the SPS. Such information may include sub-picture identification, the position of the sub-picture (e.g., the offset distance between the top-left luma sample of the sub-picture and the top-left luma sample of the picture), and the size of the sub-picture. Further, the SPS may also be used to signal whether each sub-picture is a motion-constrained sub-picture, and a motion-constrained sub-picture is a sub-picture that includes MCTS. Profile, tier, and level information for each sub-picture may be included in the bitstream unless such information can be derived separately. Such information may be used for the profile, tier, and level information of an extraction bitstream created from extracting sub-pictures from the original bitstream including the entire picture. The profile and tier of each sub-picture may be derived to be the same as the original profile and tier. The level for each sub-picture may be signaled explicitly. Such signaling may exist in the above loop. The sequence-level hypothetical reference decoder (HRD) parameters may be signaled in the video usability information (VUI) portion of the SPS for each sub-picture (or equivalently each sub-picture sequence).

[0127] When a picture is not partitioned into two or more sub - pictures, the characteristics of the sub - pictures (e.g., position, size, etc.) may not be signaled in the bitstream, except for the sub - picture ID. When sub - pictures within a picture in a CVS are extracted, each access unit in the new bitstream may not contain sub - pictures. This is because the resulting picture data within each AU in the new bitstream is not partitioned into multiple sub - pictures. Therefore, sub - picture characteristics such as position and size may be omitted from the SPS since such information can be derived from the picture characteristics. However, sub - picture identification is still signaled because this ID may be referenced by the video coding layer (VCL) NAL unit / tile group included in the extracted sub - pictures. To reduce resource usage, when extracting sub - pictures, changing the sub - picture ID should be avoided.

[0128] The position of a sub-picture within a picture (x-offset and y-offset) can be signaled in units of luma samples and may represent the distance between the top-left luma sample of the sub-picture and the top-left luma sample of the picture. In other examples, the position of a sub-picture within a picture can be signaled in units of the minimum coding luma block size (MinCbSizeY) and may represent the distance between the top-left luma sample of the sub-picture and the top-left luma sample of the picture. In other examples, the unit of the sub-picture position offset may be explicitly indicated by a syntax element within a parameter set and the unit may be CtbSizeY, MinCbSizeY, luma samples or other values. The codec may require that when the right boundary of the sub-picture does not coincide with the right boundary of the picture, the width of the sub-picture be an integer multiple of the luma CTU size (CtbSizeY). Similarly, the codec may further require that when the bottom boundary of the sub-picture does not coincide with the bottom boundary of the picture, the height of the sub-picture be an integer multiple of CtbSizeY. The codec may also require that when the width of the sub-picture is not an integer multiple of the luma CTU size, the sub-picture be located at the rightmost position of the picture. Similarly, the codec may require that when the height of the sub-picture is not an integer multiple of the luma CTU size, the sub-picture be located at the bottommost position of the picture. The width of the sub-picture is signaled in units of the luma CTU size and when the width of the sub-picture is not an integer multiple of the luma CTU size, the actual width in luma samples may be derived based on the offset position of the sub-picture, the width of the sub-picture in luma CTU size, and the width of the picture in luma samples. Similarly, the height of the sub-picture is signaled in units of the luma CTU size and when the height of the sub-picture is not an integer multiple of the luma CTU size, the actual height in luma samples can be derived based on the offset position of the sub-picture, the height of the sub-picture in luma CTU size, and the height of the picture in luma samples.

[0129] For any sub - picture, the sub - picture ID may be different from the sub - picture index. The sub - picture index may be the index of the sub - picture as signaled in the loop of sub - pictures within the SPS. Alternatively, the sub - picture index may be the index assigned in the raster scan order of sub - pictures for the picture. When the value of the sub - picture ID of each sub - picture is the same as its sub - picture index, the sub - picture ID may be signaled or derived. When the sub - picture ID of each sub - picture is different from its sub - picture index, the sub - picture ID is explicitly signaled. The number of bits for signaling the sub - picture ID may be signaled in the same parameter set (e.g., in the SPS) that includes the sub - picture characteristics. Some values for the sub - picture ID may be reserved for specific purposes. Such reservation of values can be as follows. When the tile group / slice header includes a sub - picture ID to specify which sub - pictures are included in the tile group, to ensure that the first few bits of the tile group / slice header are not all 0 and to avoid generating emulation prevention codes, the value 0 may be reserved and may not be used for sub - pictures. When the sub - pictures of a picture do not cover the entire area of the picture without overlap and gaps, a value (e.g., value 1) may be reserved for tile groups that are not part of any sub - picture. Alternatively, the sub - picture IDs of the remaining areas may be explicitly signaled. The number of bits for signaling the sub - picture ID may be restricted as follows. The range of values should include the reserved values of the sub - picture ID and be sufficient to uniquely identify all sub - pictures within the picture. For example, the minimum number of bits for the sub - picture ID can be the value of Ceil(Log2(number of sub - pictures in the picture+number of reserved sub - picture IDs)).

[0130] The grouping of sub - pictures within the loop may be required to cover the entire picture without gaps and without overlap. When this constraint applies, there is a flag for each sub - picture to specify whether the sub - picture is a motion - constrained sub - picture, which means the sub - picture can be extracted. Alternatively, the grouping of sub - pictures may not need to cover the entire picture. However, there may be no overlap between the sub - pictures of the picture.

[0131] The sub-picture ID may be present immediately after the NAL unit header so that the extractor does not need to understand the rest of the NAL unit bits to assist the sub-picture extraction process. For VCL NAL units, the sub-picture ID may be present in the first bit of the tile group header. For non-VCL NAL units, the following conditions may apply. The sub-picture ID may not need to be present immediately after the NAL unit header for SPS. For PPS, when all tile groups of the same picture are constrained to refer to the same PPS, the sub-picture ID does not need to be present immediately after the NAL unit header. On the other hand, when tile groups of the same picture are allowed to refer to different PPSs, the sub-picture ID may be present in the first bit of the PPS (e.g., immediately after the PPS NAL unit header). In this case, it is not allowed for two different tile groups of one picture to share the same PPS. Alternatively, when tile groups of the same picture are allowed to refer to different PPSs and different tile groups of the same picture are also allowed to share the same PPS, the sub-picture ID does not exist in the PPS syntax. Alternatively, when tile groups of the same picture are allowed to refer to different PPSs and different tile groups of the same picture are also allowed to share the same PPS, a list of sub-picture IDs exists in the PPS syntax. The list indicates the sub-pictures to which the PPS applies. For other non-VCL NAL units, when the non-VCL unit applies at the picture level (e.g., access unit delimiter, end of sequence, end of bitstream, etc.) or above, the sub-picture ID does not need to be present immediately after the NAL unit header. Otherwise, the sub-picture ID may be present immediately after the NAL unit header.

[0132] The tile partitioning within each sub-picture may be signaled in the PPS, but tile groups within the same picture are allowed to refer to different PPSs. In this case, tiles are grouped within each sub-picture rather than across the picture. Thus, the concept of tile grouping in such cases includes partitioning the sub-picture into tiles. Alternatively, a Sub-Picture Parameter Set (SPPS) may be used to describe the tile partitioning within each individual sub-picture. The SPPS refers to the SPS by using a syntax element that references the SPS ID. The SPPS may include a sub-picture ID. For the purpose of sub-picture extraction, the syntax element that references the sub-picture ID is the first syntax element within the SPPS. The SPPS includes a tile structure indicating the number of columns, the number of rows, a uniform tile spacing, etc. The SPPS may also include a flag to indicate whether the loop filter is valid across the sub-picture boundaries it is associated with. Alternatively, the sub-picture characteristics for each sub-picture may be signaled with the SPPS instead of the SPS. The tile partitioning within each individual sub-picture may be signaled in the PPS, but tile groups within the same picture are allowed to refer to different PPSs. Once activated, the SPPS may continue over a sequence of consecutive AUs in decode order, but may be deactivated / activated at AUs that are not the start of a CVS. Multiple SPPSs may be active at any point during the decoding process of a single-layer bitstream having multiple sub-pictures, and the SPPS may be shared by different sub-pictures of an AU. Alternatively, the SPPS and the PPS may be merged into one parameter set. For this to occur, all tile groups included in the same sub-picture may be constrained to refer to the same parameter set resulting from the merge between the SPPS and the PPS.

[0133] The number of bits used to signal the sub-picture ID may be signaled in the NAL unit header. Such information, when present, assists the sub-picture extraction process in parsing the sub-picture ID value for the start of the payload of the NAL unit (e.g., the first few bits immediately following the NAL unit header). For such signaling, some of the spare bits in the NAL unit header may be used to avoid increasing the length of the NAL unit header. The number of bits for such signaling should cover the value of sub-picture-ID-bit-len. For example, 4 out of the 7 spare bits in the NAL unit header of VVC may be used for this purpose.

[0134] When decoding a subpicture, the positions of each coding tree block, indicated as the vertical CTB position (xCtb) and the horizontal CTB position (yCtb), are adjusted to the actual luma sample positions in the picture rather than the luma sample positions within the subpicture. In this way, since everything is decoded as if it were located in the picture rather than the subpicture, extraction of subpictures at the same position from each reference picture can be avoided. To adjust the positions of the coding tree blocks, the variables SubpictureXOffset and SubpictureYOffset are derived based on the subpicture positions (subpic_x_offset and subpic_y_offset). The values of the variables are added to the values of the x and y coordinates of the luma sample positions of each coding tree block within the subpicture, respectively. The subpicture extraction process can be defined as follows. The input to the process includes the target subpicture to be extracted. This can be input in the form of a subpicture ID or a subpicture position. When the input is the position of the subpicture, the associated subpicture ID can be resolved by analyzing the subpicture information in the SPS. For non-VCL NAL units, the following applies. The syntax elements in the SPS related to the picture size and level are updated with the subpicture size and level information. The following non-VCL NAL units, namely, PPS, access unit delimiter (AUD), end of sequence (EOS), end of bitstream (EOB), and any other non-VCL NAL unit applicable at the picture level and above, are not changed by extraction. The remaining non-VCL NAL units with a subpicture ID not equal to the target subpicture ID are removed. VCL NAL units with a subpicture ID not equal to the target subpicture ID are also removed.

[0135] The sub-picture nested SEI message may be used for nesting of AU-level or sub-picture-level SEI messages about a set of sub-pictures. The data carried in the sub-picture nested SEI message may include a buffering period, picture timing, and non-HRD SEI messages. The syntax and semantics of this SEI message can be as follows. For system operations such as in an omnidirectional media format (OMAF) environment, a set of sub-picture sequences covering a viewport may be requested and decoded by an OMAF player. Therefore, the sequence-level SEI message may carry information about a set of sub-picture sequences that together include rectangular or square picture regions. This information can be used by the system, and this information indicates the minimum decoding capability and the bitrate of the set of sub-picture sequences. This information includes the level of the bitstream that contains only the set of sub-picture sequences, the bitrate of the bitstream, and optionally a sub-bitstream extraction process specified for the set of sub-picture sequences.

[0136] The above implementation method contains several problems. The signaling of the width and height of the picture, and / or the width / height / offset of the sub-picture is not efficient. More bits can be saved to signal such information. When the sub-picture size and position information are signaled in the SPS, the PPS contains the tile configuration. Furthermore, the PPS is allowed to be shared by multiple sub-pictures of the same picture. Therefore, the value ranges for num_tile_columns_minus1 and num_tile_rows_minus1 should be specified more clearly. Furthermore, the semantics of the flag indicating whether the sub-picture is motion-constrained are not clearly specified. The level is compulsorily signaled for each sub-picture sequence. However, when the sub-picture sequence cannot be decoded independently, signaling the level of the sub-picture is not useful. Furthermore, in some applications, some sub-picture sequences should be decoded and rendered together with at least one other sub-picture sequence. Therefore, signaling the level for a single such sub-picture sequence may not be useful. Furthermore, determining the level value for each sub-picture may impose a burden on the encoder.

[0137] By introducing independently decodable sub-picture sequences, scenarios that require independent extraction and decoding of specific regions of a picture may not function based on tile groups. Therefore, explicit signaling of tile group IDs may not be useful. Furthermore, the respective values of the PPS syntax elements pps_seq_parameter_set_id and loop_filter_across_tiles_enabled_flag should be the same for all PPSs referenced by the tile group headers of the coded picture. This is because the active SPS should not change within the CVS, and the value of the loop_filter_across_tiles_enabled_flag should be the same for all tiles within the picture for parallel processing based on tiles. Whether to allow a mixture of rectangular and raster scan tile groups within a picture should be clearly specified. Whether to allow sub-pictures that are part of different pictures and use the same sub-picture ID within the CVS for using different tile group modes should also be specified. The derivation process for temporal luma motion vector prediction may not allow treating sub-picture boundaries as picture boundaries in temporal motion vector prediction (TMVP). Furthermore, the luma sample bilinear interpolation process, the luma sample 8-tap interpolation filtering process, and the chroma sample interpolation process may not be configured to treat sub-picture boundaries as picture boundaries in motion compensation. Also, a mechanism for controlling the deblocking, SAO, and ALF filtering operations at sub-picture boundaries should be specified.

[0138] The introduction of independently decodable sub-picture sequences may render the loop_filter_across_tile_groups_enabled_flag less useful. This is because turning off the in-loop filtering operation for parallel processing purposes may also be achieved by setting the loop_filter_across_tile_groups_enabled_flag equal to 0. Furthermore, turning off the in-loop filtering operation to enable independent extraction and decoding of a specific region of a picture may also be achieved by setting the loop_filter_across_sub_pic_enabled_flag equal to 0. Therefore, further specifying a process to turn off the in-loop filtering operation across tile group boundaries based on the loop_filter_across_tile_groups_enabled_flag unnecessarily burdens the decoder and wastes bits. Additionally, the above decoding process may not allow turning off the ALF filtering operation across tile boundaries.

[0139] Accordingly, the present disclosure includes a design for supporting sub-picture based video coding. A sub-picture is a rectangular or square region within a picture that may or may not be independently decoded using the same decoding process as the picture. The technical description is based on the Versatile Video Coding (VVC) standard. However, the technology may also be applied to other video codec specifications.

[0140] In some examples, for the syntax elements of picture width and height and the list of syntax elements of sub-picture width / height / offset_x / offset_y, the size unit is signaled. All syntax elements are signaled in the form of xxx_minus1. For example, when the size unit is 64 luma samples, a width value of 99 specifies a picture width of 6400 luma samples. The same example applies to other of these syntax elements. In other examples, one or more of the following apply. The size unit may be signaled in the form of xxx_minus1 for the syntax elements of picture width and height. Such a size unit signaled for the list of syntax elements of sub-picture width / height / offset_x / offset_y may also be in the form of xxx_minus1. In other examples, one or more of the following apply. The size unit for the list of syntax elements of picture width and sub-picture width / offset_x may be signaled in the form of xxx_minus1. The size unit for the list of syntax elements of picture height and sub-picture height / offset_y may be signaled in the form of xxx_minus1. In other examples, one or more of the following apply. The syntax elements of picture width and height in the form of xxx_minus1 may be signaled in units of the minimum coding unit. The syntax elements of sub-picture width / height / offset_x / offset_y in the form of xxx_minus1 may be signaled in units of CTU or CTB. The sub-picture width for each sub-picture at the right picture boundary may be derived. The sub-picture height for each sub-picture at the bottom picture boundary may be derived. All other values of sub-picture width / height / offset_x / offset_y may also be signaled in the bitstream. In other examples, a mode for signaling the width and height of the sub-picture and their positions within the picture may be added for the case where the sub-pictures have a uniform size.Sub - pictures have a uniform size when they include the same sub - picture rows and sub - picture columns. In this mode, the number of sub - picture rows, the number of sub - picture columns, the width of each sub - picture column, and the height of each sub - picture row may all be signaled.

[0141] In other examples, the signaling of sub - picture width and height may not be included in the PPS. num_tile_columns_minus1 and num_tile_rows_minus1 should be in a range of one integer value or less, such as 0 to 1024. In other examples, when a sub - picture referring to the PPS has more than one tile, two syntax elements conditioned by a presence flag may be signaled in the PPS. These syntax elements are used to signal the sub - picture width and height in units of CTB and specify the sizes of all sub - pictures referring to the PPS.

[0142] In other examples, further information describing individual sub-pictures may also be signaled. A flag such as sub_pic_treated_as_pic_flag[i] may be signaled for each sub-picture sequence to indicate whether the sub-picture of the sub-picture sequence is treated as a picture in the decoding process for purposes other than the in-loop filtering operation. The level to which the sub-picture sequence conforms may be signaled only when sub_pic_treated_as_pic_flag[i] is equal to 1. A sub-picture sequence is a CVS of sub-pictures having the same sub-picture ID. When sub_pic_treated_as_pic_flag[i] is equal to 1, the level of the sub-picture sequence may also be signaled. This can be controlled by a flag for all sub-picture sequences or by one flag per sub-picture sequence. In other examples, sub-bitstream extraction may be enabled without modifying the VCL NAL unit. This can be achieved by removing the signaling of the explicit tile group ID from the PPS. The semantics of tile_group_address are specified when rect_tile_group_flag indicates a rectangular tile group. Tile_group_address may include the tile group index of the tile group within the tile group in the sub-picture.

[0143] In another example, it is assumed that the respective values of the PPS syntax elements pps_seq_parameter_set_id and loop_filter_across_tiles_enabled_flag are the same in all PPSs referred to by the tile group header of the coded picture. Other PPS syntax elements may be different for different PPSs referred to by the tile group header of the coded picture. The value of single_tile_in_pic_flag may be different for different PPSs referred to by the tile group header of the coded picture. Thus, some pictures in the CVS may have only one tile, while some other pictures in the CVS may have multiple tiles. This also allows some sub-pictures of a picture (e.g., very large ones) to have multiple tiles, while other sub-pictures of the same picture (e.g., very small ones) have only one tile.

[0144] In another example, a picture may include a mixture of rectangular and raster scan tile groups. Thus, some sub-pictures of a picture use the rectangular tile group mode, while other sub-pictures use the raster scan tile group mode. This flexibility is beneficial for bitstream merge scenarios. Alternatively, the constraint may require that all sub-pictures of a picture use the same tile group mode. Sub-pictures from different pictures with the same sub-picture ID in the CVS may not use different tile group modes. Sub-pictures from different pictures with the same sub-picture ID in the CVS may use different tile group modes.

[0145] In another example, when sub_pic_treated_as_pic_flag[i] for a sub-picture is equal to 1, the motion vectors at the same position for temporal motion vector prediction for the sub-picture are restricted to originate within the boundaries of the sub-picture. Thus, the temporal motion vector prediction for the sub-picture is treated as if the sub-picture boundary were the picture boundary. Further, to enable treating the sub-picture boundary as a picture boundary in motion compensation for sub-pictures where sub_pic_treated_as_pic_flag[i] is equal to 1, the clipping operation is specified as part of the luma sample bilinear interpolation process, the luma sample 8-tap interpolation filtering process, and the chroma sample interpolation process.

[0146] In other examples, each sub-picture is associated with a signaled flag such as loop_filter_across_sub_pic_enabled_flag. The flag is used to control the loop filtering operation at the sub-picture boundary and the filtering operation in the corresponding decoding process. The deblocking filter process may not be applied to coding sub-block edges and transform block edges that coincide with the sub-picture boundary where loop_filter_across_sub_pic_enabled_flag is equal to 0. Alternatively, the deblocking filter process may not be applied to coding sub-block edges and transform block edges that coincide with the upper or left boundary of the sub-picture where loop_filter_across_sub_pic_enabled_flag is equal to 0. Alternatively, the deblocking filter process may not be applied to coding sub-block edges and transform block edges that coincide with the sub-picture boundary where sub_pic_treated_as_pic_flag[i] is equal to 1 or 0. Alternatively, the deblocking filter process may not be applied to coding sub-block edges and transform block edges that coincide with the upper or left boundary of the sub-picture. The clipping operation may be specified to turn off the SAO filtering operation across the sub-picture boundary when loop_filter_across_sub_pic_enabled_flag for the sub-picture is equal to 0. The clipping operation may be specified to turn off the ALF filtering operation across the sub-picture boundary when loop_filter_across_sub_pic_enabled_flag for the sub-picture is equal to 0. loop_filter_across_tile_groups_enabled_flag may also be removed from the PPS. Thus, when loop_filter_across_tiles_enabled_flag is equal to 0, the loop filtering operation across the tile group boundary that is not the sub-picture boundary is not turned off. The loop filter operation may include deblocking, SAO, and ALF.In another example, the clipping operation is specified to turn off the ALF filtering operation across tile boundaries when the loop_filter_across_tiles_enabled_flag for the tile is equal to 0.

[0147] One or more of the above examples may be implemented as follows. A sub-picture may be defined as a rectangular or square region of one or more tile groups or slices within a picture. The following partitioning of processing elements, namely, the partitioning into components of each picture, the partitioning of each component into CTBs, the partitioning of each picture into sub-pictures, the partitioning of each sub-picture into tile columns within the sub-picture, the partitioning of each sub-picture into tile rows within the sub-picture, the partitioning of each tile column within the sub-picture into tiles, the partitioning of each tile row within the sub-picture into tiles, and the partitioning of each sub-picture into tile groups may form a spatial or component-wise partition.

[0148] The process for the CTB raster and tile scan process within a sub-picture may be as follows. A list ColWidth[i] for i in the range of 0 to num_tile_columns_minus1, specifying the width of the i-th tile column in units of CTBs, may be derived as follows.

Number

[0149] A list RowHeight[j] for j in the range of 0 to num_tile_rows_minus1, specifying the height of the j-th tile row in units of CTBs, is derived as follows.

Number

[0150] A list ColBd[i] for i in the range of 0 to num_tile_columns_minus1 + 1, specifying the position of the i-th tile column boundary in units of CTBs, is derived as follows. [Number]

[0151] For j in the range of 0 to num_tile_rows_minus1 + 1 that specifies the position of the j-th tile row boundary in units of CTB, the list RowBd[j] is derived as follows. [Number]

[0152] For ctbAddrRs in the range of 0 to SubPicSizeInCtbsY - 1 that specifies the conversion from the CTB address in the CTB raster scan of the sub-picture to the CTB address in the tile scan of the sub-picture, the list CtbAddrRsToTs[ctbAddrRs] is derived as follows. [Number]

[0153] For ctbAddrTs in the range of 0 to SubPicSizeInCtbsY - 1 that specifies the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the sub-picture, the list CtbAddrTsToRs[ctbAddrTs] is derived as follows. [Number]

[0154] For ctbAddrTs in the range of 0 to SubPicSizeInCtbsY - 1 that specifies the conversion from the CTB address in the tile scan of the sub-picture to the tile ID, the list TileId[ctbAddrTs] is derived as follows. [Number]

[0155] The list NumCtusInTile[tileIdx] for tileIdx in the range from 0 to NumTilesInSubPic - 1, which specifies the conversion from the tile index to the number of CTUs within the tile, is derived as follows.

Number

[0156] The list FirstCtbAddrTs[tileIdx] for tileIdx in the range from 0 to NumTilesInSubPic - 1, which specifies the conversion from the tile ID to the CTB address in the tile scan of the first CTB within the tile, is derived as follows.

Number

[0157] The value of ColumnWidthInLumaSamples[i] that specifies the width of the i-th tile column in units of luma samples is set equal to ColWidth[i] << CtbLog2SizeY for i in the range from 0 to num_tile_columns_minus1. The value of RowHeightInLumaSamples[j] that specifies the height of the j-th tile row in units of luma samples is set equal to RowHeight[j] << CtbLog2SizeY for j in the range from 0 to num_tile_rows_minus1.

[0158] The syntax of an exemplary sequence parameter set RBSP is as follows.

Table 1

[0159] The syntax of an exemplary picture parameter set RBSP is as follows.

Table 2

[0160] The syntax of an exemplary general tile group header is as follows. [Table 3]

[0161] The syntax of an exemplary coding tree unit is as follows. [Table 4]

[0162] The semantics of an exemplary sequence parameter set RBSP are as follows.

[0163] bit_depth_chroma_minus8 specifies the bit depth of samples of the chroma array BitDepthC, and the value of the chroma quantization parameter range offset QpBdOffsetC is as follows. [Number] bit_depth_chroma_minus8 shall be in the range from 0 to 8 inclusive.

[0164] One plus num_sub_pics_minus1 specifies the number of sub-pictures in each coded picture in the CVS. The value of num_sub_pics_minus1 shall be in the range of 0 to 1024 inclusive. One plus sub_pic_id_len_minus1 specifies the number of bits used to represent the syntax element sub_pic_id[i] in the SPS and the syntax element tile_group_sub_pic_id in the tile group header. The value of sub_pic_id_len_minus1 shall be in the range of Ceil(Log2(num_sub_pic_minus1 + 1)) - 1 to 9 inclusive. sub_pic_level_present_flag is set to 1 to specify that the syntax element sub_pic_level_idc[i] may be present. sub_pic_level_present_flag is set to 0 to specify that the syntax element sub_pic_level_idc[i] is not present. sub_pic_id[i] specifies the sub-picture ID of the i-th sub-picture of each coded picture in the CVS. The length of sub_pic_id[i] is sub_pic_id_len_minus1 + 1 bits.

[0165] sub_pic_treated_as_pic_flag[i] is set equal to 1 to specify that the i-th sub-picture of each coded picture in the CVS is to be treated as a picture in the decoding process, excluding the in-loop filtering operation. sub_pic_treated_as_pic_flag[i] is set equal to 0 to specify that the i-th sub-picture of each coded picture in the CVS is not to be treated as a picture in the decoding process, excluding the in-loop filtering operation. sub_pic_level_idc[i] indicates the level to which the i-th sub-picture sequence conforms, where the i-th sub-picture sequence consists only of the VCL NAL units of the sub-pictures having a sub-picture ID equal to sub_pic_id[i] in the CVS and the non-VCL NAL units associated therewith. sub_pic_x_offset[i] specifies the horizontal offset, in units of luma samples, of the top-left luma sample of the i-th sub-picture with respect to the top-left luma sample of each picture in the CVS. When it does not exist, the value of sub_pic_x_offset[i] is presumed to be equal to 0. sub_pic_y_offset[i] specifies the vertical offset, in units of luma samples, of the top-left luma sample of the i-th sub-picture with respect to the top-left luma sample of each picture in the CVS. When it does not exist, the value of sub_pic_y_offset[i] is presumed to be equal to 0. sub_pic_width_in_luma_samples[i] specifies the width of the i-th sub-picture of each picture in the CVS, in units of luma samples. When the sum of sub_pic_x_offset[i] and sub_pic_width_in_luma_samples[i] is less than pic_width_in_luma_samples, the value of sub_pic_width_in_luma_samples[i] shall be an integer multiple of CtbSizeY. When it does not exist, the value of sub_pic_width_in_luma_samples[i] is presumed to be equal to pic_width_in_luma_samples.sub_pic_height_in_luma_samples[i] specifies the height of the i-th sub-picture in each picture in the CVS in units of luma samples. When the sum of sub_pic_y_offset[i] and sub_pic_height_in_luma_samples[i] is less than pic_height_in_luma_samples, the value of sub_pic_height_in_luma_samples[i] shall be an integer multiple of CtbSizeY. When it does not exist, the value of sub_pic_height_in_luma_samples[i] is presumed to be equal to pic_height_in_luma_samples.

[0166] Regarding bitstream compliance, the following constraints apply. For any integer values of i and j, when i is equal to j, the values of sub_pic_id[i] and sub_pic_id[j] shall not be the same. For any two sub-pictures subpicA and subpicB, when the sub-picture ID of subpicA is smaller than the sub-picture ID of subpicB, any coded tile group NAL unit of subPicA shall follow any coded tile group NAL unit of subPicB in decoding order. The shape of the sub-picture shall be such that when decoded, each sub-picture has its overall left boundary and overall upper boundary composed of the picture boundary or the boundary of a previously decoded sub-picture.

[0167] The list SubPicIdx[spId] equal to sub_pic_id[i] for i in the range from 0 to num_sub_pics_minus1, which specifies the conversion from the sub-picture ID to the sub-picture index, is derived as follows.

Number

[0168] log2_max_pic_order_cnt_lsb_minus4 specifies the value of the variable MaxPicOrderCntLsb, which is used in the decoding process for the picture order count as follows. [Number] The value of log2_max_pic_order_cnt_lsb_minus4 shall be in the range from 0 to 12, inclusive.

[0169] The semantics of an exemplary picture parameter set RBSP are as follows.

[0170] When present, the values of the PPS syntax elements pps_seq_parameter_set_id and loop_filter_across_tiles_enabled_flag shall be the same for all PPSs referenced by the tile group header of the coded picture. pps_pic_parameter_set_id identifies the PPS referenced by other syntax elements. The value of pps_pic_parameter_set_id shall be in the range from 0 to 63, inclusive. pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id for the active SPS. The value of pps_seq_parameter_set_id shall be in the range from 0 to 15, inclusive. loop_filter_across_sub_pic_enabled_flag is set equal to 1 to specify that the loop filter operation may be performed across the boundaries of sub-pictures that reference the PPS. loop_filter_across_sub_pic_enabled_flag is set equal to 0 to specify that the loop filter operation is not performed across the boundaries of sub-pictures that reference the PPS.

[0171] The single_tile_in_sub_pic_flag is set equal to 1 to specify that only one tile exists in each sub-picture that refers to the PPS. The single_tile_in_sub_pic_flag is set equal to 0 to specify that more than one tile exists in each sub-picture that refers to the PPS. Adding 1 to num_tile_columns_minus1 specifies the number of tile columns that divide the sub-picture. It is assumed that num_tile_columns_minus1 is in the range of 0 to 1024. When it does not exist, the value of num_tile_columns_minus1 is assumed to be equal to 0. Adding 1 to num_tile_rows_minus1 specifies the number of tile rows that divide the sub-picture. It is assumed that num_tile_rows_minus1 is in the range of 0 to 1024. When it does not exist, the value of num_tile_rows_minus1 is assumed to be equal to 0. The variable NumTilesInSubPic is set equal to (num_tile_columns_minus1 + 1)*(num_tile_rows_minus1 + 1). When single_tile_in_sub_pic_flag is equal to 0, NumTilesInSubPic is assumed to be greater than 1.

[0172] The uniform_tile_spacing_flag is set equal to 1 to specify that the tile column boundaries and likewise the tile row boundaries are uniformly distributed among the sub - pictures. The uniform_tile_spacing_flag is set equal to 0 to specify that the tile column boundaries and likewise the tile row boundaries are not uniformly distributed among the sub - pictures, but are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When not present, the value of the uniform_tile_spacing_flag is assumed to be equal to 1. One plus tile_column_width_minus1[i] specifies the width of the i - th tile column in units of CTB. One plus tile_row_height_minus1[i] specifies the height of the i - th tile row in units of CTB. The single_tile_per_tile_group is set equal to 1 to specify that each tile group that references this PPS contains one tile. The single_tile_per_tile_group is set equal to 0 to specify that the tile groups that reference this PPS may contain more than one tile.

[0173] The rect_tile_group_flag is set equal to 0 to specify that the tiles within each tile group of a subpicture are in raster scan order and the tile group information is not signaled in the PPS. The rect_tile_group_flag is set equal to 1 to specify that the tiles within each tile group cover a rectangular or square region of the subpicture and the tile group information is signaled in the PPS. When the single_tile_per_tile_group_flag is set to 1, the rect_tile_group_flag is assumed to be equal to 1. One plus num_tile_groups_in_sub_pic_minus1 specifies the number of tile groups within each subpicture that references the PPS. The value of num_tile_groups_in_sub_pic_minus1 shall be in the range from 0 to NumTilesInSubPic - 1. In the absence and when the single_tile_per_tile_group_flag is equal to 1, the value of num_tile_groups_in_sub_pic_minus1 is assumed to be equal to NumTilesInSubPic - 1.

[0174] top_left_tile_idx[i] specifies the tile index of the tile located at the upper left corner of the i-th tile group of the subpicture. The value of top_left_tile_idx[i] shall be unequal to the value of top_left_tile_idx[j] for any i not equal to j. When it does not exist, the value of top_left_tile_idx[i] is presumed to be equal to i. The length of the top_left_tile_idx[i] syntax element is Ceil(Log2(NumTilesInSubPic)) bits. bottom_right_tile_idx[i] specifies the tile index of the tile located at the lower right corner of the i-th tile group of the subpicture. When single_tile_per_tile_group_flag is set to 1, bottom_right_tile_idx[i] is presumed to be equal to top_left_tile_idx[i]. The length of the bottom_right_tile_idx[i] syntax element is Ceil(Log2(NumTilesInSubPic)) bits.

[0175] It is a requirement for bitstream compliance that any particular tile be included in only one tile group. The variable NumTilesInTileGroup[i] specifying the number of tiles within the i-th tile group of the subpicture and related variables are derived as follows. [Number]

[0176] The loop_filter_across_tiles_enabled_flag is set equal to 1 to specify that the loop filter operation may be performed across tile boundaries within a subpicture that references the PPS. The loop_filter_across_tiles_enabled_flag is set equal to 0 to specify that the loop filter operation is not performed across tile boundaries within a subpicture that references the PPS. The loop filter operation includes deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, the value of the loop_filter_across_tiles_enabled_flag is presumed to be equal to 1. One plus num_ref_idx_default_active_minus1[i] specifies the estimated value of the variable NumRefIdxActive[0] for a P or B tile group with num_ref_idx_active_override_flag equal to 0 when i is equal to 0, and specifies the estimated value of NumRefIdxActive[1] for a B tile group with num_ref_idx_active_override_flag equal to 0 when i is equal to 1. The value of num_ref_idx_default_active_minus1[i] shall be in the range of 0 to 14, inclusive.

[0177] The semantics of an exemplary general tile group header are as follows. When present, the values of the syntax elements tile_group_pic_order_cnt_lsb and tile_group_temporal_mvp_enabled_flag of the tile group header shall be the same for all tile group headers of the coded picture. When present, the value of tile_group_pic_parameter_set_id shall be the same for all tile group headers of the coded subpicture. tile_group_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id for the PPS in use. The value of tile_group_pic_parameter_set_id shall be in the range from 0 to 63 inclusive. It is a bitstream conformity requirement that the value of the TemporalId of the current picture be greater than or equal to the value of the TemporalId of each PPS referenced by the tile group of the current picture. tile_group_sub_pic_id identifies the subpicture to which the tile group belongs. The length of tile_group_sub_pic_id is sub_pic_id_len_minus1 + 1 bits. The value of tile_group_sub_pic_id shall be the same for all tile group headers of the coded subpicture.

[0178] The variables SubPicWidthInCtbsY, SubPicHeightInCtbsY and SubPicSizeInCtbsY are derived as follows.

Number

[0179] The following variables, namely, the list ColWidth[i] for i in the range from 0 to num_tile_columns_minus1 that specifies the width of the i-th tile column in units of CTB, the list RowHeight[j] for j in the range from 0 to num_tile_rows_minus1 that specifies the height of the j-th tile row in units of CTB, the list ColBd[i] for i in the range from 0 to num_tile_columns_minus1 + 1 that specifies the position of the i-th tile column boundary in units of CTB, the list RowBd[j] for j in the range from 0 to num_tile_rows_minus1 + 1 that specifies the position of the j-th tile row boundary in units of CTB, the list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs in the range from 0 to SubPicSizeInCtbsY - 1 that specifies the conversion from the CTB address in the CTB raster scan of the subpicture to the CTB address in the tile scan of the subpicture, the list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs in the range from 0 to SubPicSizeInCtbsY - 1 that specifies the conversion from the CTB address in the tile scan of the subpicture to the CTB address in the CTB raster scan of the subpicture, the list TileId[ctbAddrTs] for ctbAddrTs in the range from 0 to SubPicSizeInCtbsY - 1 that specifies the conversion from the CTB address in the tile scan of the subpicture to the tile ID, the list NumCtusInTile[tileIdx] for tileIdx in the range from 0 to NumTilesInSubPic - 1 that specifies the conversion from the tile index to the number of CTUs in the tile, the list FirstCtbAddrTs[tileIdx] for tileIdx in the range from 0 to NumTilesInSubPic - 1 that specifies the conversion from the tile ID to the CTB address in the tile scan of the first CTB in the tile, the list ColumnWidthInLumaSamples[i] for i in the range from 0 to num_tile_columns_minus1 that specifies the width of the i-th tile column in units of luma samples.And the list RowHeightInLumaSamples[j] for j in the range from 0 to num_tile_rows_minus1, which specifies the height of the j-th tile row in luma samples, is derived by calling the CTB raster and tile scan conversion processes.

[0180] The values of ColumnWidthInLumaSamples[i] for i in the range from 0 to num_tile_columns_minus1 and RowHeightInLumaSamples[j] for j in the range from 0 to num_tile_rows_minus1 shall all be greater than 0. The variables SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows.

Number

[0181] For each tile with index i = 0..NumTilesInSubPic - 1 within the current subpicture, the variables TileLeftBoundaryPos[i], TileTopBoundaryPos[i], TileRightBoundaryPos[i], and TileBotBoundaryPos[i] are derived as follows.

Number

[0182] The tile_group_address specifies the tile address of the first tile within the tile group. When it does not exist, the value of tile_group_address is assumed to be equal to 0. When rect_tile_group_flag is equal to 0, the following applies. The tile address is the tile ID, the length of tile_group_address is Ceil(Log2(NumTilesInSubPic)) bits, and the value of tile_group_address is in the range from 0 to NumTilesInSubPic - 1. Otherwise (when rect_tile_group_flag is equal to 1), the following applies. The tile address is the tile group index of the tile group within the tile group in the subpicture, the length of tile_group_address is Ceil(Log2(num_tile_groups_in_sub_pic_minus1 + 1)) bits, and the value of tile_group_address is in the range from 0 to num_tile_groups_in_sub_pic_minus1.

[0183] The following constraints shall apply, which are the requirements for bitstream conformance. The value of tile_group_address shall not be equal to the value of tile_group_address of any other coded tile group NAL unit of the same coded subpicture. The tile groups of the subpicture shall be in ascending order of these tile_group_address values. The shape of the tile groups of the subpicture shall be such that each tile, when decoded, has its entire left boundary and entire upper boundary composed of the subpicture boundary or the boundary of a previously decoded tile.

[0184] num_tiles_in_tile_group_minus1, if present, specifies the number of tiles in the tile group minus 1. The value of num_tiles_in_tile_group_minus1 shall be in the range from 0 to NumTilesInSubPic - 1. If not present, the value of num_tiles_in_tile_group_minus1 is assumed to be equal to 0. The variable NumTilesInCurrTileGroup that specifies the number of tiles in the current tile group and the variable TgTileIdx[i] that specifies the tile index of the i-th tile in the current tile group are derived as follows.

Number

[0185] An exemplary derivation process for temporal luma motion vector prediction is as follows. The variables mvLXCol and availableFlagLXCol are derived as follows. If tile_group_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0. Otherwise (if tile_group_temporal_mvp_enabled_flag is equal to 1), the following steps in order apply. The motion vectors at the same position in the lower right, as well as the lower and right boundary sample positions are derived as follows.

Number

[0186] An exemplary luma sample bilinear interpolation process is as follows. The luma position at all sample units (xInti,yInti) is derived as follows for i = 0..1. If sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 1, the following applies.

Number

Number

[0187] An exemplary luma sample 8-tap interpolation filtering process is as follows. The luma positions at all sample units (xInti, yInti) are derived as follows for i = 0..7. When sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 1, the following applies. [Number] Otherwise (when sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 0), the following applies. [Number]

[0188] An exemplary chroma sample interpolation process is as follows. The variable xOffset is set equal to (sps_ref_wraparound_offset_minus1 + 1)*MinCbSizeY) / SubWidthC. The chroma positions at all sample units (xInti, yInti) are derived as follows for i = 0..3.

[0189] When sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 1, the following applies. [Number] Otherwise (when sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 0), the following applies. [Number]

[0190] An exemplary deblocking filter process is as follows. The deblocking filter process is applied to all coding sub-block edges and transform block edges of a picture, except for edges of the following types, namely, edges at the boundary of a picture, edges that coincide with the boundary of a sub-picture where loop_filter_across_sub_pic_enabled_flag is equal to 0, edges that coincide with the boundary of a tile where loop_filter_across_tiles_enabled_flag is equal to 0, edges that coincide with the boundary of a tile group or the upper or left boundary within a tile group where tile_group_deblocking_filter_disabled_flag is equal to 1, edges that do not correspond to the 8×8 sample grid boundary of the component being considered, edges within a chroma component where both sides of the edge use inter prediction, edges of a chroma transform block that are not edges of the associated transform unit, and edges across the luma transform block of a coding unit having an IntraSubPartitionsSplit value not equal to ISP_NO_SPLIT.

[0191] An exemplary deblocking filter process in one direction is as follows. For each coding unit having a coding block width log2CbW, a coding block height log2CbH, and the position (xCb, yCb) of the top-left sample of the coding block, when edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, the edge is filtered by the following steps in order. The coding block width nCbW is set equal to 1 << log2CbW, and the coding block height nCbH is set equal to 1 << log2CbH. The variable filterEdgeFlag is derived as follows. When edgeType is equal to EDGE_VER and one or more of the following conditions are true, filterEdgeFlag is set equal to 0. The left boundary of the current coding block is the left boundary of the picture. The left boundary of the current coding block is the left or right boundary of a sub-picture and loop_filter_across_sub_pic_enabled_flag is equal to 0. The left boundary of the current coding block is the left boundary of a tile and loop_filter_across_tiles_enabled_flag is equal to 0. Otherwise, when edgeType is equal to EDGE_HOR and one or more of the following conditions are true, the variable filterEdgeFlag is set equal to 0. The top boundary of the current luma coding block is the top boundary of the picture. The top boundary of the current coding block is the top or bottom boundary of a sub-picture and loop_filter_across_sub_pic_enabled_flag is equal to 0. The top boundary of the current coding block is the top boundary of a tile and loop_filter_across_tiles_enabled_flag is equal to 0. Otherwise, filterEdgeFlag is set equal to 1.

[0192] An exemplary CTB change process is as follows. Depending on the values of pcm_loop_filter_disabled_flag, pcm_flag[xYi][yYj], and cu_transquant_bypass_flag of the coding unit including the coding block covering recPicture[xSi][ySj], for all sample positions (xSi, ySj) and (xYi, yYj) where i = 0..nCtbSw-1 and j = 0..nCtbSh-1, the following applies. If one or more of the following conditions are true for all sample positions (xSik’, ySjk’) and (xYik’, yYjk’) where k = 0..1, edgeIdx is set equal to 0. The sample at position (xSik’, ySjk’) is outside the picture boundary. The sample at position (xSik’, ySjk’) belongs to a different subpicture and loop_filter_across_sub_pic_enabled_flag in the tile group to which the sample recPicture[xSi][ySj] belongs is equal to 0. loop_filter_across_tiles_enabled_flag is equal to 0 and the sample at position (xSik’, ySjk’) belongs to a different tile.

[0193] An exemplary filtering process for a coding tree block of luma samples is as follows. For the derivation of the reconstructed luma sample alfPictureL[x][y] after filtering, each reconstructed luma sample in the current luma coding tree block recPictureL[x][y] is filtered as follows for x, y = 0..CtbSizeY-1. The respective positions (hx, vy) of the corresponding luma samples (x, y) in a given array recPicture of luma samples are derived as follows. If loop_filter_across_tiles_enabled_flag for the tile tileA containing the luma sample at position (hx, vy) is equal to 0, and variable tileIdx is the tile index of tileA, the following applies. [Number] Otherwise, if the loop_filter_across_sub_pic_enabled_flag in the sub-picture containing the luma sample at position (hx, vy) is equal to 0, the following applies.

Number

Number

[0194] An exemplary derivation process for the ALF transposition and filter index for luma samples is as follows. The respective positions (hx, vy) of the corresponding luma samples (x, y) in a given array recPicture of luma samples are derived as follows. If the loop_filter_across_tiles_enabled_flag for the tile tileA containing the luma sample at position (hx, vy) is equal to 0, and tileIdx is the tile index of tileA, the following applies.

Number

Number

Number

[0195] The filtering process of an exemplary coding tree block for chroma samples is as follows. For the derivation of the reconstructed chroma sample alfPicture[x][y] after filtering, each reconstructed chroma sample within the current chroma coding tree block recPicture[x][y] is filtered as follows for x,y = 0..ctbSizeC-1. The respective positions (hx,vy) of the corresponding chroma samples (x,y) within a given array recPicture of chroma samples are derived as follows. If the loop_filter_across_tiles_enabled_flag for the tile tileA containing the chroma sample at position (hx,vy) is equal to 0, and tileIdx is the tile index of tileA, the following applies. [Number] Otherwise, if the loop_filter_across_sub_pic_enabled_flag for the sub-picture containing the chroma sample at position (hx,vy) is equal to 0, the following applies. [Number] Otherwise, the following applies. [Number]

[0196] The sum of the variables is derived as follows. [Number] The reconstructed chroma picture sample alfPicture[xCtbC+x][yCtbC+y] after modified filtering is derived as follows. [Number]

[0197] FIG. 12 is a schematic diagram of an exemplary video coding device 1200. The video coding device 1200 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 1200 includes a downstream port 1220, an upstream port 1250, and / or a transceiver unit (Tx / Rx) 1210 that includes a transmitter and / or a receiver for communicating data upstream and / or downstream over a network. The video coding device 1200 also includes a processor 1230 that includes a logic unit and / or a central processing unit (CPU) for processing data, and a memory 1232 for storing data. The video coding device 1200 may also include electrical, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components coupled to the upstream port 1250 and / or the downstream port 1220 for communicating data over an electrical, optical, or wireless communication network. The video coding device 1200 may also include an input and / or output (I / O) device 1260 for communicating data to and from a user. The I / O device 1260 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 1260 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0198] Processor 1230 is implemented by hardware and software. Processor 1230 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 1230 communicates with downstream port 1220, Tx / Rx 1210, upstream port 1250, and memory 1232. Processor 1230 includes coding module 1214. Coding module 1214 implements the disclosed embodiments described herein, such as methods 100, 1300, and 1400, which may use current blocks 801 and / or 901 coded according to one-way inter prediction 600 and / or bidirectional inter prediction 700 based on a candidate list generated according to loop filter 1000, bitstream 1100, picture 500, and / or pattern 900. Coding module 1214 may also implement any other method / mechanism described herein. Further, coding module 1214 may implement codec system 200, encoder 300, and / or decoder 400. For example, coding module 1214 may implement the implementation manners of the first, second, third, fourth, fifth, and / or sixth examples as described above. Thus, coding module 1214 provides additional functionality and / or coding efficiency to video coding device 1200 when coding video data. Thus, coding module 1214 improves the functionality of video coding device 1200 and addresses problems specific to video coding technology. Further, coding module 1214 results in the conversion of video coding device 1200 to different states.Alternatively, the coding module 1214 is implemented as instructions stored in the memory 1232 and can be executed by the processor 1230 (e.g., as a computer program product stored in a non-transitory medium).

[0199] The memory 1232 includes one or more memory types such as a disk, a tape drive, a solid state drive, a read only memory (ROM), a random access memory (RAM), a flash memory, a ternary content-addressable memory (TCAM), a static random-access memory (SRAM), etc. The memory 1232 may be used as an overflow data storage device to store such programs when a program is selected for execution and to store instructions and data read during program execution.

[0200] FIG. 13 is a flowchart of an exemplary method 1300 for encoding a video sequence into a bitstream such as bitstream 1100 while applying a clipping function in an interpolation filter such as interpolation filter 913 when a subpicture such as subpicture 510 is treated as a picture such as picture 500. The method 1300 may be used by an encoder such as codec system 200, encoder 300, and / or video coding device 1200 when executing method 100 for encoding current blocks 801 and / or 901 according to unidirectional inter prediction 600 and / or bidirectional inter prediction 700 with respect to a candidate list generated by using loop filter 1000 and / or according to pattern 900.

[0201] Method 1300 may start when an encoder receives a video sequence including a plurality of pictures and determines to encode the video sequence into a bitstream, for example, based on user input. In step 1301, the encoder divides the current picture into sub-pictures. The encoder also divides the sub-pictures into blocks. In step 1303, the encoder determines to encode the blocks according to inter prediction. Accordingly, the encoder selects a motion vector for encoding the blocks.

[0202] In step 1305, the encoder applies a clipping function to the sample positions in the reference block pointed to by the motion vector. The clipping function is applied to support the application of an interpolation filter when the motion vector points outside the sub-picture. This process occurs when a flag is set to indicate that the sub-picture is treated as a picture, and thus should be coded to support independent extraction from other sub-pictures within the picture. In this context, when a sub-picture is coded without referring to data within other sub-pictures, the sub-picture is treated as a picture and can thus be extracted separately.

[0203] In step 1307, the encoder applies an interpolation filter to the result of the clipping function to obtain predicted sample values. In one example, the interpolation filter includes a luma sample bilinear interpolation process. In such a case, the block includes a block of luma samples. Further, the predicted sample values include predicted luma sample values. In this case, the luma sample bilinear interpolation process receives an input including the luma position at all sample units (xIntL, yIntL). The luma sample bilinear interpolation process outputs predicted luma sample values (predSampleLXL). In this case, the clipping function of step 1305 is applied to the sample positions as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that the sub - picture is treated as a picture, SubPicIdx is the index of the sub - picture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the sub - picture, SubPicLeftBoundaryPos is the position of the left boundary of the sub - picture, SubPicTopBoundaryPos is the position of the upper boundary of the sub - picture, SubPicBotBoundaryPos is the position of the lower boundary of the sub - picture, Clip3 is a clipping function that follows:

Number

[0204] In another example, the interpolation filter includes a luma sample 8 - tap interpolation filtering process. In this case, the block includes a block of luma samples, and the predicted sample value includes the predicted luma sample value. Also, the luma sample 8 - tap interpolation filtering process receives an input including the luma position at the entire sample unit (xIntL, yIntL). The luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). Further, the clipping function in step 1305 is applied to the sample position as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i - 3), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i - 3) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, Clip3 is a clipping function that follows the following,

Number

[0205] In other examples, the interpolation filter includes a chroma sample interpolation process. In this case, the block includes a block of chroma samples, and the predicted sample value includes the predicted chroma sample value. In this case, the chroma sample interpolation process receives an input including the chroma positions at all sample units (xIntC, yIntC). Further, the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC). In this case, the clipping function of step 1305 is applied to the sample positions as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xIntC + i), and yInti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yIntC + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, SubWidthC and SubHeightC indicate the horizontal and vertical sampling rate ratios between luma samples and chroma samples, Clip3 is a clipping function that follows:

Number

[0206] In step 1309, the encoder encodes the block into a bitstream based on the predicted sample value and the motion vector. Then, the bitstream may be stored for communication towards the decoder.

[0207] FIG. 14 is a flowchart of an exemplary method 1400 for decoding a video sequence from a bitstream such as bitstream 1100 while applying a clipping function in an interpolation filter such as interpolation filter 913 when a subpicture such as subpicture 510 is treated as a picture such as picture 500. Method 1400 may be used by a decoder such as codec system 200, decoder 400, and / or video coding device 1200 when executing method 100 for decoding current blocks 801 and / or 901 according to one-directional inter prediction 600 and / or bi-directional inter prediction 700 with respect to a candidate list generated by using in-loop filter 1000 and / or according to pattern 900.

[0208] Method 1400 may begin when the decoder begins to receive a bitstream of coded data representing a video sequence, e.g., as a result of method 1300. In step 1401, the decoder receives a bitstream including a current picture including a subpicture. Pictures, subpictures, slices, tiles, CTUs, and / or other sub-regions thereof are coded according to inter prediction. In step 1402, the decoder determines a motion vector for a block of the subpicture.

[0209] In step 1403, the decoder applies a clipping function to a sample position within a reference block pointed to by the motion vector. The clipping function is applied to support the application of the interpolation filter when the motion vector points outside the subpicture. This process occurs when a flag is set to indicate that the subpicture is treated as a picture and should thus be coded to support independent extraction from other subpictures within the picture.

[0210] In step 1405, the decoder applies an interpolation filter to the result of the clipping function to obtain a predicted sample value. In one example, the interpolation filter includes a luma sample bilinear interpolation process. In such a case, the block includes a block of luma samples. Further, the predicted sample value includes a predicted luma sample value. In this case, the luma sample bilinear interpolation process receives an input including the luma position at all sample units (xIntL, yIntL). The luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). In this case, the clipping function of step 1403 is applied to the sample position as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, Clip3 is a clipping function that follows:

Number

[0211] In another example, the interpolation filter includes a luma sample 8-tap interpolation filtering process. In this case, the block includes a block of luma samples, and the predicted sample value includes a predicted luma sample value. Also, the luma sample 8-tap interpolation filtering process receives an input including the luma position at all sample units (xIntL, yIntL). The luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). Further, the clipping function in step 1403 is applied to the sample position as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL + i - 3), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i - 3) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, and Clip3 is a clipping function that follows

Number

[0212] In other examples, the interpolation filter includes a chroma sample interpolation process. In this case, the block includes a block of chroma samples, and the predicted sample values include predicted chroma sample values. In this case, the chroma sample interpolation process receives an input including the chroma position at all sample units (xIntC, yIntC). Further, the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC). In this case, the clipping function of step 1403 is applied to the sample position as follows. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies. xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xIntC + i), and yInti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yIntC + i) Here, subpic_treated_as_pic_flag is a flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, SubWidthC and SubHeightC indicate the horizontal and vertical sampling rate ratios between luma samples and chroma samples, Clip3 is a clipping function that follows the following, [Number] Here, according to the fact that x, y, and z are numerical input values, it is applied to the sample position.

[0213] In step 1407, the decoder decodes the block based on the predicted sample value. Then, the decoder can transfer the block for display as part of the decoded video sequence.

[0214] FIG. 15 is a schematic diagram of an exemplary system 1500 for coding a video sequence of an image in a bitstream 1100 such as a bitstream while applying a clipping function in an interpolation filter such as interpolation filter 913 when a subpicture such as subpicture 510 is treated as a picture such as picture 500. System 1500 may be implemented by an encoder and decoder such as codec system 200, encoder 300, decoder 400 and / or video coding device 1200. Further, system 1500 may be used when implementing methods 100, 1300 and / or 1400 for coding current blocks 801 and / or 901 according to one-way inter prediction 600 and / or bidirectional inter prediction 700 with respect to a candidate list generated by using in-loop filter 1000 and / or according to pattern 900.

[0215] System 1500 includes a video encoder 1502. The video encoder 1502 includes a partitioning module 1503 for partitioning a current picture into sub-pictures and partitioning the sub-pictures into blocks. The video encoder 1502 further includes a determination module 1504 for determining to encode a block according to inter prediction. The video encoder 1502 further includes a selection module 1505 for selecting a motion vector for encoding a block. The video encoder 1502 further includes an application module 1506 for applying a clipping function to a sample position in a reference block to support the application of an interpolation filter when the motion vector points outside the sub-picture and a flag is set to indicate that the sub-picture is treated as a picture. The application module 1506 is also for applying an interpolation filter to the result of the clipping function to obtain a predicted sample value. The video encoder 1502 further includes an encoding module 1507 for encoding a block into a bitstream based on the predicted sample value and the motion vector. The video encoder 1052 further includes a storage module 1508 for storing the bitstream for communication to a decoder. The video encoder 1502 further includes a transmission module 1509 for transmitting the bitstream towards a video decoder 1510. The video encoder 1502 may be further configured to execute any of the steps of method 1300.

[0216] System 1500 also includes a video decoder 1510. The video decoder 1510 includes a receiving module 1511 for receiving a bitstream including a current picture that includes sub-pictures coded according to inter prediction. The video decoder 1510 further includes a determining module 1512 for determining motion vectors for blocks of sub-pictures. The video decoder 1510 further includes an applying module 1513 for applying a clipping function to sample positions in a reference block to support the application of an interpolation filter when the motion vector points outside the sub-picture and a flag is set to indicate that the sub-picture is to be treated as a picture. The applying module 1513 is also for applying an interpolation filter to the result of the clipping function to obtain predicted sample values. The video decoder 1510 further includes a decoding module 1514 for decoding a block based on the predicted sample values. The video decoder 1510 includes a transferring module 1515 for transferring the block for display as part of a decoded video sequence. The video decoder 1510 may be further configured to perform any of the steps of method 1400.

[0217] The first component is directly coupled to the second component when there are no intervening components other than the line, trace, or other medium between the first and second components. The first component is indirectly coupled to the second component when there are intervening components other than the line, trace, or other medium between the first and second components. The term "coupled" and its variants include both being directly coupled and being indirectly coupled. The use of the term "about" means a range including ± 10% of the subsequent number unless otherwise specifically noted.

[0218] It should also be understood that the steps of the exemplary methods described herein need not necessarily be performed in the order described, and that the order of such method steps should be understood to be merely exemplary. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in ways consistent with various embodiments of the present disclosure.

[0219] Although several embodiments are provided in the present disclosure, it can be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. This example is to be considered illustrative and not restrictive, and the intention is not to be limited to the details given herein. For example, various elements or components may be combined or integrated into other systems, or certain features may be omitted or not implemented.

[0220] Furthermore, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of modifications, substitutions, and changes may be ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. A method implemented in a decoder, comprising: receiving, by a receiver of the decoder, a bitstream including a current picture including a sub-picture coded according to inter prediction, wherein the sub-picture is a rectangular or square region of two or more slices / tile groups within the picture; determining, by a processor of the decoder, a reference block for a block of the sub-picture; indicating, by a flag, that the sub-picture is to be treated as a picture, and when a motion vector for the block points outside the sub-picture, applying a clipping function to an integer part of a sample position within the reference block to limit the integer part to be within the sub-picture, wherein the sample position is used for an interpolation filter, the interpolation filter includes a luma sample bilinear interpolation process, and the block includes a block of luma samples; applying, by the processor, the interpolation filter to a result of the clipping function to obtain a predicted sample value, wherein the predicted sample value includes a predicted luma sample value; decoding, by the processor, the block based on the predicted sample value; and a method comprising the steps of:

2. The luma sample bilinear interpolation process receives an input including the luma positions in all sample units (xInt L , yInt L ), the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLX L ), and the clipping function follows as follows, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L + i), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i) wherein subpic_treated_as_pic_flag is the flag set to indicate that the sub-picture is to be treated as a picture, SubPicIdx is an index of the sub-picture, xInti and yInti are clipped sample positions at index i, SubPicRightBoundaryPos is a position of a right boundary of the sub-picture, SubPicLeftBoundaryPos is a position of a left boundary of the sub-picture, SubPicTopBoundaryPos is a position of an upper boundary of the sub-picture, SubPicBotBoundaryPos is a position of a lower boundary of the sub-picture, and Clip3 is the clipping function according to the following, 【Number 49】 wherein x, y, and z are numerical input values, and the method according to claim 1, applied to the integer part.

3. The interpolation filter includes a luma sample 8-tap interpolation filtering process, the block includes a block of luma samples, and the predicted sample value includes a predicted luma sample value. The method according to claim 1 or 2.

4. The luma sample 8-tap interpolation filtering process receives an input including the luma position in all sample units (xInt L , yInt L ), the luma sample 8-tap interpolation filtering process outputs a predicted luma sample value (predSampleLX L ), and the clipping function is as follows, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L + i - 3), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i - 3) Here, subpic_treated_as_pic_flag is the flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, Clip3 is the clipping function according to the following, 【Number 50】 Here, according to x, y, and z being numerical input values, the method according to claim 3 applied to the integer part.

5. The interpolation filter includes a chroma sample interpolation process, the block includes a block of chroma samples, and the predicted sample value includes a predicted chroma sample value. The method according to any one of claims 1 to 4.

6. The chroma sample interpolation process receives an input including chroma positions in all sample units (xInt C , yInt C ), the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLX C ), and the clipping function follows the following, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xInt C + i), and yInti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yInt C + i) Here, subpic_treated_as_pic_flag is the flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, SubWidthC and SubHeightC indicate the horizontal and vertical sampling rate ratios between luma samples and chroma samples, and Clip3 is the clipping function according to the following, 【Number 51】 A method according to claim 5, wherein x, y, and z are numerical input values, and the method is applied to the integer part. **Claim 7** A method implemented in an encoder, comprising: a step of dividing a current picture into sub-pictures and further dividing the sub-pictures into blocks by a processor of the encoder, wherein the sub-pictures are rectangular or square regions of two or more slice / tile groups within the picture; a step of obtaining a reference block for encoding the block by the processor; a step of indicating by a flag that the sub-picture is to be treated as a picture, and when a motion vector for the block points outside the sub-picture, applying a clipping function to the integer part of a sample position within the reference block to limit the integer part to be within the sub-picture, wherein the sample position is used for an interpolation filter, the interpolation filter includes a luma sample bilinear interpolation process, and the block includes a block of luma samples; a step of applying the interpolation filter to the result of the clipping function to obtain a predicted sample value by the processor, wherein the predicted sample value includes a predicted luma sample value; a step of encoding the block into a bitstream by the processor based on the predicted sample value and the motion vector and a method including the above steps. **Claim 8** The luma sample bilinear interpolation process receives an input including luma positions in all sample units (xInt L , yInt L ), the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLX L ), and the clipping function follows the following, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L + i), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i) Here, subpic_treated_as_pic_flag is the flag set to indicate that the sub-picture is to be treated as a picture, SubPicIdx is the index of the sub-picture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the sub-picture, SubPicLeftBoundaryPos is the position of the left boundary of the sub-picture, SubPicTopBoundaryPos is the position of the upper boundary of the sub-picture, SubPicBotBoundaryPos is the position of the lower boundary of the sub-picture, and Clip3 is the clipping function according to the following: 【Number 52】 The method according to claim 7, wherein x, y, and z are numerical input values, and are applied to the integer part accordingly.

9. The method according to claim 7 or 8, wherein the interpolation filter includes a luma sample 8-tap interpolation filtering process, the block includes a block of luma samples, and the predicted sample value includes a predicted luma sample value.

10. The luma sample 8-tap interpolation filtering process receives an input including the luma position in all sample units (xInt L , yInt L ), the luma sample 8-tap interpolation filtering process outputs a predicted luma sample value (predSampleLX L ), and the clipping function is as follows, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xInt L + i - 3), and yInti = Clip3(SubPicTopBoundaryPos, SubPicBotBoundaryPos, yInt L + i - 3) Here, subpic_treated_as_pic_flag is the flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, and Clip3 is the clipping function according to the following, 【Number 53】 The method according to claim 9, wherein x, y, and z are numerical input values, and are applied to the integer part accordingly.

11. The method according to any one of claims 7 to 10, wherein the interpolation filter includes a chroma sample interpolation process, the block includes a block of chroma samples, and the predicted sample value includes a predicted chroma sample value.

12. The chroma sample interpolation process receives an input including chroma positions in all sample units (xInt C , yInt C ), the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLX C ), and the clipping function follows the following, that is, when subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies, xInti = Clip3(SubPicLeftBoundaryPos / SubWidthC, SubPicRightBoundaryPos / SubWidthC, xInt C + i), and yInti = Clip3(SubPicTopBoundaryPos / SubHeightC, SubPicBotBoundaryPos / SubHeightC, yInt C + i) Here, subpic_treated_as_pic_flag is the flag set to indicate that the subpicture is treated as a picture, SubPicIdx is the index of the subpicture, xInti and yInti are the clipped sample positions at index i, SubPicRightBoundaryPos is the position of the right boundary of the subpicture, SubPicLeftBoundaryPos is the position of the left boundary of the subpicture, SubPicTopBoundaryPos is the position of the upper boundary of the subpicture, SubPicBotBoundaryPos is the position of the lower boundary of the subpicture, SubWidthC and SubHeightC indicate the horizontal and vertical sampling rate ratios between luma samples and chroma samples, Clip3 is the clipping function that follows the following, 【Number 54】 Here, according to the fact that x, y, and z are numerical input values, the method according to claim 11 applied to the integer part.

13. A video decoding device, including a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to execute the method of any one of claims 1 to 6, a video decoding device.

14. A video encoding device, including a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to execute the method of any one of claims 7 to 12, a video encoding device.

15. A non-transitory computer-readable medium including a computer program product for use by a video decoding device, wherein the computer program product includes computer-executable instructions stored in the non-transitory computer-readable medium that, when executed by a processor, cause the video decoding device to execute the method of any one of claims 1 to 6, a non-transitory computer-readable medium.

16. A non-transitory computer-readable medium including a computer program product for use by a video encoding device, wherein the computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video encoding device to execute the method according to any one of claims 7 to 12. **Claim 17** A decoder, a receiving unit configured to receive a bitstream including a current picture including a sub-picture encoded according to inter prediction, wherein the sub-picture is a rectangular or square region of two or more slice / tile groups within the picture; a determination unit configured to determine a reference block for a block of the sub-picture; an application unit configured such that a flag indicates that the sub-picture is to be treated as a picture, and when a motion vector for the block points outside the sub-picture, a clipping function is applied to an integer part of a sample position within the reference block to limit the integer part to be within the sub-picture, the sample position being used for an interpolation filter, the interpolation filter including a luma sample bilinear interpolation process, the block including a block of luma samples, and the interpolation filter is applied to a result of the clipping function to obtain a predicted sample value including a predicted luma sample value; a decoding unit configured to decode the block based on the predicted sample value and including the decoder. **Claim 18** The decoder according to claim 17, further configured to execute the method according to any one of claims 1 to 6. **Claim 19** An encoder, a partitioning unit configured to partition a current picture into sub-pictures and partition the sub-pictures into blocks, wherein the sub-pictures are rectangular or square regions of two or more slice / tile groups within the picture; an acquisition unit configured to acquire a reference block for encoding the block A flag indicates that the sub-picture is to be treated as a picture. When the motion vector for the block points outside the sub-picture, a clipping function is applied to the integer part of the sample position within the reference block to limit the integer part to be within the sub-picture. The sample position is used for an interpolation filter, the interpolation filter includes a luma sample bilinear interpolation process, the block includes a block of luma samples, and an application unit configured to apply the interpolation filter to the result of the clipping function to obtain a predicted sample value including a predicted luma sample value. An encoding unit configured to encode the block into a bitstream based on the predicted sample value and the motion vector. An encoder including the above.

20. The encoder according to claim 19, further configured to execute the method according to any one of claims 7 to 12.

21. A computer program including program code for executing the method according to any one of claims 1 to 6 when executed on a computer or a processor.

22. A computer program including program code for executing the method according to any one of claims 7 to 12 when executed on a computer or a processor.