Interpolation filter clipping for sub-image motion vectors

By introducing the sub-image_treated_as_pic_flag flag and the use of the limiting function in video decoding, the error problem caused by motion vector pointing to external data during independent extraction and decoding of sub-images is solved, and a more stable and efficient video encoding and decoding process is achieved.

CN118631988BActive Publication Date: 2025-05-16HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410744572.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-29
Filing Date
2020-03-11
Publication Date
2025-05-16
Estimated Expiration
2040-03-11

AI Technical Summary

Technical Problem

In video decoding, there are errors in independent extraction and decoding of sub-images, especially when the motion vector points to outside the sub-image, the lack of decoded data leads to errors.

Method used

A flag sub-image _treated_as_pic_flag is introduced to indicate that the sub-image can be treated as an image. When this flag is set and the motion vector points outside the sub-image, a limiting function is applied to support the use of an interpolation filter, ensuring independent extraction and decoding of the sub-image.

Benefits of technology

Through the combination of this flag and the limiting function, errors caused by motion vector pointing to external data during sub-image extraction are avoided, and the stability and efficiency of the video codec are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118631988B_ABST
    Figure CN118631988B_ABST
Patent Text Reader

Abstract

The present invention discloses a video decoding mechanism. The mechanism includes: receiving a code stream including a current image, wherein the current image includes a sub-image decoded according to inter-frame prediction; determining a motion vector of a block in the sub-image; when the motion vector points outside the sub-image and when a flag is set to indicate that the sub-image is treated as an image, applying a clipping function to a sample position in a reference block to support the application of an interpolation filter; applying the interpolation filter to the result of the clipping function to obtain a predicted sample value; decoding the block according to the predicted sample value; and forwarding the block for display as part of a decoded video sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202080020571.5, and the original application date is March 11, 2020. The entire contents of the original application are incorporated into this application by reference.

[0002] Cross-reference to related applications

[0003] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 816,751, filed by Ye-Kui Wang et al. on March 11, 2019, entitled “Sub-Picture Based Video Coding”, and the benefit of U.S. Provisional Patent Application No. 62 / 826,659, filed by Ye-Kui Wang et al. on March 29, 2019, entitled “Sub-Picture Based Video Coding”, the entire contents of which are incorporated herein by reference. Technical Field

[0004] The present invention relates generally to video coding, and more particularly to coding sub-pictures in pictures in video coding. Background Art

[0005] Even if a video is relatively short, a large amount of video data is required to describe it, which can cause difficulties when the data is to be streamed or otherwise transmitted over a communication network with limited bandwidth capacity. Therefore, video data is often compressed before being transmitted over modern telecommunications networks. Since memory resources may be limited, the size of the video may also be an issue when storing the video on a storage device. Video compression devices typically use software and / or hardware on the source side to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. A video decompression device that decodes the video data then receives the compressed data on the destination side. With limited network resources and a growing demand for higher video quality, there is a need for improved compression and decompression techniques that can increase the compression ratio with little impact on image quality. Summary of the invention

[0006] In one embodiment, the present invention includes a method implemented in a decoder. The method includes: a receiver of the decoder receives a code stream including a current image, wherein the current image includes a sub-image decoded according to inter-frame prediction; a processor of the decoder determines a motion vector of a block in the sub-image; when the motion vector points outside the sub-image and when a flag is set to indicate that the sub-image is processed as an image, the processor applies a clipping function to a sample position in a reference block to support the application of an interpolation filter; the processor applies the interpolation filter to the result of the clipping function to obtain a predicted sample value; the processor decodes the block according to the predicted sample value. Inter-frame prediction can be performed according to one of several inter-frame prediction modes. Some inter-frame prediction modes generate candidate lists of motion vector prediction values ​​on both the encoder and decoder sides. In this way, the encoder can indicate a motion vector by indicating an index in the candidate list instead of indicating the motion vector itself. In addition, some systems decode sub-images for independent extraction. Independent extraction allows the current sub-image to be decoded and displayed without decoding information in other sub-images. This situation may cause errors when using motion vectors pointing outside the sub-image, because the data pointed to by the motion vector may not be decoded and therefore cannot be used. The present invention includes a flag indicating that a sub-image can be processed as an image. When the current sub-image is processed as an image, the current sub-image can be extracted without reference to other sub-images. Specifically, this example uses a limiting function applied when applying an interpolation filter. This limiting function ensures that the interpolation filter does not rely on data in adjacent sub-images to keep the sub-images separate, thereby supporting separate extraction. Therefore, when the flag is set and the motion vector points outside the current sub-image, the limiting function is applied. Then, the interpolation filter is applied to the result of the limiting function. Therefore, this example provides an additional function for the video codec, namely preventing errors when performing sub-image extraction.

[0007] Optionally, according to any of the above aspects, in another implementation of the aspect, the interpolation filter includes a luma sample bilinear interpolation process, the block includes a luma sample block, and the predicted sample value includes a predicted luma sample value.

[0008] Optionally, according to any of the above aspects, in another implementation of the aspect, the luma sample bilinear interpolation process receives an input including a luma position (xIntL, yIntL) having an integer number of sample units, the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), and the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0009] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i),

[0010] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i),

[0011] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula:

[0012]

[0013] Where x, y, and z are numeric input values.

[0014] Optionally, according to any of the above aspects, in another implementation of the aspect, the interpolation filter includes a luma sample 8-tap interpolation filtering process, the block includes a luma sample block, and the predicted sample value includes a predicted luma sample value.

[0015] Optionally, according to any of the above aspects, in another implementation of the aspect, the luma sample 8-tap interpolation filtering process receives an input including a luma position (xIntL, yIntL) having an integer number of sample units, and the luma sample 8-tap interpolation filtering process outputs a predicted luma sample value (predSampleLXL), and the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0016] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i–3),

[0017] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i–3),

[0018] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 is the clipping function according to the following formula:

[0019]

[0020] Where x, y, and z are numeric input values.

[0021] Optionally, according to any of the above aspects, in another implementation of the aspect, the interpolation filter includes a chroma sample interpolation process, the block includes a chroma sample block, and the predicted sample value includes a predicted chroma sample value.

[0022] Optionally, according to any of the above aspects, in another implementation of the aspect, the chroma sample interpolation process receives an input including a chroma position (xIntC, yIntC) having an integer number of sample units, the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC), and the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0023] xInti=Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xIntC+i),

[0024] yInti=Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yIntC+i),

[0025] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as an image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, SubWidthC and SubHeightC represent the horizontal sampling rate ratio and the vertical sampling rate ratio between the luminance sample and the chrominance sample, and Clip3 represents the clipping function according to the following formula:

[0026]

[0027] Where x, y, and z are numeric input values.

[0028] In one embodiment, the present invention includes a method implemented in an encoder. The method includes: a processor of the encoder divides a current image into sub-images and divides the sub-images into blocks; the processor determines to encode the blocks according to inter-frame prediction; the processor selects a motion vector to encode the blocks; when the motion vector points outside the sub-image and when a flag is set to indicate that the sub-image is treated as an image, the processor applies a clipping function to a sample position in a reference block to support the application of an interpolation filter; the processor applies the interpolation filter to the result of the clipping function to obtain a predicted sample value; the processor encodes the block into a code stream according to the predicted sample value and the motion vector; a memory coupled to the processor stores the code stream for transmission to a decoder. Inter-frame prediction can be performed according to one of several inter-frame prediction modes. Some inter-frame prediction modes generate candidate lists of motion vector prediction values ​​at both the encoder and decoder sides. In this way, the encoder can indicate the motion vector by indicating an index in the candidate list instead of indicating the motion vector itself. In addition, some systems encode sub-images for independent extraction. Independent extraction allows the current sub-image to be decoded and displayed without encoding the information in other sub-images. This situation may cause errors when using motion vectors pointing outside the sub-image, because the data pointed to by the motion vector may not be encoded and therefore cannot be used. The present invention includes a flag indicating that the sub-image can be processed as an image. When the current sub-image is processed as an image, the current sub-image can be extracted without reference to other sub-images. Specifically, this example uses a limiting function applied when applying an interpolation filter. This limiting function ensures that the interpolation filter does not rely on data in adjacent sub-images to keep the sub-images separate, thereby supporting separate extraction. Therefore, when the flag is set and the motion vector points outside the current sub-image, the limiting function is applied. Then, the interpolation filter is applied to the result of the limiting function. Therefore, this example provides an additional function for the video codec, namely preventing errors when performing sub-image extraction.

[0029] Optionally, according to any of the above aspects, in another implementation of the aspect, the interpolation filter includes a luma sample bilinear interpolation process, the block includes a luma sample block, and the predicted sample value includes a predicted luma sample value.

[0030] Optionally, according to any of the above aspects, in another implementation of the aspect, the luma sample bilinear interpolation process receives an input including a luma position (xIntL, yIntL) having an integer number of sample units, the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), and the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0031] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i),

[0032] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i),

[0033] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula:

[0034]

[0035] Where x, y, and z are numeric input values.

[0036] Optionally, according to any of the above aspects, in another implementation of the aspect, the interpolation filter includes a luma sample 8-tap interpolation filtering process, the block includes a luma sample block, and the predicted sample value includes a predicted luma sample value.

[0037] Optionally, according to any of the above aspects, in another implementation of the aspect, the luma sample 8-tap interpolation filtering process receives an input including a luma position (xIntL, yIntL) having an integer number of sample units, and the luma sample 8-tap interpolation filtering process outputs a predicted luma sample value (predSampleLXL), and the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0038] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i–3),

[0039] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i–3),

[0040] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula:

[0041]

[0042] Where x, y, and z are numeric input values.

[0043] Optionally, according to any of the above aspects, in another implementation of the aspect, the interpolation filter includes a chroma sample interpolation process, the block includes a chroma sample block, and the predicted sample value includes a predicted chroma sample value.

[0044] Optionally, according to any of the above aspects, in another implementation of the aspect, the chroma sample interpolation process receives an input including a chroma position (xIntC, yIntC) having an integer number of sample units, the chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC), and the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0045] xInti=Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xIntC+i),

[0046] yInti=Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yIntC+i),

[0047] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as an image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, SubWidthC and SubHeightC represent the horizontal sampling rate ratio and the vertical sampling rate ratio between the luminance sample and the chrominance sample, and Clip3 represents the clipping function according to the following formula:

[0048]

[0049] Where x, y, and z are numeric input values.

[0050] In one embodiment, the present invention includes a video decoding device, which includes a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to execute the method according to any of the above aspects.

[0051] In one embodiment, the present invention includes a non-transitory computer-readable medium. The non-transitory computer-readable medium includes a computer program product for use by a video decoding device, the computer program product includes computer executable instructions stored in the non-transitory computer-readable medium, and when a processor executes the computer executable instructions, the video decoding device performs the method according to any of the above aspects.

[0052] In one embodiment, the present invention includes a decoder. The decoder includes: a receiving module for receiving a code stream including a current image, wherein the current image includes a sub-image decoded according to inter-frame prediction; a determining module for determining a motion vector of a block in the sub-image; an applying module for: applying a clipping function to a sample position in a reference block to support application of an interpolation filter when the motion vector points outside the sub-image and when a flag is set to indicate that the sub-image is to be processed as an image; applying the interpolation filter to a result of the clipping function to obtain a predicted sample value; a decoding module for decoding the block according to the predicted sample value; and a forwarding module for forwarding the block for display as part of a decoded video sequence.

[0053] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the decoder is also used to execute the method according to any of the above aspects.

[0054] In one embodiment, the present invention includes an encoder. The encoder includes: a segmentation module for segmenting a current image into sub-images and segmenting the sub-images into blocks; a determination module for determining to encode the blocks according to inter-frame prediction; a selection module for selecting a motion vector to encode the blocks; an application module for: when the motion vector points outside the sub-image and when a flag is set to indicate that the sub-image is treated as an image, applying a clipping function to a sample position in a reference block to support application of an interpolation filter; applying the interpolation filter to a result of the clipping function to obtain a predicted sample value; an encoding module for encoding the blocks into a bitstream according to the predicted sample value and the motion vector; and a storage module for storing the bitstream for sending to a decoder.

[0055] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the encoder is also used to execute the method according to any of the above aspects.

[0056] Optionally, according to any of the above aspects, in another implementation manner of the aspect, the encoder is also used to execute the method according to any of the above aspects.

[0057] For the sake of clarity, any of the above-described embodiments may be combined with any one or more of the other embodiments described above to create new embodiments within the scope of the invention.

[0058] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] For a more complete understanding of the present invention, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0060] Figure 1 A flow chart of an exemplary method of decoding a video signal.

[0061] Figure 2 is a schematic diagram of an exemplary encoding and decoding (codec) system for video coding.

[0062] Figure 3 is a schematic diagram of an exemplary video encoder.

[0063] Figure 4 is a schematic diagram of an exemplary video decoder.

[0064] Figure 5A Schematic diagram of an exemplary image segmented into sub-images.

[0065] Figure 5B Schematic diagram of an exemplary sub-image divided into strips.

[0066] Figure 5C A schematic diagram of an exemplary strip divided into blocks.

[0067] Figure 5D A schematic diagram of an exemplary slice partitioned into coding tree units (CTUs).

[0068] Figure 6 A schematic diagram of an example of unidirectional inter-frame prediction.

[0069] Figure 7 A schematic diagram of an example of bidirectional inter-frame prediction.

[0070] Figure 8 FIG. 4 is a diagram illustrating an example of coding a current block according to candidate motion vectors from neighboring coded blocks.

[0071] Fig. 9 A schematic diagram of an exemplary mode for determining a motion vector candidate list.

[0072] Fig.10 is a block diagram of an exemplary in-loop filter.

[0073] Fig.11 A schematic diagram of an exemplary code stream including decoding tool parameters to support decoding of a sub-image in an image.

[0074] Fig.12 is a schematic diagram of an exemplary video decoding device.

[0075] Fig.13 A flow chart of an exemplary method for encoding a video sequence into a bitstream when a clipping function is applied to an interpolation filter when sub-images are processed as images.

[0076] Fig.14 A flow chart of an exemplary method for decoding a video sequence from a bitstream when a clipping function is applied to an interpolation filter when sub-images are processed as images.

[0077] Fig.15 Schematic diagram of an exemplary system for decoding a video sequence composed of images in a codestream when a clipping function is applied to an interpolation filter when sub-images are processed as images. DETAILED DESCRIPTION

[0078] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.

[0079] The following abbreviations are used in this document: Adaptive Loop Filter (ALF), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Joint Video Experts Team (JVET), Motion-Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sample Adaptive Offset (SAO), Sequence Parameter Set (SPS), Temporal Motion Vector Prediction (TMVP), Versatile Video Coding (VVC), and Working Draft (WD).

[0080] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques can include performing spatial (e.g., intra-frame) prediction and / or temporal (e.g., inter-frame) prediction to reduce or remove data redundancy in a video sequence. For block-based video decoding, a video slice (e.g., a video image or a portion of a video image) can be divided into video blocks, which can also be referred to as treeblocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in intra-frame decoded (I) slices in an image are decoded using spatial prediction relative to reference samples in adjacent blocks in the same image, while video blocks in inter-frame decoded unidirectional prediction (P) or bidirectional prediction (B) slices in an image can be decoded using spatial prediction relative to reference samples in adjacent blocks in the same image, or can be decoded using temporal prediction relative to reference samples in other reference images. A picture / image may be referred to as a frame, and a reference picture may be referred to as a reference frame. Spatial prediction or temporal prediction produces a prediction block representing an image block. The residual data represents the pixel difference between the original image block and the prediction block. Thus, an inter-coded block is encoded based on a motion vector pointing to a block of reference samples forming a prediction block and residual data representing the difference between the decoded block and the prediction block; while an intra-coded block is encoded based on an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to the transform domain. These produce residual transform coefficients (referred to as quantized transform coefficients for short) that may be quantized. The quantized transform coefficients may initially be arranged in a two-dimensional array. The quantized transform coefficients may be scanned in order to produce a one-dimensional vector of transform coefficients. Entropy coding may be used to achieve further compression. These video compression techniques are discussed in more detail below.

[0081] To ensure that the encoded video can be correctly decoded, the video is encoded and decoded according to the corresponding video coding standards. Video coding standards include International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extended versions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three-dimensional (3D) AVC (3D-AVC). HEVC includes extended versions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC has begun to develop a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD), which includes JVET-M1001-v6, which provides algorithm descriptions, encoding end descriptions in VVC WD, and reference software.

[0082] In order to decode a video image, the image is first segmented, and the partitions are then decoded into a bitstream. There are various image segmentation schemes. For example, an image can be segmented into regular slices, non-independent slices, tiles, and / or segmented according to Wavefront Parallel Processing (WPP). For simplicity, HEVC restricts the encoder so that only regular slices, non-independent slices, tiles, WPP, and combinations thereof can be used when segmenting slices into CTB groups for video decoding. This segmentation can be used to support maximum transfer unit (MTU) size matching, parallel processing, and reduce end-to-end latency. MTU represents the maximum amount of data that can be sent in a single data packet. If the data packet payload exceeds the MTU, the payload is divided into two data packets through a process called fragmentation.

[0083] A regular strip, also referred to as a strip, is a segmented portion of an image that can be reconstructed independently of other regular strips in the same image, but there are still dependencies due to the presence of loop filtering operations. Each regular strip is encapsulated in its own Network Abstraction Layer (NAL) unit for transmission. In addition, to support independent reconstruction, in-picture prediction / intra sample prediction (motion information prediction, decoding mode prediction) and entropy decoding dependencies across strip boundaries can be disabled. This independent reconstruction supports parallel processing. For example, parallel processing based on regular strips reduces inter-processor or inter-core communication. However, since each regular strip is independent, each strip is associated with a separate strip header. Since there is a bit cost for the strip header of each strip and the lack of cross-strip boundary prediction, the use of regular strips may result in a large amount of decoding overhead. In addition, regular strips can be used to support MTU size matching requirements. Specifically, since regular slices are encapsulated in separate NAL units and can be decoded independently, each regular slice needs to be smaller than the MTU in the MTU scheme to avoid splitting the slice into multiple packets. Therefore, in order to achieve parallel processing and MTU size matching, the slice layout in the image may contradict each other.

[0084] A non-independent slice is similar to a regular slice, but the slice header is shorter and the image tree block boundary can be segmented without affecting intra-frame prediction. Therefore, a non-independent slice can split a regular slice into multiple NAL units, which can reduce end-to-end latency by first completing the encoding of the entire regular slice and then sending part of the regular slice.

[0085] A tile is a partition of an image created by horizontal and vertical boundaries that create tile columns and tile rows. Tiles can be decoded in raster scan order (right to left, top to bottom). The scan order of CTBs is the order in which the scan is performed within a tile. Therefore, the CTBs in the first tile are first decoded in raster scan order, and then proceed to the CTBs in the next tile. Similar to regular slices, tiles affect intra prediction dependencies as well as entropy decoding dependencies. However, tiles may not be included in individual NAL units, so tiles may not be used for MTU size matching. Each tile can be processed by one processor / core, and inter-processor / inter-core communication for intra prediction between processing units for adjacent tiles may be limited to sending shared slice headers (when adjacent tiles are in the same slice) and performing loop filtering related sharing of reconstructed samples and metadata. When a slice includes multiple partitions, in addition to the first entry point offset in the slice, the entry point byte offset of each partition can be indicated (signaled) in the slice header. For each slice and partition, at least one of the following conditions needs to be met: (1) all decoded tree blocks in the slice belong to the same partition; (2) all decoded tree blocks in the partition belong to the same slice.

[0086] In WPP, the image is partitioned into multiple single rows of CTBs. Entropy decoding and prediction mechanisms can use data from CTBs in other rows. Parallel processing may be achieved through parallel decoding of CTB rows. For example, the current row can be decoded in parallel with the previous row. However, the decoding of the current row is delayed by two CTBs compared to the decoding process of the previous rows. This delay ensures that data related to the upper CTB and the upper right CTB of the current CTB in the current row are available before the current CTB is decoded. When this method is represented graphically, it looks like a wavefront. This staggered start decoding can achieve parallel processing using as many processors / cores as the number of CTB rows included in the image. Since intra-frame prediction is supported between adjacent tree blocks in each row in the image, a large amount of inter-processor / inter-core communication may be required to achieve intra-frame prediction. WPP partitioning does not take into account the NAL unit size. Therefore, WPP does not support MTU size matching. However, conventional stripes can be used in conjunction with WPP, resulting in certain decoding overhead, so as to achieve MTU size matching as needed.

[0087] Tiles may also include motion constrained tilesets. A motion constrained tileset (MCTS) is a set of tiles designed so that the associated motion vectors are constrained to point to whole sample positions within the MCTS and to fractional sample positions that only require interpolation of whole sample positions within the MCTS. In addition, motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction are prohibited. In this way, each MCTS can be decoded independently without the need for tiles outside the MCTS. Temporal MCTS supplemental enhancement information (SEI) messages can be used to indicate the presence of MCTS in the bitstream and indicate these MCTS. MCTS SEI messages provide supplemental information (expressed as part of the SEI message semantics) that can be used for MCTS substream extraction to generate a consistent bitstream for the MCTS set. The information includes multiple extraction information sets, each of which defines multiple MCTS sets and includes raw bytes sequence payload (RBSP) bytes of replacement video parameter sets (VPS), sequence parameter sets (SPS) and picture parameter sets (PPS) to be used in the MCTS substream extraction process. When extracting a substream according to the MCTS substream extraction process, since one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) can use different values ​​in the extracted substream, the parameter sets (VPS, SPS and PPS) can be rewritten or replaced and the slice header can be updated.

[0088] The image may also be segmented into one or more sub-images. Segmenting the image into sub-images may allow different portions of the image to be processed differently from a decoding perspective. For example, a sub-image may be extracted and displayed without extracting other sub-images. As another example, different sub-images may be displayed at different resolutions, repositioned relative to each other (e.g., in teleconferencing applications), or may be decoded as separate images even if the sub-images collectively include data from the same image.

[0089] An exemplary implementation of a sub-image is as follows. An image can be divided into one or more sub-images. A sub-image is a rectangular or square set of strips / blocks starting with a strip / block group with an address equal to 0. Each sub-image can refer to a different PPS, so each sub-image can adopt a different segmentation mechanism. During the decoding process, the sub-image can be treated as an image. The current reference image used to decode the current sub-image can be generated by extracting an area juxtaposed with the current sub-image from a reference image in a decoded image buffer. The extracted area can be a decoded sub-image, so inter-frame prediction can occur between sub-images of the same size and position in the image. The block group can be a series of blocks under a block raster scan of a sub-image. The following content can be derived to determine the position of a sub-image in an image. Each sub-image can be included in the next unoccupied position in the image arranged in the CTU raster scan order, which is large enough to fit a sub-image at the image boundary.

[0090] The sub-image schemes adopted by various video decoding systems have various problems that reduce decoding efficiency and / or functionality. The present invention includes various technical solutions to solve these problems. In a first exemplary problem, inter-frame prediction can be performed according to one of several inter-frame prediction modes. Some inter-frame prediction modes generate candidate lists of motion vector prediction values ​​on both the encoder and decoder sides. In this way, the encoder can indicate the motion vector by indicating an index in the candidate list, rather than indicating the motion vector itself. In addition, some systems encode sub-images for independent extraction. Independent extraction allows the current sub-image to be decoded and displayed without decoding the information in other sub-images. This situation may cause errors when using motion vectors pointing outside the sub-image, because the data pointed to by the motion vector may not be decoded and therefore cannot be used.

[0091] Therefore, in a first example, this document discloses a flag indicating that a sub-image can be processed as an image. The flag is set to support separate extraction of sub-images. When the flag is set, the motion vector prediction values ​​obtained from the collocated block only include motion vectors pointing to the inside of the sub-image, and any motion vector prediction values ​​pointing to the outside of the sub-image are excluded. This ensures that motion vectors pointing to the outside of the sub-image are not selected, avoiding related errors. A collocated block is a block in an image that is different from the current image. Motion vector prediction values ​​from blocks in the current image (non-collocated blocks) can point to the outside of the sub-image because other processes such as interpolation filters can avoid errors caused by these motion vector prediction values. Therefore, this example provides an additional function for a video encoder / decoder (codec), namely preventing errors from occurring when performing sub-image extraction.

[0092] In a second example, this document discloses a flag indicating that a sub-image can be processed as an image. When the current sub-image is processed as an image, the current sub-image can be extracted without reference to other sub-images. Specifically, this example uses a clipping function used when applying an interpolation filter. This clipping function ensures that the interpolation filter does not rely on data in adjacent sub-images to keep the sub-images separate, thereby supporting separate extraction. Therefore, when the flag is set and the motion vector points outside the current sub-image, the clipping function is applied. Then, the interpolation filter is applied to the result of the clipping function. Therefore, this example provides an additional function for the video codec, namely preventing errors when performing sub-image extraction. In this way, the first example and the second example solve the first exemplary problem.

[0093] In the second exemplary problem, the video decoding system divides the image into sub-images, slices, tiles and / or coding tree units, which are then divided into blocks. These blocks are then encoded for sending to the decoder. Decoding these blocks may produce decoded images including various noises. To address these problems, the video decoding system can apply various filters across block boundaries. These filters can remove blocking, quantization noise and other bad decoding artifacts. As described above, some systems encode sub-images for independent extraction. Independent extraction allows the current sub-image to be encoded and displayed without encoding the information in other sub-images. In these systems, sub-images can be divided into blocks for encoding. Therefore, block boundaries along the edges of sub-images can be aligned with sub-image boundaries. In some cases, block boundaries can also be aligned with partition boundaries. Filters can be applied across these block boundaries, and therefore can also be applied across sub-image boundaries and / or partition boundaries. This may cause errors when the current sub-image is extracted independently, because the filtering process may be performed in an unexpected way when data in adjacent sub-images is not available.

[0094] In a third example, a flag is disclosed herein that controls sub-image level filtering. When the flag is set for a sub-image, filters may be applied across sub-image boundaries. When the flag is not set, filters are not applied across sub-image boundaries. In this way, filters may be disabled for sub-images that are encoded for individual extraction, and enabled for sub-images that are encoded for grouped display. Thus, this example provides additional functionality to the video codec, namely preventing filter-related errors when performing sub-image extraction.

[0095] In a fourth example, a flag that can be set to control block-level filtering is disclosed herein. When the flag is set for a block, the filter can be applied across block boundaries. When the flag is not set, the filter will not be applied across block boundaries. In this way, the filter can be disabled or enabled for use at block boundaries (e.g., while continuing to filter the internal portion of the block). Therefore, this embodiment provides additional functionality for the video codec, namely, supporting selective filtering across block boundaries. In this way, the third and fourth examples solve the second exemplary problem.

[0096] In a third exemplary problem, a video decoding system may partition an image into sub-images. In this way, different sub-images may be processed in different ways when decoding a video. For example, sub-images may be extracted and displayed separately, resized independently based on application-level changes, and so on. In some cases, sub-images may be created by partitioning an image into blocks and assigning the blocks to sub-images. Some video decoding systems describe sub-image boundaries based on the blocks included in the sub-image. However, the blocking scheme may not be applicable to some images. Therefore, these boundary descriptions may limit the use of sub-images for images that employ blocking.

[0097] In the fifth example, a mechanism for indicating a sub-image boundary according to CTB and / or CTU is disclosed herein. Specifically, the width and height of the sub-image can be indicated in units of CTB. In addition, the position of the upper left CTU in the sub-image can be indicated as an offset from the upper left CTU in the image, and the measurement unit is CTB. The size of the CTU and CTB can be set to a predetermined value. Therefore, indicating the sub-image size and position according to the CTB and CTU provides sufficient information to the decoder to locate the sub-image for display. In this way, the sub-image can also be used in the case where blocking is not adopted. In addition, this indication mechanism avoids complexity and can be decoded using relatively few bits. Therefore, this example provides an additional function for the video codec, that is, the sub-image is used independently of the block. In addition, this example improves the decoding efficiency, thereby reducing the use of processor resources, memory resources and / or network resources on the encoder and decoder sides. In this way, the fifth example solves the third exemplary problem.

[0098] In the fourth exemplary problem, an image may be divided into multiple strips for encoding. In some video decoding systems, these strips are addressed according to their position relative to the image. Still other video decoding systems employ the concept of sub-images. As described above, from a decoding perspective, sub-images may be processed differently from other sub-images. For example, sub-images may be extracted and displayed independently of other sub-images. In this case, the strip addresses generated based on the image position may stop functioning properly due to the omission of a large number of expected strip addresses. Some video decoding systems solve this problem by dynamically rewriting the strip header upon request to change the strip address to support sub-image extraction. This process may require a lot of resources because it may occur every time a user requests to view a sub-image.

[0099] In a sixth example, the present invention discloses addressing a strip relative to a sub-image that includes the strip. For example, a strip header may include a sub-image identifier (ID) and an address of each strip included in the sub-image. In addition, a sequence parameter set (SPS) may include dimensions of the sub-image to which the sub-image ID may refer. Therefore, when a request is made to extract a sub-image separately, there is no need to rewrite the strip header. The strip header and SPS include sufficient information to support positioning the strip in the sub-image for display. Therefore, this example improves decoding efficiency and / or avoids redundancy in strip header rewriting, thereby reducing the use of processor resources, memory resources, and / or network resources on the encoder and / or decoder side. In this way, the sixth example solves the fourth exemplary problem.

[0100] Figure 1 Flowchart of an exemplary method 100 for decoding a video signal. Specifically, the video signal is encoded at the encoder side. The encoding process compresses the video signal by using various mechanisms to reduce the size of the video file. The file size is smaller, and the compressed video file can be used to send to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is generally the inverse of the encoding process so that the video signal reconstructed by the decoder can be consistent with the video signal at the encoder side.

[0101] In step 101, a video signal is input into an encoder. For example, the video signal may be an uncompressed video file stored in a memory. As another example, the video file may be captured by a video capture device such as a camera and encoded to support live streaming of the video. The video file may include an audio component and a video component. The video component includes a series of image frames. When these image frames are viewed in sequence, they give a visual effect of motion. These frames include pixels represented by light, referred to herein as brightness components (or brightness samples), and pixels represented by color, referred to as chrominance components (or chrominance samples). In some examples, these frames may also include depth values ​​to support three-dimensional viewing.

[0102] In step 103, the video is segmented into blocks. Segmentation includes subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), the frames are first divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). These CTUs include luma samples and chroma samples. The coding tree can be used to divide the CTU into blocks, and then the blocks are repeatedly subdivided until a configuration that supports further encoding is obtained. For example, the luma component of a frame can be subdivided into blocks including relatively uniform luma values. In addition, the chroma component of a frame can be subdivided into blocks including relatively uniform color values. Therefore, the segmentation mechanism varies depending on the content of the video frame.

[0103] In step 105, various compression mechanisms are used to compress the image blocks obtained by segmentation in step 103. For example, inter-frame prediction and / or intra-frame prediction can be used. Inter-frame prediction is designed to take advantage of the fact that objects in a scene often appear in consecutive frames. In this way, the blocks describing the objects in the reference frame do not need to be described repeatedly in adjacent frames. Specifically, an object (such as a table) can remain in a fixed position in multiple frames. Therefore, the table is described once, and adjacent frames can refer back to the reference frame. Pattern matching mechanisms can be used to match objects over multiple frames. In addition, moving objects can be represented across multiple frames due to reasons such as object movement or camera movement. In a specific example, a video can show a car moving across the screen over multiple frames. Motion vectors can be used to describe this movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in one frame to the coordinates of the object in a reference frame. Therefore, inter-frame prediction can encode the image blocks in the current frame as a set of motion vectors, representing the offset of the image blocks in the current frame from the corresponding blocks in the reference frame.

[0104] Intra prediction encodes blocks in a common frame. Intra prediction exploits the fact that luminance and chrominance components tend to be clustered in one frame. For example, a patch of green in a part of a tree tends to be adjacent to similar patches of green. Intra prediction uses a variety of directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. These directional modes indicate that the samples of the current block are similar / identical to the samples of the neighboring blocks in the corresponding direction. The planar mode indicates that a series of blocks on a row / column (e.g., a plane) can be interpolated based on the neighboring blocks on the edge of the row. The planar mode actually represents the smooth transition of light / color across rows / columns by taking a relatively constant slope of the changing value. The DC mode is used for boundary smoothing and indicates that the block is similar / identical to the average of the samples of all neighboring blocks, which are related to the angular direction of the directional prediction mode. Therefore, the intra prediction block can represent the image block as various relational prediction mode values ​​instead of the actual value. In addition, the inter prediction block can represent the image block as a motion vector value instead of the actual value. In either case, the prediction block may not accurately represent the image block in some cases. All differences are stored in residual blocks. Transformations can be applied to the residual blocks to further compress the file.

[0105] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may produce blocky images on the decoder side. In addition, a block-based prediction scheme can encode the block and then reconstruct the encoded block for subsequent use as a reference block. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to the block / frame. These filters reduce block artifacts so that the encoded file can be accurately reconstructed. In addition, these filters reduce artifacts in the reconstructed reference block, so that the artifacts are less likely to produce other artifacts in subsequent blocks encoded according to the reconstructed reference block.

[0106] Once the video signal has been segmented, compressed and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the data described above and any indicative data required to support appropriate reconstruction of the video signal on the decoder side. For example, these data may include segmentation data, prediction data, residual blocks, and various flags that provide decoding instructions to the decoder. The bitstream can be stored in a memory so that it can be sent to the decoder upon request. The bitstream can also be broadcast and / or multicast to multiple decoders. Creating a bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 can be performed continuously and / or simultaneously in multiple frames and blocks. Figure 1The order shown is for purposes of clarity and ease of discussion and is not intended to limit the video coding process to a particular order.

[0107] In step 111, the decoder receives the code stream and begins the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the code stream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data in the code stream to determine the segmented parts of the frame. The segmentation should match the result of the block segmentation in step 103. The entropy encoding / decoding used in step 111 is described below. The encoder makes many choices during the compression process, for example, selecting a block segmentation scheme from several possible choices based on the spatial placement of values ​​in one or more input images. A large number of bits (bins) may be used to indicate the exact choice. "Bit" as used herein is a binary value that is a variable (for example, a bit value that may vary depending on the content). Entropy encoding causes the encoder to discard any options that are obviously not suitable for a particular situation, leaving a set of available options. Then, a codeword is assigned to each available option. The length of the codeword depends on the number of allowable options (for example, one bit corresponds to two options, two bits correspond to three or four options, and so on). Then, the encoder encodes the codeword corresponding to the selected option. This scheme reduces the codeword because the codeword is as large as expected, uniquely indicating a selection from a small subset of permissible options rather than a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of permissible options in a similar manner to the encoder. By determining the set of permissible options, the decoder can read the codeword and determine the selection made by the encoder.

[0108] In step 113, the decoder performs block decoding. Specifically, the decoder uses an inverse transform to generate a residual block. The decoder then uses the residual block and the corresponding prediction block to reconstruct the image block according to the segmentation. The prediction block may include an intra-frame prediction block and an inter-frame prediction block generated by the encoder in step 105. Next, the reconstructed image block is placed in a frame of the reconstructed video signal according to the segmentation data determined in step 111. The syntax for step 113 may also be indicated in the bitstream by entropy coding as described above.

[0109] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107 on the encoder side. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and a SAO filter can be applied to the frame to remove blocking artifacts. Once the frame is filtered, in step 117, the video signal can be output to a display for viewing by an end user.

[0110] Figure 2Schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the codec system 200 provides functions to support the implementation of the operating method 100. The codec system 200 is used broadly to describe components used on both the encoder and decoder sides. The codec system 200 receives a video signal and segments the video signal, as described in conjunction with steps 101 and 103 in the operating method 100, to obtain segmented video signals 201. Then, the codec system 200 compresses the segmented video signal 201 into an encoded bitstream when acting as an encoder, as described in conjunction with steps 105, 107, and 109 in the method 100. The codec system 200 generates an output video signal from the bitstream when acting as a decoder, as described in conjunction with steps 111, 113, 115, and 117 in the operating method 100. The codec system 200 includes a general decoder control component 211, a transform scaling and quantization component 213, an intra-frame estimation component 215, an intra-frame prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded image buffer component 223, and a header format and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. Figure 2 In FIG. 2 , the black lines represent the movement of the data to be encoded / decoded, and the dotted lines represent the movement of the control data that controls the operation of other components. The components in the codec system 200 may all be present in the encoder. The decoder may include a subset of the components in the codec system 200. For example, the decoder may include an intra-frame prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded image buffer component 223. These components are described below.

[0111] The segmented video signal 201 is a captured video sequence that has been segmented into pixel blocks by a coding tree. The coding tree uses various partitioning modes to subdivide pixel blocks into smaller pixel blocks. These blocks can then be further subdivided into smaller blocks. These blocks can be called nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. In some cases, the divided blocks are included in a coding unit (coding unit, CU). For example, a CU can be a sub-part of a CTU, including a luminance block, one or more red difference chrominance (Cr) blocks and one or more blue difference chrominance (Cb) blocks and corresponding syntax instructions of the CU. The partitioning mode may include a binary tree (binary tree, BT), a ternary tree (triple tree, TT) and a quad tree (quad tree, QT), which are used to divide the nodes into two, three or four child nodes of different shapes according to the partitioning mode adopted. The segmented video signal 201 is forwarded to the general decoder control component 211, the transform scaling and quantization component 213, the intra-frame estimation component 215, the filter control analysis component 227 and the motion estimation component 221 for compression.

[0112] The universal decoder control component 211 is used to make decisions related to encoding images in a video sequence into a bitstream according to application constraints. For example, the universal decoder control component 211 manages the optimization of the bitrate / bitstream size relative to the reconstruction quality. These decisions can be made based on storage space / bandwidth availability and image resolution requests. The universal decoder control component 211 also manages buffer utilization based on transmission speed to alleviate buffer underload and overload problems. To address these problems, the universal decoder control component 211 manages segmentation, prediction and filtering performed by other components. For example, the universal decoder control component 211 can dynamically increase the compression complexity to increase resolution and increase bandwidth utilization, or reduce the compression complexity to reduce resolution and bandwidth utilization. Therefore, the universal decoder control component 211 controls other components in the codec system 200 to balance the video signal reconstruction quality and bitrate issues. The universal decoder control component 211 generates control data, which is used to control the operation of other components. The control data is also forwarded to the header format and CABAC component 231 for encoding into the codestream to instruct the decoder on the parameters to be used for decoding.

[0113] The segmented video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter-frame prediction. The frames or slices of the segmented video signal 201 can be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-frame prediction decoding on the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple decoding passes to select an appropriate decoding mode for each block of video data, and so on.

[0114] The motion estimation component 221 and the motion compensation component 219 can be highly integrated, but for conceptual purposes, they are described separately. The motion estimation performed by the motion estimation component 221 is the process of generating motion vectors, which are used to estimate the motion of video blocks. For example, a motion vector can represent the displacement of a coded object relative to a prediction block. A prediction block is a block that is found to be highly matched to a block to be coded in terms of pixel difference. A prediction block can also be referred to as a reference block. This pixel difference can be determined by the sum of absolute difference (SAD), the sum of square difference (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into multiple CTBs, and then the CTBs can be divided into CBs to be included in a CU. A CU can be encoded as a prediction unit (PU) including prediction data and / or a transform unit (TU) including transform residual data of a CU. The motion estimation component 221 uses rate-distortion analysis as part of a rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. of the current block / frame, and may select a reference block, motion vector, etc. having an optimal rate-distortion characteristic. The optimal rate-distortion characteristic balances the quality of video reconstruction (e.g., the amount of data lost due to compression) and decoding efficiency (e.g., the size of the final encoding).

[0115] In some examples, the codec system 200 may calculate values ​​for sub-integer pixel positions of a reference image stored in the decoded image buffer component 223. For example, the video codec system 200 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation component 221 may perform motion search relative to integer pixel positions and fractional pixel positions, and output motion vectors with fractional pixel precision. The motion estimation component 221 calculates the motion vector of a PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a prediction block of the reference image. The motion estimation component 221 outputs the calculated motion vector as motion data to the header format and CABAC component 231 for encoding, and outputs it as motion data to the motion compensation component 219.

[0116] The motion compensation performed by the motion compensation component 219 may involve obtaining or generating a prediction block based on the motion vector determined by the motion estimation component 221. Likewise, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector of the PU of the current video block, the motion compensation component 219 may locate the prediction block to which the motion vector points. The pixel values ​​of the prediction block are then subtracted from the pixel values ​​of the current video block being decoded to obtain pixel differences, thereby forming a residual video block. In general, the motion estimation component 221 performs motion estimation with respect to the luma component, while the motion compensation component 219 uses the motion vector calculated based on the luma component for the chroma component and the luma component. The prediction block and the residual block are forwarded to the transform scaling and quantization component 213.

[0117] The segmented video signal 201 is also sent to the intra-frame estimation component 215 and the intra-frame prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-frame estimation component 215 and the intra-frame prediction component 217 can be highly integrated, but for conceptual purposes, they are described separately. The intra-frame estimation component 215 and the intra-frame prediction component 217 perform intra-frame prediction on the current block relative to the blocks in the current frame, replacing the inter-frame prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra-frame estimation component 215 determines an intra-frame prediction mode to encode the current block. In some examples, the intra-frame estimation component 215 selects a suitable intra-frame prediction mode from a plurality of tested intra-frame prediction modes to encode the current block. The selected intra-frame prediction mode is then forwarded to the header format and CABAC component 231 for encoding.

[0118] For example, the intra-frame estimation component 215 performs rate-distortion analysis on various tested intra-frame prediction modes to calculate rate-distortion values, and selects the intra-frame prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to generate the encoded block, and determines the code rate (e.g., number of bits) used to generate the encoded block. The intra-frame estimation component 215 calculates a ratio based on the distortion and rate of the various encoded blocks to determine the intra-frame prediction mode that exhibits the best rate-distortion value for the block. In addition, the intra-frame estimation component 215 can be used to decode the depth block of the depth image using the depth modeling mode (DMM) according to rate-distortion optimization (RDO).

[0119] The intra prediction component 217 can generate a residual block from the prediction block according to the selected intra prediction mode determined by the intra estimation component 215 when implemented on an encoder, or can read the residual block from the bitstream when implemented on a decoder. The residual block includes the difference between the prediction block and the original block, represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra estimation component 215 and the intra prediction component 217 can operate on the luminance component and the chrominance component.

[0120] The transform scaling and quantization component 213 is used to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform to the residual block, thereby generating a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used. The transform may convert the residual information from a pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component 213 is also used to scale the transform residual information according to frequency, etc. This scaling involves applying a scaling factor to the residual information so as to quantize different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then scan the matrix including the quantized transform coefficients. The quantized transform coefficients are forwarded to the header format and CABAC component 231 for encoding into the codestream.

[0121] The scaling and inverse transform component 229 performs operations opposite to the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, inverse transform and / or inverse quantization to reconstruct a residual block in the pixel domain, for example, for subsequent use as a reference block. The reference block can become a prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate a reference block by adding the residual block back to the corresponding prediction block for motion estimation of subsequent blocks / frames. Filters are applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization and transformation. These artifacts may make the prediction inaccurate (and generate additional artifacts) when predicting subsequent blocks.

[0122] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 can be combined with the corresponding prediction block from the intra-frame prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter can then be applied to the reconstructed image block. In some examples, the filter can be applied to the residual block instead. Figure 2 Like other components in, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and can be implemented together, but for conceptual purposes, they are described separately. The filters applied to the reconstructed reference block are applied to specific spatial regions, and these filters include multiple parameters to adjust the way these filters are used. The filter control analysis component 227 analyzes the reconstructed reference block to determine the locations where these filters need to be used and sets the corresponding parameters. This data is forwarded to the header format and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies these filters based on the filter control data. These filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. These filters can be applied in the spatial domain / pixel domain (for example, for reconstructed pixel blocks) or the frequency domain according to the example.

[0123] When operating as an encoder, the filtered reconstructed image blocks, residual blocks and / or prediction blocks are stored in the decoded image buffer component 223 for subsequent use in motion estimation, as described above. When operating as a decoder, the decoded image buffer component 223 stores the reconstructed and filtered blocks and forwards them to the display as part of the output video signal. The decoded image buffer component 223 can be any storage device capable of storing prediction blocks, residual blocks and / or reconstructed image blocks.

[0124] The header format and CABAC component 231 receives data from various components of the codec system 200 and encodes the data into a decoded bitstream for sending to the decoder. Specifically, the header format and CABAC component 231 generates various headers to encode control data (such as general control data and filter control data). In addition, prediction data (including intra-frame prediction data and motion data) and residual data in the form of quantized transform coefficient data are encoded into the bitstream. The final bitstream includes all the information required by the decoder to reconstruct the original segmented video signal 201. This information may also include an intra-frame prediction mode index table (also called a codeword mapping table), definitions of coding contexts for various blocks, indications of the most likely intra-frame prediction mode, indications of segmentation information, etc. This data can be encoded using entropy coding. For example, the information may be encoded using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding or other entropy coding techniques. After entropy coding, the encoded bitstream may be sent to another device (e.g., a video decoder) or archived for subsequent transmission or retrieval.

[0125] Figure 3 is a block diagram of an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or perform step 101, step 103, step 105, step 107, and / or step 109 in the operation method 100. The encoder 300 segments the input video signal to obtain segmented video signals 301 that are substantially similar to the segmented video signals 201. Then, the segmented video signals 301 are compressed by the components in the encoder 300 and encoded into a bitstream.

[0126] Specifically, the segmented video signal 301 is forwarded to the intra prediction component 317 for intra prediction. The intra prediction component 317 can be substantially similar to the intra estimation component 215 and the intra prediction component 217. The segmented video signal 301 is also forwarded to the motion compensation component 321 for inter prediction based on the reference block in the decoded image buffer component 323. The motion compensation component 321 can be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction block and the residual block from the intra prediction component 317 and the motion compensation component 321 are forwarded to the transform and quantization component 313 for transform and quantization of the residual block. The transform and quantization component 313 can be substantially similar to the transform scaling and quantization component 213. The transform quantized residual block and the corresponding prediction block (together with the related control data) are forwarded to the entropy coding component 331 for decoding into the bitstream. The entropy coding component 331 can be substantially similar to the header format and CABAC component 231.

[0127] The transformed quantized residual block and / or the corresponding prediction block are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 to be reconstructed as a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 can be substantially similar to the scaling and inverse transform component 229. According to an example, the in-loop filter in the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block. The in-loop filter component 325 can be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 can include a plurality of filters, as described in conjunction with the in-loop filter component 225. The filtered block is then stored in the decoded image buffer component 323 for use as a reference block for the motion compensation component 321. The decoded image buffer component 323 can be substantially similar to the decoded image buffer component 223.

[0128] Figure 4 4 is a block diagram of an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or perform step 111, step 113, step 115 and / or step 117 in the operation method 100. The decoder 400 receives a bitstream from the encoder 300, etc., and generates a reconstructed output video signal according to the bitstream for display to an end user.

[0129] The code stream is received by the entropy decoding component 433. The entropy decoding component 433 is used to perform an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE decoding or other entropy decoding techniques. For example, the entropy decoding component 433 can use header information to provide context to parse additional data encoded as code words in the code stream. The decoded information includes any information required to decode the video signal, such as general control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients in the residual block. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 to reconstruct the residual block. The inverse transform and quantization component 429 can be similar to the inverse transform and quantization component 329.

[0130] The reconstructed residual block and / or prediction block are forwarded to the intra prediction component 417 to be reconstructed into an image block according to the intra prediction operation. The intra prediction component 417 can be similar to the intra estimation component 215 and the intra prediction component 217. Specifically, the intra prediction component 417 uses the prediction mode to locate the reference block in the frame, and adds the residual block to the above result to reconstruct the intra prediction image block. The reconstructed intra prediction image block and / or residual block and the corresponding inter prediction data are forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 and the in-loop filter component 425 can be substantially similar to the decoded image buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block and / or prediction block. This information is stored in the decoded image buffer component 423. The reconstructed image block from the decoded image buffer component 423 is forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 can be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vector of the reference block to generate a prediction block, and applies the residual block to the above result to reconstruct the image block. The obtained reconstructed block can also be forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 continues to store other reconstructed image blocks. These reconstructed image blocks can be reconstructed into frames through segmentation information. These frames can also be placed in a sequence. The sequence is output to the display as a reconstructed output video signal.

[0131] Figure 5A Schematic diagram of an exemplary image 500 segmented into sub-images 510. For example, the image 500 may be segmented by the codec system 200 and / or the encoder 300 for encoding and by the codec system 200 and / or the decoder 400 for decoding. For another example, the image 500 may be segmented by the encoder in step 103 of the method 100 for use by the decoder in step 111.

[0132] Image 500 is an image that describes a complete visual portion of a video sequence at a specified time position. Image (picture / image) 500 can also be called a frame. Image 500 can be represented by a picture order count (POC). POC is an index that represents the output order / display order of image 500 in a video sequence. Image 500 can be divided into sub-images 510. Sub-image 510 is a rectangular area or square area of ​​one or more strips / block groups in image 500. Sub-image 510 is optional, so some video sequences include sub-images 510, while other video sequences do not. Although four sub-images 510 are shown, image 500 can be divided into any number of sub-images 510. The division of sub-images 510 can be consistent throughout the entire coded video sequence (coded video sequence, CVS).

[0133] Sub-images 510 can be used to process different areas in image 500 in different ways. For example, specified sub-images 510 can be extracted independently and sent to a decoder. In a specific example, a user can see a subset of image 500 using a virtual reality (VR) helmet, which can give the user a sense of being in the space described by image 500. In this case, streaming the sub-images 510 that are likely to be displayed to the user may improve decoding efficiency. For another example, different sub-images 510 can be processed differently in certain applications. In a specific example, a teleconferencing application can display an active speaker at a higher resolution in a more prominent position than the user who is not currently speaking. Placing different users in different sub-images 510 supports real-time configuration of the displayed image to support such functionality.

[0134] Each sub-image 510 can be identified by a unique sub-image ID, which can be consistent for the entire CVS. For example, a sub-image 510 located in the upper left corner of the image 500 can have a sub-image ID of 0. In this case, the upper left sub-image 510 of any image 500 in the sequence can be referenced by a sub-image ID of 0. In addition, each sub-image 510 can include a defined configuration, which can be consistent for the entire CVS. For example, a sub-image 510 can include a height, a width, and / or an offset. The height and width describe the size of the sub-image 510, and the offset describes the position of the sub-image 510. For example, the sum of the widths of all sub-images 510 in a row is equal to the width of the image 500. In addition, the sum of the heights of all sub-images 510 in a column is equal to the height of the image 500. In addition, the offset represents the position of the upper left corner of the sub-image 510 relative to the upper left corner of the image 500. The height, width, and offset of the sub-image 510 provide sufficient information to locate the corresponding sub-image 510 in the image 500. Since the segmentation of the sub-image 510 may remain consistent across the entire CVS, the parameters associated with the sub-image may be included in a sequence parameter set (SPS).

[0135] Figure 5B Schematic diagram of an exemplary sub-image 510 divided into slices 515. As shown, a sub-image 510 in an image 500 may include one or more slices 515. A slice 515 includes an integer number of complete blocks or an integer number of consecutive complete CTU rows within a block, which are included only in a single network abstraction layer (NAL) unit. Although four slices 515 are shown, a sub-image 510 may include any number of slices 515. A slice 515 includes visual data specific to the image 500 of a specified POC. Therefore, parameters related to the slice 515 may be included in a picture parameter set (PPS) and / or a slice header.

[0136] Figure 5CSchematic diagram of an exemplary strip 515 divided into tiles 517. As shown, a strip 515 in an image 500 may include one or more tiles 517. Tiles 517 may be created by dividing the image 500 into rectangular rows and columns. Therefore, a tile 517 is a rectangular or square area consisting of CTUs within a specific tile column and a specific tile row in the image. Tiles are optional, so some video sequences include tiles 517 while other video sequences do not. Although four tiles 517 are shown, a strip 515 may include any number of tiles 517. Tiles 517 may include visual data specific to a strip 515 in an image 500 of a specified POC. In some cases, a strip 515 may also be included in a tile 517. Therefore, parameters related to a block 517 may be included in a PPS and / or a slice header.

[0137] Figure 5D Schematic diagram of an exemplary slice 515 partitioned into CTUs 519. As shown, a slice 515 (or a block 517 in a slice 515) in an image 500 may include one or more CTUs 519. A CTU 519 is a region in an image 500 that is subdivided by a coding tree to generate a decoding block for encoding / decoding. A CTU 519 may include a luma sample in a black and white image 500 or a combination of a luma sample and a chroma sample in a color image 500. A group of luma samples or chroma samples that can be partitioned by a coding tree is called a coding tree block (CTB) 518. Therefore, a CTU 519 includes a CTB 518 composed of luma samples in an image 500 having three sample arrays and two corresponding CTBs 518 composed of chroma samples, or a CTB 518 composed of samples in a black and white image or an image that is decoded using three separate color planes and syntax structures, wherein these syntax structures are used to decode these samples.

[0138] As described above, the image 500 can be divided into sub-images 510, strips 515, blocks 517, CTUs 519, and / or CTBs 518, which are then divided into blocks. These blocks are then encoded for sending to a decoder. Decoding these blocks may produce decoded images that include various noises. To address these problems, the video decoding system can apply various filters across block boundaries. These filters can remove blocking effects, quantization noise, and other bad decoding artifacts. As described above, sub-images 510 can be used when performing independent extraction. In this case, the current sub-image 510 can be decoded and displayed without decoding the information in other sub-images 510. Therefore, the block boundaries along the edges of the sub-images 510 can be aligned with the sub-image boundaries. In some cases, the block boundaries can also be aligned with the partition boundaries. Filters can be applied across these block boundaries, and therefore can also be applied across sub-image boundaries and / or partition boundaries. This may cause errors when the current sub-image 510 is extracted independently, because the filtering process may be performed in an unexpected manner when data in the adjacent sub-images 510 is not available.

[0139] To address these issues, a flag may be used to control sub-image 510 level filtering. For example, the flag may be represented as loop_filter_across_subpic_enabled_flag. When the flag is set for a sub-image 510, the filter may be applied across the corresponding sub-image boundary. When the flag is not set, the filter will not be applied across the corresponding sub-image boundary. In this way, the filter may be disabled for sub-images 510 encoded for individual extraction and enabled for sub-images 510 encoded for grouped display. Another flag may be set to control tile 517 level filtering. The flag may be represented as loop_filter_across_tiles_enabled_flag. When the flag is set for a tile 517, the filter may be applied across the tile boundary. When the flag is not set, the filter will not be applied across the tile boundary. In this way, the filter may be disabled or enabled so as to be used at the tile boundary (e.g., while continuing to filter the internal portion of the tile). As described herein, the filter is applied across the boundary of a sub-image 510 or a tile 517 when applied to samples on both sides of the boundary.

[0140] As also described above, tiling is optional. However, some video decoding systems describe sub-image boundaries according to the tiles 517 included in the sub-image 510. In these systems, the image 500 using the tiles 517 is restricted to use the sub-image 510 according to the sub-image boundary description of the tiles 517. In order to expand the applicability of the sub-image 510, the sub-image 510 can be described according to the boundaries, CTBs 518, and / or CTUs 519. Specifically, the width and height of the sub-image 510 can be indicated in units of CTBs 518. In addition, the position of the upper left CTU 519 in the sub-image 510 can be indicated as an offset from the upper left CTU 519 in the image 500, and the measurement unit is CTB 518. The size of the CTU 519 and the CTB 518 can be set to a predetermined value. Therefore, indicating the sub-image size and position according to the CTBs 518 and CTUs 519 provides sufficient information to the decoder to position the sub-image 510 for display. In this way, sub-image 510 can also be used when blocking 517 is not used.

[0141] In addition, some video decoding systems address the slices 515 according to their positions relative to the image 500. This creates a problem when the sub-image 510 is decoded for independent extraction and display. In this case, the slice 515 and the corresponding address associated with the omitted sub-image 510 are also omitted. Omitting the address of the slice 515 may prevent the decoder from correctly locating the slice 515. Some video decoding systems solve this problem by dynamically rewriting the address in the slice header associated with the slice 515. Since the user can request any sub-image, this rewriting occurs every time the user requests a video, which requires a lot of resources. To solve this problem, when the sub-image 510 is used, the slice 515 is addressed relative to the sub-image 510 including the slice 515. For example, the slice 515 can be identified by an index or other value unique to the sub-image 510 including the slice 515. The slice address can be decoded into the slice header associated with the slice 515. The sub-image ID of the sub-image 510 including the slice 515 can also be encoded into the slice header. In addition, the size / configuration of the sub-image 510 may be decoded into the SPS together with the sub-image ID. Therefore, the decoder may obtain the configuration of the sub-image 510 from the SPS according to the sub-image ID, and locate the slice 515 in the sub-image 510 without referring to the complete image 500. Therefore, when the sub-image 510 is extracted, the slice header rewriting may be omitted, which greatly reduces the resource usage on the encoder, decoder, and / or corresponding slicer side.

[0142] Once the image 500 is segmented into CTBs 518 and / or CTUs 519, the CTBs 518 and / or CTUs 519 can be further divided into coded blocks. The coded blocks can then be coded according to intra-frame prediction and / or inter-frame prediction. The present invention also includes improvements related to inter-frame prediction mechanisms. Inter-frame prediction can be performed in several different modes, which can operate according to unidirectional inter-frame prediction and / or bidirectional inter-frame prediction.

[0143] Figure 6 Schematic diagram of an example of unidirectional inter prediction 600. For example, at the block compression step 105, the block decoding step 113, the motion estimation component 221, the motion compensation component 219, the motion compensation component 321 and / or the motion compensation component 421, the unidirectional inter prediction 600 is performed to determine a motion vector (MV). For example, the unidirectional inter prediction 600 can be used to determine the motion vector of the encoded block and / or the decoded block generated when the image (e.g., the image 500) is segmented.

[0144] The unidirectional inter-frame prediction 600 uses a reference frame 630 including a reference block 631 to predict a current block 611 in a current frame 610. The reference frame 630 may be temporally located after the current frame 610 (e.g., as a next reference frame) as shown in the figure, but in some examples, may also be temporally located before the current frame 610 (e.g., as a previous reference frame). The current frame 610 is an exemplary frame / image that is encoded / decoded at a specific time. The current frame 610 includes an object in the current block 611, which matches an object in the reference block 631 in the reference frame 630. The reference frame 630 is a reference frame used to encode the current frame 610, and the reference block 631 is a block in the reference frame 630, and the object included in this block is also included in the current block 611 in the current frame 610.

[0145] The current block 611 is any decoding unit that is encoded / decoded at a specified time point in the decoding process. The current block 611 can be an entire segmented block or a sub-block when affine inter-frame prediction is used. The current frame 610 is separated from the reference frame 630 by a certain temporal distance (TD) 633. TD 633 represents the amount of time between the current frame 610 and the reference frame 630 in the video sequence, and the measurement unit can be a frame. The prediction information of the current block 611 can refer to the reference frame 630 and / or the reference block 631 through a reference index representing the direction and temporal distance between each frame. During the time period represented by TD 633, the object in the current block 611 moves from one position in the current frame 610 to another position in the reference frame 630 (for example, the position of the reference block 631). For example, the object can move along a trajectory 613, which represents the direction in which the object moves over time. The motion vector 635 describes the direction and magnitude of the object moving along the trajectory 613 within TD 633. Therefore, the encoded motion vector 635 , the reference block 631 , and the residual including the difference between the current block 611 and the reference block 631 provide sufficient information to reconstruct the current block 611 and locate the current block 611 in the current frame 610 .

[0146] Figure 7 Schematic diagram of an example of bidirectional inter-frame prediction 700. For example, at the block compression step 105, the block decoding step 113, the motion estimation component 221, the motion compensation component 219, the motion compensation component 321 and / or the motion compensation component 421, the bidirectional inter-frame prediction 700 is performed to determine the MV. For example, the bidirectional inter-frame prediction 700 can be used to determine the motion vector of the encoded block and / or the decoded block generated when the image (e.g., the image 500) is segmented.

[0147] Bidirectional inter prediction 700 is similar to unidirectional inter prediction 600, but uses a pair of reference frames to predict a current block 711 in a current frame 710. Therefore, the current frame 710 and the current block 711 are substantially similar to the current frame 610 and the current block 611, respectively. The current frame 710 is temporally located between a previous reference frame 720, which occurs before the current frame 710 in the video sequence, and a next reference frame 730, which occurs after the current frame 710 in the video sequence. The previous reference frame 720 and the next reference frame 730 are substantially similar to the reference frame 630 in other respects.

[0148] The current block 711 matches the previous reference block 721 in the previous reference frame 720 and the next reference block 731 in the next reference frame 730. This matching indicates that, during the playback of the video sequence, an object moves along the motion path 713 from a position in the previous reference block 721 through the current block 711 to a position in the next reference block 731. The current frame 710 is separated from the previous reference frame 720 by a certain previous temporal distance (TD0) 723, and is separated from the next reference frame 730 by a certain next temporal distance (TD1) 733. TD0 723 represents the amount of time between the previous reference frame 720 and the current frame 710 in the video sequence, in units of frames. TD1 733 represents the amount of time between the current frame 710 and the next reference frame 730 in the video sequence, in units of frames. Therefore, the object moves along the motion path 713 from the previous reference block 721 to the current block 711 within the time period represented by TD0 723. The object also moves from the current block 711 to the next reference block 731 along the motion path 713 within the time period represented by TD1 733. The prediction information of the current block 711 can refer to the previous reference frame 720 and / or the previous reference block 721 and the next reference frame 730 and / or the next reference block 731 through a pair of reference indexes representing the direction and time distance between the frames.

[0149] The previous motion vector (MV0) 725 describes the direction and magnitude of the object moving along the motion path 713 in TD0 723 (e.g., between the previous reference frame 720 and the current frame 710). The next motion vector (MV1) 735 describes the direction and magnitude of the object moving along the motion path 713 in TD1 733 (e.g., between the current frame 710 and the next reference frame 730). Therefore, in the bidirectional inter-frame prediction 700, the current block 711 can be decoded and reconstructed by the previous reference block 721 and / or the next reference block 731, MV0 725, and MV1 735.

[0150] In the merge mode and the advanced motion vector prediction (AMVP) mode, the candidate list is generated by adding the candidate motion vectors to the candidate list in the order defined by the candidate list determination mode. These candidate motion vectors may include motion vectors according to unidirectional inter-frame prediction 600, bidirectional inter-frame prediction 700, or a combination thereof. Specifically, the motion vectors are generated for the adjacent blocks when they are encoded. These motion vectors are added to the candidate list of the current block, and the motion vector of the current block is selected from the candidate list. The motion vector can then be indicated as the index of the selected motion vector in the candidate list. The decoder can construct the candidate list using the same process as the encoder, and can determine the selected motion vector from the candidate list according to the indicated index. Therefore, the candidate motion vectors include motion vectors generated according to unidirectional inter-frame prediction 600 and / or bidirectional inter-frame prediction 700, depending on which method is used when encoding these adjacent blocks.

[0151] Figure 8 Schematic diagram of an example 800 of decoding a current block 801 based on candidate motion vectors from neighboring decoded blocks 802. The encoder 300 and / or decoder 400 operating the method 100 and / or using the functions of the codec system 200 can use the neighboring decoded blocks 802 to generate a candidate list. Such a candidate list can be used in inter-frame prediction according to unidirectional inter-frame prediction 600 and / or bidirectional inter-frame prediction 700. Then, the candidate list can be used to encode / decode the current block 801, which can be generated by segmenting an image (e.g., image 500).

[0152] The current block 801 is a block encoded at the encoder side or decoded at the decoder side according to an example within a specified time. The decoded block 802 is a block that has been encoded at a specified time. Therefore, the decoded block 802 is likely to be used to generate a candidate list. The current block 801 and the decoded block 802 may be included in the same frame and / or may be included in a temporally adjacent frame. When the decoded block 802 is included in the same frame as the current block 801, the decoded block 802 includes a border that is immediately adjacent to (e.g., adjacent to) the border of the current block 801. When the decoded block 802 is included in a temporally adjacent frame, the position of the decoded block 802 in the temporally adjacent frame is the same as the position of the current block 801 in the current frame. The candidate list may be generated by adding the motion vector from the decoded block 802 as a candidate motion vector. Then, the current block 801 may be decoded by selecting a candidate motion vector from the candidate list and indicating the index of the selected candidate motion vector.

[0153] Fig. 9Schematic diagram of an exemplary mode 900 for determining a motion vector candidate list. Specifically, the encoder 300 and / or the decoder 400 operate the method 100 and / or adopt the function of the codec system 200, and can adopt the candidate list determination mode 900 to generate a candidate list 911, so as to encode the current block 801 segmented from the image 500. The obtained candidate list 911 can be a fusion candidate list or an AMVP candidate list, which can be used in inter-frame prediction according to the unidirectional inter-frame prediction 600 and / or the bidirectional inter-frame prediction 700.

[0154] When encoding the current block 901, the candidate list determination mode 900 searches for valid candidate motion vectors in positions 905 (denoted as A0, A1, B0, B1 and / or B2) in the same image / frame where the current block 901 is located. The candidate list determination mode 900 may also search for valid candidate motion vectors in collocated blocks 909. Collocated blocks 909 are blocks that are located at the same position as the current block 901, but are included in temporally adjacent images / frames. The candidate motion vectors may then be placed in a candidate list 911 in a predetermined checking order. Thus, the candidate list 911 is a list of procedurally generated indexed candidate motion vectors.

[0155] The candidate list 911 can be used to select a motion vector to perform inter-frame prediction of the current block 901. For example, the encoder can obtain samples in the reference block pointed to by the candidate motion vector in the candidate list 911. Then, the encoder can select a candidate motion vector pointing to the reference block that best matches the current block 901. Then, the index of the selected candidate motion vector can be encoded to represent the current block 901. In some cases, one or more candidate motion vectors point to a reference block including a portion of the reference sample 915. In this case, the interpolation filter 913 can be used to reconstruct the complete reference sample 915 to support motion vector selection. The interpolation filter 913 is a filter that can upsample a signal. Specifically, the interpolation filter 913 is a filter that can accept a partial / low-quality signal as input and determine an approximation of a complete / high-quality signal. Therefore, the interpolation filter 913 can be used in some cases to obtain a complete set of reference samples 915 for selecting a reference block for the current block 901, and therefore for selecting a motion vector to encode the current block 901.

[0156] The above-described mechanism for decoding blocks using a candidate list according to inter-frame prediction may result in certain errors when using sub-images (e.g., sub-image 510). Specifically, these problems may occur when the current block 901 is included in the current sub-image, but the motion vector points to a reference block that is at least partially located in an adjacent sub-image. In this case, the current sub-image may be extracted for presentation without the adjacent sub-image. When this occurs, portions of the reference block in the adjacent sub-image may not be sent to the decoder, and therefore the reference block may not be available for decoding the current block 901. When this occurs, the decoder cannot access sufficient data to decode the current block 901.

[0157] The present invention provides a mechanism to solve this problem. In one example, a flag is used to indicate that the current sub-image can be treated as an image. The flag can be set to support the separate extraction of sub-images. Specifically, when the flag is set, the current sub-image can be encoded without reference to data in other sub-images. In this case, the current sub-image is treated as an image because the current sub-image is decoded separately from other sub-images and can be used to display as a separate image. Therefore, the flag can be represented as subpic_treated_as_pic_flag[i], where i represents the index of the current sub-image. When the flag is set, the motion vector candidates (also called motion vector prediction values) obtained from the juxtaposition block 909 only include motion vectors pointing to the inside of the current sub-image. Any motion vector prediction value pointing to the outside of the current sub-image is excluded from the candidate list 911. This ensures that the motion vector pointing to the outside of the current sub-image is not selected, avoiding correlation errors. This example is specifically applied to motion vectors from the juxtaposition block 909. The motion vector from the search position 905 in the same image / frame can be modified by different mechanisms as described below.

[0158] Another example may be used to address the search position 905 when the current sub-image is treated as an image (e.g., when subpic_treated_as_pic_flag[i] is set). When the current sub-image is treated as an image, the current sub-image may be extracted without reference to other sub-images. An example mechanism involves an interpolation filter 913. The interpolation filter 913 may be applied to samples at one position to interpolate (e.g., predict) related samples at another position. In this example, as long as the interpolation filter 913 can interpolate reference samples 915 outside the current sub-image only based on reference samples 915 in the current sub-image, the motion vector from the decoded block located at the search position 905 may point to the reference sample 915. Therefore, this example employs a clipping function used when applying the interpolation filter 913 to the motion vector candidate from the search position 905 in the same image. In determining the reference sample 915 pointed to by the motion vector candidate, the clipping function clips the data in the adjacent sub-image and thus removes the data input to the interpolation filter 913. This approach keeps sub-images separate during encoding to support separate extraction and decoding when the sub-images are treated as images. The clipping function may be applied to the luma sample bilinear interpolation process, the luma sample 8-tap interpolation filter process, and / or the chroma sample interpolation process.

[0159] Fig.10 1 is a block diagram of an exemplary in-loop filter 1000. The in-loop filter 1000 can be used to implement the in-loop filter components 225, 325 and / or 425. In addition, when performing method 100, the in-loop filter 1000 can be applied to the encoder and decoder sides. In addition, the in-loop filter 1000 can be used to filter the current block 801 segmented from the image 500, and the current block 801 can be decoded according to the candidate list generated according to the mode 900 according to the unidirectional inter-frame prediction 600 and / or the bidirectional inter-frame prediction 700. The in-loop filter 1000 includes a deblocking filter 1043, a SAO filter 1045 and an adaptive loop filter (adaptive loop filter, ALF) 1047. The filters in the in-loop filter 1000 are sequentially applied to the reconstructed image blocks at the encoder side (e.g., before being used as a reference block) and the decoder side (before display).

[0160] The deblocking filter 1043 is used to remove blocky edges resulting from block-based inter- and intra-frame prediction. The deblocking filter 1043 scans image portions (e.g., image strips) to determine discontinuities in chrominance values ​​and / or luminance values ​​that occur at segmentation boundaries. The deblocking filter 1043 then applies a smoothing function to the block boundaries to remove these discontinuities. The strength of the deblocking filter 1043 can vary depending on spatial activity (e.g., changes in luminance / chrominance components) occurring in areas adjacent to block boundaries.

[0161] The SAO filter 1045 is used to remove artifacts related to sample distortion caused by the encoding process. The SAO filter 1045 on the encoder side classifies the deblocking filtered samples in the reconstructed image into several categories based on the associated deblocking filtered edge shape and / or direction. An offset is then determined based on the category and added to the sample. The offset is then encoded into the bitstream and used by the SAO filter 1045 on the decoder side. The SAO filter 1045 removes banding artifacts (bands of values ​​instead of smooth transitions) and ringing artifacts (noisy signals near sharp edges).

[0162] The ALF 1047 on the encoder side is used to compare the reconstructed image with the original image. The ALF 1047 determines the coefficients describing the difference between the reconstructed image and the original image through a Wiener-based adaptive filter, etc. These coefficients are encoded into the bitstream and used by the ALF 1047 on the decoder side to remove the difference between the reconstructed image and the original image.

[0163] The image data filtered by the in-loop filter 1000 is output to the image buffer 1023, which is substantially similar to the decoded image buffer components 223, 323 and / or 423. As described above, the deblocking filter 1043, the SAO filter 1045 and / or the ALF 1047 can be disabled at sub-picture boundaries and / or tile boundaries via flags such as loop_filter_across_subpic_enabled_flag and / or loop_filter_across_tiles_enabled_flag.

[0164] Fig.11Schematic diagram of an exemplary bitstream 1100 including decoding tool parameters to support decoding of a sub-image in an image. For example, the bitstream 1100 may be generated by the codec system 200 and / or the encoder 300 to be decoded by the codec system 200 and / or the decoder 400. For another example, the bitstream 1100 may be generated by the encoder in step 109 of the method 100 to be used by the decoder in step 111. In addition, the bitstream 1100 may include the encoded image 500, the corresponding sub-image 510 and / or the related decoded blocks, such as the current block 801 and / or 901, which may be decoded according to the unidirectional inter-frame prediction 600 and / or the bidirectional inter-frame prediction 700 using the candidate list generated according to the mode 900. The bitstream 1100 may also include parameters for configuring the in-loop filter 1000.

[0165] The code stream 1100 includes a sequence parameter set (SPS) 1110, a plurality of picture parameter sets (PPS) 1111, a plurality of slice headers 1115, and image data 1120. The SPS 1110 includes sequence data common to all images in the video sequence included in the code stream 1100. Such data may include image size, bit depth, decoding tool parameters, bit rate constraints, etc. The PPS 1111 includes parameters applied to the entire image. Therefore, each image in the video sequence may refer to the PPS 1111. It should be noted that, although each image refers to the PPS 1111, in some examples, a single PPS 1111 may include data for a plurality of images. For example, a plurality of similar images may be decoded according to similar parameters. In this case, a single PPS 1111 may include data for such similar images. The PPS 1111 may indicate decoding tools, quantization parameters, offsets, etc. that may be used for the slices in the corresponding image. The slice header 1115 includes parameters specific to each slice in the image. Therefore, each slice in the video sequence may have a slice header 1115. The slice header 1115 may include slice type information, picture order count (POC), reference picture list, prediction weight, partition entry point, deblocking filter parameters, etc. It should be noted that in some contexts, the slice header 1115 may also be called a partition group header.

[0166] The image data 1120 includes video data encoded according to inter-frame prediction and / or intra-frame prediction and corresponding transformed quantized residual data. For example, a video sequence includes multiple images decoded as image data. An image is a single frame in a video sequence, so it is usually displayed as a single unit when the video sequence is displayed. However, sub-images can be used for display to implement certain technologies, such as virtual reality, picture-in-picture, etc. The images are all referenced to PPS1111. As described above, the image is divided into sub-images, blocks and / or stripes. In some systems, a strip is referred to as a block group including a block. The strip and / or the block group including the block refer to the strip header 1115. These strips are further divided into CTUs and / or CTBs. CTU / CTBs are further divided into decoding blocks according to the coding tree. The decoding blocks can then be encoded / decoded according to the prediction mechanism.

[0167] The parameter set in the codestream 1100 includes various data that can be used to implement the examples described herein. To support the first exemplary implementation, the SPS 1110 in the codestream 1100 includes a sub-pic treated as a pic flag 1131 associated with a specified sub-picture. In some examples, the sub-picture treated as a pic flag 1131 is represented as subpic_treated_as_pic_flag[i], where i represents the index of the sub-picture associated with the flag. For example, the sub-picture treated as a pic flag 1131 can be set to 1, indicating that the i-th sub-picture of each decoded picture in the encoded video sequence (in the picture data 1120) is treated as a picture during the decoding process except for the in-loop filtering operation. The sub-picture treated as a picture flag 1131 can be used when the current sub-picture in the current picture has been decoded according to inter-frame prediction. When the sub-image is processed as an image flag 1131 is set to indicate that the current sub-image is processed as an image, the candidate list of candidate motion vectors of the current block can be determined by excluding the collocated motion vectors included in the collocated block and pointing to the outside of the current sub-image from the candidate list. This ensures that when the current sub-image is extracted separately from other sub-images, the motion vector pointing to the outside of the current sub-image is not selected, avoiding correlation errors.

[0168] In some examples, the candidate list of motion vectors for the current block is determined based on temporal luma motion vector prediction. For example, temporal luma motion vector prediction may be used in the following situations: the current block is a luma block composed of luma samples, the selected current motion vector of the current block is a temporal luma motion vector pointing to a reference luma sample in a reference block, and the current block is decoded based on the reference luma sample. In this case, the temporal luma motion vector prediction is performed according to the following:

[0169] xColBr = xCb + cbWidth;

[0170] yColBr=yCb+cbHeight;

[0171] rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]?

[0172] SubPicRightBoundaryPos:pic_width_in_luma_samples–1;

[0173] botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]?

[0174] SubPicBotBoundaryPos:pic_height_in_luma_samples–1,

[0175] Where xColBr and yColBR represent the position of the collocated block, xCb and yCb represent the upper left sample of the current block relative to the upper left sample of the current image, cbWidth represents the width of the current block, cbHeight represents the height of the current block, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicBotBoundaryPos represents the position of the bottom boundary of the sub-image, pic_width_in_luma_samples represents the width of the current image measured in luma samples, pic_height_in_luma_samples represents the height of the current image measured in luma samples, botBoundaryPos represents the calculated position of the bottom boundary of the sub-image, rightBoundaryPos represents the calculated position of the right boundary of the sub-image, SubPicIdx represents the index of the sub-image; when yCb>>CtbLog2SizeY is not equal to yColBr>>CtbLog2SizeY, the collocated motion vector is excluded, where CtbLog2SizeY represents the size of the coding tree block.

[0176] The sub-image treated as image processing flag 1131 can also be used in the second exemplary implementation. As in the first example, the sub-image treated as image processing flag 1131 can be used in the case where the current sub-image in the current image has been decoded according to inter-frame prediction. In this example, the motion vector can be determined for the current block in the sub-image (for example, from a candidate list). When the sub-image treated as image processing flag 1131 is set, a limiting function can be applied to the sample position in the reference block. The sample position is a position in the image and can include a single sample consisting of a luminance value and / or a pair of chrominance values. Then, an interpolation filter can be applied in the case where the motion vector points outside the current sub-image. This limiting function ensures that the interpolation filter does not rely on data in adjacent sub-images to keep the sub-images separate, thereby supporting separate extraction.

[0177] The clipping function can be applied to the luma sample bilinear interpolation process. The luma sample bilinear interpolation process can receive an input including a luma position (xIntL, yIntL) with an integer number of sample units. The luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). The clipping function is applied to the sample position as described below. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0178] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i),

[0179] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i),

[0180] Wherein, subpic_treated_as_pic_flag indicates a flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx indicates the index of the sub-image, xInti and yInti indicate the position of the clipped sample at index i, SubPicRightBoundaryPos indicates the position of the right boundary of the sub-image, SubPicLeftBoundaryPos indicates the position of the left boundary of the sub-image, SubPicTopBoundaryPos indicates the position of the upper boundary of the sub-image, SubPicBotBoundaryPos indicates the position of the lower boundary of the sub-image, and Clip3 indicates the clipping function according to the following formula:

[0181]

[0182] Where x, y, and z are numeric input values.

[0183] The clipping function can also be applied to the luma sample 8-tap interpolation filtering process. The luma sample 8-tap interpolation filtering process receives an input including a luma position (xIntL, yIntL) with an integer number of sample units. The luma sample 8-tap interpolation filtering process outputs a predicted luma sample value (predSampleLXL). The clipping function is applied to the sample position as described below. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0184] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i–3),

[0185] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i–3),

[0186] Among them, subpic_treated_as_pic_flag represents a flag set to indicate that the sub-image is treated as an image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the limited sample at index i, SubPicRightBoundaryPos represents the position of the right border of the sub-image, SubPicLeftBoundaryPos represents the position of the left border of the sub-image, SubPicTopBoundaryPos represents the position of the upper border of the sub-image, SubPicBotBoundaryPos represents the position of the lower border of the sub-image, and Clip3 is as described above.

[0187] The clipping function can also be applied in the chroma sample interpolation process. The chroma sample interpolation process receives an input including a chroma position (xIntC, yIntC) with an integer number of sample units. The chroma sample interpolation process outputs a predicted chroma sample value (predSampleLXC). The clipping function is applied to the sample positions as described below. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0188] xInti=Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xIntC+i),

[0189] yInti=Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yIntC+i),

[0190] Among them, subpic_treated_as_pic_flag represents a flag set to indicate that the sub-image is treated as an image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the limited sample at index i, SubPicRightBoundaryPos represents the position of the right border of the sub-image, SubPicLeftBoundaryPos represents the position of the left border of the sub-image, SubPicTopBoundaryPos represents the position of the upper border of the sub-image, SubPicBotBoundaryPos represents the position of the lower border of the sub-image, SubWidthC and SubHeightC represent the horizontal sampling rate ratio and the vertical sampling rate ratio between the luminance samples and the chrominance samples, and Clip3 is as described above.

[0191] The cross-sub-image loop filter enable flag 1132 in the SPS 1110 can be used for the third exemplary implementation. The cross-sub-image loop filter enable flag 1132 can be set to control whether filtering is applied across the boundary of a specified sub-image. For example, the cross-sub-image loop filter enable flag 1132 can be represented as loop_filter_across_subpic_enabled_flag. The cross-sub-image loop filter enable flag 1132 can be set to 1 when indicating that the in-loop filtering operation can be performed across the boundary of the sub-image, or can be set to 0 when indicating that the in-loop filtering operation is not performed across the boundary of the sub-image. Therefore, depending on the value of the cross-sub-image loop filter enable flag 1132, the filtering process can be performed across the boundary of the sub-image or not across the boundary of the sub-image. The filtering operation can include applying a deblocking filter 1043, an ALF 1047, and / or an SAO filter 1045. In this way, the filter can be disabled for sub-images encoded for individual extraction and enabled for sub-images encoded for group display.

[0192] The cross-block loop filter enable flag 1134 in PPS1111 can be used for the fourth exemplary implementation. The cross-block loop filter enable flag 1134 can be set to control whether filtering is used across the boundary of a specified block. For example, the cross-block loop filter enable flag 1134 can be represented as loop_filter_across_tiles_enabled_flag. The cross-block loop filter enable flag 1134 can be set to 1 when indicating that the in-loop filtering operation can be performed across the boundary of the block, or it can be set to 0 when indicating that the in-loop filtering operation is not performed across the boundary of the block. Therefore, according to the value of the cross-block loop filter enable flag 1134, the filtering process can be performed across the boundary of the block or not. The filtering operation may include applying a deblocking filter 1043, an ALF 1047, and / or an SAO filter 1045.

[0193] The sub-image data 1133 in the SPS 1110 may be used for the fifth exemplary implementation. The sub-image data 1133 may include the width, height, and offset of each sub-image in the image data 1120. For example, the width and height of each sub-image may be described in units of CTB in the sub-image data 1133. In some examples, the width and height of the sub-image are stored in the sub-image data 1133 as subpic_width_minus1 and subpic_height_minus1, respectively. In addition, the offset of each sub-image may be described in units of CTU in the sub-image data 1133. For example, the offset of each sub-image may be expressed as the vertical position and horizontal position of the upper left CTU in the sub-image. Specifically, the offset of the sub-image may be expressed as the difference between the upper left CTU in the image and the upper left CTU in the sub-image. In some examples, the vertical position and horizontal position of the upper left CTU in the sub-image are stored in the sub-image data 1133 as subpic_ctu_top_left_y and subpic_ctu_top_left_x, respectively. This exemplary implementation describes the sub-image in the sub-image data 1133 in units of CTB / CTU rather than in units of blocks. In this way, the sub-image can also be used in the case where the blocks are not used in the corresponding image / sub-image.

[0194] The sub-image data 1133 in the SPS 1110, the slice address 1136 in the slice header 1115, and the slice sub-image ID 1135 in the slice header 1115 may be used for the sixth exemplary implementation. The sub-image data 1133 may be implemented as described in the fifth exemplary implementation. The slice address 1136 may include a sub-image level slice index of a slice (e.g., in the image data 1120) associated with the slice header 1115. For example, the slice is indexed according to the position of the slice in the sub-image rather than according to the position of the slice in the image. The slice address 1136 may be stored in a variable slice_address. The slice sub-image ID 1135 includes an ID of a sub-image that includes the slice associated with the slice header 1115. Specifically, the slice sub-image ID 1135 may refer to a description (e.g., width, height, and offset) of a corresponding sub-image in the sub-image data 1133. The slice sub-image ID 1135 may be stored in the variable slice_subpic_id. Thus, the slice address 1136 is indicated as an index according to the position of the slice in the sub-image indicated by the slice sub-image ID 1135 as described in the sub-image data 1133. In this way, the position of the slice in the sub-image may also be determined when the sub-image is extracted individually and the other sub-images are omitted from the codestream 1100. This is because this addressing scheme separates the address of each sub-image from the addresses of the other sub-images. Therefore, the slice header 1115 does not need to be rewritten when the sub-image is extracted, which would be necessary in an addressing scheme where the slice is addressed according to the position of the slice in the image. It should be noted that this approach may be used when the slice is a rectangular / square strip (as opposed to a raster scan strip). For example, the rect_slice_flag in the PPS 1111 may be set to 1 to indicate that the slice is a rectangular strip.

[0195] An exemplary implementation of a sub-image used in some video decoding systems is described below. Information related to sub-images that may exist in a CVS may be indicated in an SPS. Such an indication may include the following information. The number of sub-images present in each image of a CVS may be included in an SPS. In the context of an SPS or a CVS, the collocated sub-images of all access units (AUs) may be collectively referred to as a sub-image sequence. A loop for further indicating information related to the attributes of each sub-image may also be included in an SPS. This information may include a sub-image identifier, a sub-image position (e.g., an offset distance between the upper left corner brightness sample in a sub-image and the upper left corner brightness sample in an image), and a size of the sub-image. In addition, an SPS may also be used to indicate whether each sub-image is a motion-constrained sub-image, wherein a motion-constrained sub-image is a sub-image that includes an MCTS. The profile, tier, and level information of each sub-image may be included in the bitstream, unless such information is derivable. This information may be used to extract the profile, tier, and level information of the extracted bitstream generated by extracting the sub-image from the original bitstream that includes the entire image. The level and layer of each sub-image can be derived to be the same as the level and layer of the original code stream. The level of each sub-image can be explicitly indicated. This indication can appear in the above loop. Sequence-level hypothetical reference decoder (HRD) parameters can be indicated in the video usability information (VUI) part of the SPS for each sub-image (or each sub-image sequence).

[0196] When an image is not divided into two or more sub-images, the attributes of the sub-image (e.g., position, size, etc.), except for the sub-image ID, may not be indicated in the codestream. When extracting a sub-image in an image of a CVS, each access unit in the new codestream may not include the sub-image because the image data generated in each AU in the new codestream is not divided into multiple sub-images. Therefore, sub-image attributes such as position and size can be omitted from the SPS because this information can be derived from the image attributes. However, the sub-image identification can still be indicated because this ID can be referenced by the video coding layer (VCL) NAL unit / block group included in the extracted sub-image. When extracting sub-images, it is necessary to avoid changing the sub-image ID to reduce resource usage.

[0197] For example, the position of the sub-image in the image (x offset and y offset) can be indicated in units of luma samples and can represent the distance between the upper left luma sample in the sub-image and the upper left luma sample in the image. For another example, the position of the sub-image in the image can be indicated in units of the minimum decoded luma block size (MinCbSizeY) and can represent the distance between the upper left luma sample in the sub-image and the upper left luma sample in the image. For another example, the unit of the sub-image position offset can be explicitly indicated by a syntax element in a parameter set, which can be CtbSizeY, MinCbSizeY, luma samples, or other values. The codec may require that when the right boundary of the sub-image does not coincide with the right boundary of the image, the width of the sub-image should be an integer multiple of the luma CTU size (CtbSizeY). Similarly, the codec may also require that when the lower boundary of the sub-image does not coincide with the lower boundary of the image, the height of the sub-image should be an integer multiple of CtbSizeY. The codec may also require that the sub-image be located at the rightmost position of the image when the width of the sub-image is not an integer multiple of the luma CTU size. Similarly, the codec may also require that the sub-image be located at the bottommost position of the image when the height of the sub-image is not an integer multiple of the luma CTU size. When the width of the sub-image is indicated in units of luma CTU size, and the width of the sub-image is not an integer multiple of the luma CTU size, the actual width in units of luma samples may be derived from the offset position of the sub-image, the sub-image width in units of luma CTU size, and the image width in units of luma samples. Similarly, when the height of the sub-image is indicated in units of luma CTU size, and the height of the sub-image is not an integer multiple of the luma CTU size, the actual height in units of luma samples may be derived from the offset position of the sub-image, the sub-image height in units of luma CTU size, and the image height in units of luma samples.

[0198] For any sub-image, the sub-image ID may be different from the sub-image index. The sub-image index may be an index indicating a sub-image in a sub-image loop in an SPS. Alternatively, the sub-image index may be an index assigned relative to an image in a sub-image raster scan order. When the value of the sub-image ID of each sub-image is the same as its sub-image index, the sub-image ID may be indicated or derived. When the sub-image ID of each sub-image is different from its sub-image index, the sub-image ID is explicitly indicated. The number of bits used to indicate the sub-image ID may be indicated in the same parameter set that includes the sub-image attributes (e.g., in an SPS). For certain purposes, some values ​​of the sub-image ID may be reserved. Such value reservation may be as described below. When a block group / strip header includes a sub-image ID to indicate which sub-image includes the block group, a value of 0 may be reserved and may not be used for the sub-image to ensure that the first few bits at the beginning of the block group / strip header are not all 0, thereby avoiding the generation of an anti-counterfeiting code. When the sub-images in an image do not cover the entire area of ​​the image without overlap and gaps, a value (e.g., a value of 1) may be reserved for the group of blocks that do not belong to any sub-image. Optionally, the sub-image IDs of the remaining area may be explicitly indicated. The number of bits used to indicate the sub-image ID may be constrained as follows. The range of values ​​needs to be sufficient to uniquely identify all sub-images in the image, including the reserved values ​​for the sub-image ID. For example, the minimum number of bits used for the sub-image ID may be a value of Ceil(Log2(number of sub-images in the image + number of reserved sub-image IDs).

[0199] The merging of sub-images in a loop may need to cover the entire image without gaps and overlaps. When such constraints are applied, a flag is present for each sub-image indicating whether the sub-image is a motion constrained sub-image, which indicates that the sub-image can be extracted. Optionally, the merging of sub-images may not cover the entire image. However, there may be no overlap between sub-images in the image.

[0200] The sub-image ID may be present immediately after the NAL unit header to assist the sub-image extraction process, so the extractor does not need to know the rest of the NAL unit bits. For VCL NAL units, the sub-image ID may be present in the first few bits of the tile group header. For non-VCL NAL units, the following may apply. The sub-image ID may not need to be immediately following the NAL unit header of the SPS. Regarding the PPS, when all tile groups in the same image are constrained to reference the same PPS, the sub-image ID does not need to be immediately following the NAL unit header. On the other hand, if the tile groups in the same image can reference different PPSs, the sub-image ID may be present in the first few bits of the PPS (e.g., immediately following the PPS NAL unit header). In this case, any two tile groups in an image cannot share the same PPS. Optionally, when the tile groups in the same image can reference different PPSs, and different tile groups in the same image can also share the same PPS, the sub-image ID does not exist in the PPS syntax. Optionally, when different PPSs can be referenced by groups of tiles in the same image, and different groups of tiles in the same image can also share the same PPS, a sub-image ID list is present in the PPS syntax. The list indicates the sub-images to which the PPS applies. For other non-VCL NAL units, if the non-VCL unit is applicable at the picture level (e.g., access unit delimiter, end of sequence, end of codestream, etc.) or above, the sub-image ID does not need to follow its NAL unit header. Otherwise, the sub-image ID can follow the NAL unit header.

[0201] The block segmentation within each sub-image can be indicated in the PPS, but groups of blocks in the same image can refer to different PPSs. In this case, the blocks are grouped in each sub-image, rather than across images. Therefore, the concept of block grouping in this case includes the way in which sub-images are divided into blocks. Optionally, a Sub-Picture Parameter Set (SPPS) can be used to describe the block segmentation within each sub-image. The SPPS refers to the SPS by using syntax elements that reference the SPS ID. The SPPS may include a sub-image ID. For the purpose of sub-image extraction, the syntax element that references the sub-image ID is the first syntax element in the SPPS. The SPPS includes a block structure that indicates the number of columns, rows, uniform block spacing, etc. The SPPS may include a flag to indicate whether the loop filter is enabled across the relevant sub-image boundary. Optionally, the sub-image attributes of each sub-image can be indicated in the SPPS instead of in the SPS. The block segmentation within each sub-image can be indicated in the PPS, but groups of blocks in the same image can refer to different PPSs. Once activated, the SPPS can last for a series of consecutive AUs in decoding order, but can be deactivated / activated at an AU that does not start with a CVS. During the decoding process of a single-layer codestream including multiple sub-images, multiple SPPSs can be active at any time, and the SPPS can be shared by different sub-images of an AU. Optionally, the SPPS and PPS can be combined into one parameter set. To achieve this, all block groups included in the same sub-image may be constrained to refer to the same parameter set obtained by combining the SPPS and PPS.

[0202] The number of bits used to indicate the sub-picture ID can be indicated in the NAL unit header. The presence of this information helps the sub-picture extraction process to parse the sub-picture ID value at the beginning of the NAL unit payload (for example, the first few bits immediately following the NAL unit header). For this indication, some reserved bits in the NAL unit header can be used to avoid increasing the length of the NAL unit header. The number of bits used for such an indication can override the value of sub-picture-ID-bit-len. For example, 4 of the 7 reserved bits in the VVC NAL unit header can be used for this purpose.

[0203] When decoding a sub-image, the position of each coding tree block, expressed as a vertical CTB position (xCtb) and a horizontal CTB position (yCtb), is adjusted to the actual luma sample position in the image, rather than the luma sample position in the sub-image. In this way, since everything is decoded as if it were in the image rather than in the sub-image, the extraction of the collocated sub-image from each reference image can be avoided. In order to adjust the position of the coding tree blocks, the variables SubpictureXOffset and SubpictureYOffset can be derived from the sub-image position (subpic_x_offset and subpic_y_offset). The values ​​of these two variables can be added to the values ​​of the x and y coordinates of the luma sample position of each coding tree block in the sub-image, respectively. The sub-image extraction process can be defined as follows. The input to the process includes the target sub-image to be extracted. The target sub-image to be extracted can be input in the form of a sub-image ID or a sub-image position. When the input is the position of the sub-image, the relevant sub-image ID can be parsed out by parsing the sub-image information in the SPS. For non-VCL NAL units, the following applies. Syntax elements in the SPS related to picture size and level may be updated with sub-picture size and level information. The following non-VCL NAL units are not changed by extraction: PPS, access unit delimiter (AUD), end of sequence (EOS), end of bitstream (EOB), and any other non-VCL NAL units applicable at picture level or above. Remaining non-VCL NAL units whose sub-picture ID is not equal to the target sub-picture ID may be removed. VCL NAL units whose sub-picture ID is not equal to the target sub-picture ID may also be removed.

[0204] The sub-image nesting SEI message can be used for AU-level or sub-image-level SEI messages that nest sub-image sets. The data carried by the sub-image nesting SEI message may include buffering periods, image timing, and non-HRD SEI messages. The syntax and semantics of this SEI message may be described as follows. For system operation, for example, an omnidirectional mediaformat (OMAF) environment, a sub-image sequence set covering a perspective can be requested and decoded by an OMAF player. Therefore, a sequence-level SEI message may carry information about a sub-image sequence set that uniformly includes a rectangular or square image area. This information can be used by the system, which indicates the minimum decoding capability and the bitrate of the sub-image sequence set. The information includes the level of the codestream of only the sub-image sequence set, the bitrate of the codestream, and optionally includes a sub-codestream extraction process specified for the sub-image sequence set.

[0205] There are several problems with the above implementation. The indication of the width and height of the image and / or the width / height / offset of the sub-image is not efficient. Indicating this information can save more bits. When the sub-image size and position information is indicated in the SPS, the PPS includes a block configuration. In addition, the PPS can be shared by multiple sub-images in the same image. Therefore, the value range of num_tile_columns_minus1 and num_tile_rows_minus1 can be more clearly stated. In addition, the semantics of the flag indicating whether the sub-image is a motion constrained sub-image are not clearly stated. Each sub-image sequence must indicate the level. However, when the sub-image sequence cannot be decoded independently, indicating the level of the sub-image is useless. In addition, in some applications, some sub-image sequences can be decoded and displayed together with at least one other sub-image sequence. Therefore, indicating the level of a single sub-image sequence in these sub-image sequences may be useless. In addition, determining the level value of each sub-image may burden the encoder.

[0206] With the introduction of independently decodable sub-image sequences, schemes that require independent extraction and decoding of certain areas in an image may not be implemented based on tile groups. Therefore, explicit indication of the tile group ID may not be useful. In addition, the values ​​of each syntax element in the PPS syntax elements pps_seq_parameter_set_id and loop_filter_across_tiles_enabled_flag may be the same in all PPSs referenced by the tile group header of the decoded image. This is because the active SPS cannot be changed in the CVS, and the value of loop_filter_across_tiles_enabled_flag can be the same for all tiles in the image used for parallel processing based on the tiles. Whether rectangular tile groups and raster scan tile groups can be mixed in an image needs to be clearly stated. Whether sub-images belonging to different images and using the same sub-image ID in the CVS can use different tile group modes also needs to be stated. The derivation process of temporal luminance motion vector prediction may not be able to treat sub-image boundaries as image boundaries in temporal motion vector prediction (TMVP). In addition, the luma sample bilinear interpolation process, the luma sample 8-tap interpolation filtering process, and the chroma sample interpolation process may not treat the sub-image boundary as an image boundary in motion compensation. In addition, the mechanism for controlling the deblocking filter operation, SAO filter operation, and ALF filter operation at the sub-image boundary also needs to be explained.

[0207] With the introduction of independently decodable sub-image sequences, loop_filter_across_tile_groups_enabled_flag may be less useful. This is because disabling in-loop filtering operations for parallel processing purposes can also be satisfied by setting loop_filter_across_tile_groups_enabled_flag to 0. In addition, disabling in-loop filtering operations to enable independent extraction and decoding of certain areas in the image can also be satisfied by setting loop_filter_across_sub_pic_enabled_flag to 0. Therefore, further explanation of the process of disabling in-loop filtering operations across tile group boundaries according to loop_filter_across_tile_groups_enabled_flag will unnecessarily burden the decoder and waste bits. In addition, the decoding process as described above may not be able to disable ALF filtering operations across tile boundaries.

[0208] Therefore, the present invention includes designs that support sub-image based video decoding. A sub-image is a rectangular or square region within an image that may or may not be independently decoded using the same decoding process as the image. The description of these techniques is based on the Versatile Video Coding (VVC) standard. However, these techniques are also applicable to other video codec specifications.

[0209] In some examples, size units are indicated for the image width and height syntax elements and the list of syntax elements sub-image width / height / offset_x / offset_y. All syntax elements are indicated in the form of xxx_minus1. For example, when the size unit is 64 luma samples, a width value of 99 indicates that the image width is 6400 luma samples. The same example applies to other elements in these syntax elements. As another example, one or more of the following size units apply. Size units may be indicated for the syntax elements image width and height in the form of xxx_minus1. The size units indicated for the list of syntax elements sub-image width / height / offset_x / offset_y may be in the form of xxx_minus1. As another example, one or more of the following size units apply. The size units of the list of syntax elements image width and sub-image width / offset_x may be indicated in the form of xxx_minus1. The size units of the list of syntax elements image height and sub-image height / offset_y may be indicated in the form of xxx_minus1. As another example, one or more of the following size units apply. The grammatical elements image width and height in the form of xxx_minus1 can be indicated in units of minimum coding units. The grammatical elements sub-image width / height / offset_x / offset_y in the form of xxx_minus1 can be indicated in units of CTU or CTB. The sub-image width of each sub-image at the right border of the image can be derived. The sub-image height of each sub-image at the bottom border of the image can be derived. Other values ​​of sub-image width / height / offset_x / offset_y can be indicated in the code stream. In other examples, for cases where sub-images have uniform sizes, a pattern for indicating the width and height of sub-images and their positions in the image can be added. When a sub-image includes the same sub-image rows and sub-image columns, the sub-image has a uniform size. In this mode, the number of sub-image rows, the number of sub-image columns, the width of each sub-image column, and the height of each sub-image row can all be indicated.

[0210] As another example, the indication of the sub-image width and height may not be included in the PPS. The values ​​of num_tile_columns_minus1 and num_tile_rows_minus1 may range from 0 to an integer value, such as 1024, inclusive. As another example, when a sub-image of the reference PPS includes more than one partition, two syntax elements adjusted by the presence flag may be indicated in the PPS. These syntax elements are used to indicate the sub-image width and height in units of CTBs and represent the size of all sub-images of the reference PPS.

[0211] As another example, additional information describing each sub-image may also be indicated. A flag such as sub_pic_treated_as_pic_flag[i] may be indicated for each sub-image sequence to indicate whether the sub-images in the sub-image sequence are treated as images during decoding for purposes other than in-loop filtering operations. When sub_pic_treated_as_pic_flag[i] is equal to 1, only the level to which the sub-image sequence conforms may be indicated. A sub-image sequence is a CVS consisting of sub-images with the same sub-image ID. When sub_pic_treated_as_pic_flag[i] is equal to 1, the level of the sub-image sequence may also be indicated. This may be controlled by a flag for all sub-image sequences or by a flag for each sub-image sequence. As another example, sub-codestream extraction may be achieved without changing the VCL NAL unit. This may be achieved by removing the explicit tile group ID indication from the PPS. When rect_tile_group_flag is equal to 1, indicating a rectangular tile group, the semantics of tile_group_address are specified. Tile_group_address may include a tile group index of a tile group among a plurality of tile groups within a sub-image.

[0212] As another example, the value of each of the PPS syntax elements pps_seq_parameter_set_id and loop_filter_across_tiles_enabled_flag may be the same in all PPSs referenced by the partition group headers of a coded image. Other PPS syntax elements may be different for different PPSs referenced by the partition group headers of a coded image. The value of single_tile_in_pic_flag may be different for different PPSs referenced by the partition group headers of a coded image. Thus, some images in a CVS may include only one tile, while other images in the CVS may include more than one tile. Thus, some sub-images in an image (e.g., very large sub-images) may include multiple tiles, while other sub-images in the same image (e.g., very small sub-images) may include only one tile.

[0213] For another example, an image may include a mixture of rectangular scan blocks and raster scan block groups. Therefore, some sub-images in the image use rectangular block group mode, while other sub-images use raster scan block group mode. This flexibility is beneficial to code stream fusion scenarios. Optionally, constraints may require that all sub-images in the image should use the same block group mode. Sub-images in different images with the same sub-image ID in CVS may not use different block group modes. Sub-images in different images with the same sub-image ID in CVS may use different block group modes.

[0214] As another example, when sub_pic_treated_as_pic_flag[i] is equal to 1 for a sub-picture, the collocated motion vectors used for temporal motion vector prediction of the sub-picture are constrained to be from within the boundaries of the sub-picture. Thus, the temporal motion vector prediction of the sub-picture is processed as if the sub-picture boundaries were picture boundaries. In addition, a clipping operation is specified as part of the luma sample bilinear interpolation process, the luma sample 8-tap interpolation filter process, and the chroma sample interpolation process to treat sub-picture boundaries as picture boundaries in motion compensation for sub-pictures with sub_pic_treated_as_pic_flag[i] equal to 1.

[0215] As another example, each sub-image is associated with an indicator flag such as loop_filter_across_sub_pic_enabled_flag. This flag is used to control the in-loop filtering operation at the sub-image boundary and is used to control the filtering operation in the corresponding decoding process. The deblocking filter process may not be used to decode sub-block edges and transform block edges that coincide with the boundary of the sub-image for which loop_filter_across_sub_pic_enabled_flag is equal to 0. Optionally, the deblocking filter process is not used to decode sub-block edges and transform block edges that coincide with the upper boundary or left boundary of the sub-image for which loop_filter_across_sub_pic_enabled_flag is equal to 0. Optionally, the deblocking filter process is not used to decode sub-block edges and transform block edges that coincide with the boundary of the sub-image for which sub_pic_treated_as_pic_flag[i] is equal to 1 or zero. Optionally, the deblocking filter process is not used to decode sub-block edges and transform block edges that coincide with the upper boundary or left boundary of the sub-image. When the loop_filter_across_sub_pic_enabled_flag of a sub-image is equal to zero, a clipping operation may be specified to disable SAO filtering operations across sub-image boundaries. When the loop_filter_across_sub_pic_enabled_flag of a sub-image is equal to 0, a clipping operation may be specified to disable ALF filtering operations across sub-image boundaries. The loop_filter_across_tile_groups_enabled_flag may also be removed from the PPS. Therefore, when loop_filter_across_tiles_enabled_flag is equal to 0, in-loop filtering operations across tile group boundaries that are not sub-image boundaries are not disabled. The loop filter operations may include deblocking filtering operations, SAO filtering operations, and ALF filtering operations. For example, when the loop_filter_across_tiles_enabled_flag of a tile is equal to 0, a clipping operation is specified to disable ALF filtering operations across tile boundaries.

[0216] One or more of the above examples may be implemented as follows. A sub-image may be defined as a rectangular or square region consisting of one or more block groups or strips in an image. The following divisions of processing elements may form a spatial or component segmentation: each image is divided into components, each component is divided into CTBs, each image is divided into sub-images, each sub-image is divided into block columns in a sub-image, each sub-image is divided into block rows in a sub-image, each block column in a sub-image is divided into blocks, each block row in a sub-image is divided into blocks, and each sub-image is divided into block groups.

[0217] The CTB raster and tile scanning process within a sub-image can be described as follows: List ColWidth[i] (i ranges from 0 to num_tile_columns_minus1, including the end value) represents the width of the i-th tile column in CTB units, which can be derived as follows.

[0218]

[0219] The list RowHeight[j] (the value range of j is 0 to num_tile_rows_minus1, including the end value) represents the height of the j-th tile row in CTB units, which is derived as follows.

[0220]

[0221] The list ColBd[i] (the value range of i is 0 to num_tile_columns_minus1+1, including the end value) represents the position of the i-th tile column boundary in units of CTB, which is derived as follows.

[0222] for(ColBd[0]=0,i=0;i<=num_tile_columns_minus1;i++)

[0223] ColBd[ i + 1 ] = ColBd[ i ] + ColWidth[ i ] (6-3)

[0224] The list RowBd[j] (the value range of j is 0 to num_tile_rows_minus1+1, including the end value) represents the position of the j-th tile row boundary in units of CTB, which is derived as follows.

[0225] for(RowBd[0]=0,j=0;j<=num_tile_rows_minus1;j++)

[0226] RowBd[ j + 1 ] = RowBd[ j ] + RowHeight[ j ] (6-4)

[0227] The list CtbAddrRsToTs[ctbAddrRs] (the value range of ctbAddrRs is 0 to SubPicSizeInCtbsY–1, including the end value) represents the conversion from the CTB address under the CTB raster scan in the sub-image to the CTB address under the block scan in the sub-image, which is derived as follows.

[0228]

[0229]

[0230] The list CtbAddrTsToRs[ctbAddrTs] (the value range of ctbAddrTs is 0 to SubPicSizeInCtbsY–1, including the end value) represents the conversion from the CTB address under the block scan to the CTB address under the CTB raster scan in the sub-image, which is derived as follows.

[0231] for( ctbAddrRs = 0; ctbAddrRs < SubPicSizeInCtbsY; ctbAddrRs++ ) (6-6)

[0232] CtbAddrTsToRs[CtbAddrRsToTs[ctbAddrRs]]=ctbAddrRs

[0233] The list TileId[ctbAddrTs] (the value range of ctbAddrTs is 0 to SubPicSizeInCtbsY–1, including the end value) represents the conversion from the CTB address under the block scanning in the sub-image to the block ID, which is derived as follows.

[0234]

[0235] The list NumCtusInTile[tileIdx] (tileIdx ranges from 0 to NumTilesInSubPic–1, inclusive) represents the conversion from tile index to the number of CTUs in the tile, derived as follows:

[0236]

[0237] The list FirstCtbAddrTs[tileIdx] (tileIdx ranges from 0 to NumTilesInSubPic–1, inclusive) represents the conversion from the tile ID to the CTB address of the first CTB in the tile under the tile scan, as derived as follows.

[0238]

[0239] The value of ColumnWidthInLumaSamples[i] represents the width of the i-th tile column in luma samples and is set to ColWidth[i] << CtbLog2SizeY (where i ranges from 0 to num_tile_columns_minus1, inclusive). The value of RowHeightInLumaSamples[j] represents the height of the j-th tile row in luma samples and is set to RowHeight[j] << CtbLog2SizeY (where j ranges from 0 to num_tile_rows_minus1, inclusive).

[0240] The exemplary sequence parameter set RBSP syntax is as follows.

[0241]

[0242]

[0243] The exemplary picture parameter set RBSP syntax is as follows.

[0244]

[0245]

[0246] The exemplary general tile group header syntax is as follows.

[0247]

[0248] The exemplary coding tree unit syntax is as follows.

[0249]

[0250] The exemplary sequence parameter set RBSP semantics are as follows.

[0251] bit_depth_chroma_minus8 represents the bit depth of the samples in the chroma array BitDepthC, and the value of the chroma quantization parameter range offset QpBdOffsetC is as follows:

[0252] BitDepthC = 8 + bit_depth_chroma_minus8 (7-4)

[0253] QpBdOffsetC = 6 * bit_depth_chroma_minus8 (7-5)

[0254] The value range of bit_depth_chroma_minus8 shall be from 0 to 8, inclusive.

[0255] num_sub_pics_minus1+1 indicates the number of sub-pictures in each decoded picture in the CVS. The value range of num_sub_pics_minus1 should be 0 to 1024, inclusive. sub_pic_id_len_minus1+1 indicates the number of bits used to represent the syntax element sub_pic_id[i] in the SPS and the syntax element tile_group_sub_pic_id in the tile group header. The value range of sub_pic_id_len_minus1 should be Ceil(Log2(num_sub_pic_minus1+1)–1 to 9, inclusive. sub_pic_level_present_flag is set to 1, indicating that the syntax element sub_pic_level_idc[i] may be present. sub_pic_level_present_flag is set to 0, indicating that the syntax element sub_pic_level_idc[i] does not exist. sub_pic_id[i] represents the sub-picture ID of the i-th sub-picture in each decoded picture in the CVS. The length of sub_pic_id[i] is (sub_pic_id_len_minus1+1) bits.

[0256] sub_pic_treated_as_pic_flag[i] is set to 1 to indicate that the i-th sub-picture of each decoded picture in the CVS is treated as a picture during decoding without in-loop filtering operations. sub_pic_treated_as_pic_flag[i] is set to 0 to indicate that the i-th sub-picture of each decoded picture in the CVS is not treated as a picture during decoding without in-loop filtering operations. sub_pic_level_idc[i] indicates the level to which the i-th sub-picture sequence conforms, where the i-th sub-picture sequence only includes the VCL NAL units of the sub-picture with the sub-picture ID equal to sub_pic_id[i] in the CVS and its associated non-VCL NAL units. sub_pic_x_offset[i] indicates the horizontal offset of the top-left luma sample in the i-th sub-picture relative to the top-left luma sample in each picture in the CVS, in units of luma samples. When sub_pic_x_offset[i] is not present, the value of ref_pic_list_idx[i] is inferred to be 0. sub_pic_y_offset[i] represents the vertical offset of the top-left luma sample in the ith sub-image relative to the top-left luma sample in each image in the CVS, in units of luma samples. When sub_pic_y_offset[i] is not present, the value of sub_pic_y_offset[i] is inferred to be 0. sub_pic_width_in_luma_samples[i] represents the width of the ith sub-image in each image in the CVS, in units of luma samples. When the sum of sub_pic_x_offset[i] and sub_pic_width_in_luma_samples[i] is less than pic_width_in_luma_samples, the value of sub_pic_width_in_luma_samples[i] should be an integer multiple of CtbSizeY. When sub_pic_width_in_luma_samples[i] is not present, the value of sub_pic_width_in_luma_samples[i] is inferred to be pic_width_in_luma_samples. sub_pic_height_in_luma_samples[i] represents the height of the i-th sub-picture in each picture in the CVS, in units of luma samples.When the sum of sub_pic_y_offset[i] and sub_pic_height_in_luma_samples[i] is less than pic_height_in_luma_samples, the value of sub_pic_height_in_luma_samples[i] should be an integer multiple of CtbSizeY. When sub_pic_height_in_luma_samples[i] does not exist, the value of sub_pic_height_in_luma_samples[i] is inferred to be pic_height_in_luma_samples.

[0257] For codestream conformance, the following constraints apply. For any integer values ​​of i and j, when i is equal to j, the values ​​of sub_pic_id[i] and sub_pic_id[j] shall be different. For any two sub-pictures subpicA and subpicB, when the sub-picture ID of subpicA is less than the sub-picture ID of subpicB, any decoded partition group NAL unit of subPicA shall follow any decoded partition group NAL unit of subPicB in decoding order. The shape of sub-pictures shall ensure that each sub-picture, when decoded, has its entire left and entire top borders consisting of picture boundaries or of the boundaries of one or more previously decoded sub-pictures.

[0258] The list SubPicIdx[spId] with spId values ​​equal to sub_pic_id[i] (i ranges from 0 to num_sub_pics_minus1, inclusive) represents the conversion from sub-picture ID to sub-picture index, derived as follows:

[0259] for(i=0;i<=num_sub_pics_minus1;i++)

[0260] SubPicIdx[ sub_pic_id[ i ] ] = i (7-5)

[0261] log2_max_pic_order_cnt_lsb_minus4 represents the value of the variable MaxPicOrderCntLsb used in the decoding process of the picture order number, as described below:

[0262] MaxPicOrderCntLsb = 2( log2_max_pic_order_cnt_lsb_minus4 + 4 ) (7-5)

[0263] The value range of log2_max_pic_order_cnt_lsb_minus4 should be 0 to 12, including the end value.

[0264] Exemplary picture parameter set RBSP semantics are described below.

[0265] When the PPS syntax elements pps_seq_parameter_set_id and loop_filter_across_tiles_enabled_flag are present, the value of each of them shall be the same in all PPSs referenced by the tile group headers of the decoded picture. pps_pic_parameter_set_id identifies the SPPS for reference by other syntax elements. The value range of pps_pic_parameter_set_id shall be 0 to 63, inclusive. pps_seq_parameter_set_id indicates the value of sps_seq_parameter_set_id of the active SPS. The value range of pps_seq_parameter_set_id shall be 0 to 15, inclusive. loop_filter_across_sub_pic_enabled_flag is set to 1, indicating that in-loop filtering operations can be performed across the boundaries of sub-pictures referencing the PPS. loop_filter_across_sub_pic_enabled_flag is set to 0, indicating that in-loop filtering operations are not performed across the boundaries of sub-pictures referencing the PPS.

[0266] single_tile_in_sub_pic_flag is set to 1, indicating that only one tile is included in each sub-image of the referenced PPS. single_tile_in_pic_flag is set to 0, indicating that more than one tile is included in each sub-image of the referenced PPS. num_tile_columns_minus1+1 indicates the number of tile columns of the partitioned sub-image. The value of num_tile_columns_minus1 should range from 0 to 1024, inclusive. When num_tile_columns_minus1 is not present, the value of num_tile_columns_minus1 is inferred to be 0. num_tile_rows_minus1+1 indicates the number of tile rows of the partitioned sub-image. The value of num_tile_rows_minus1 should range from 0 to 1024, inclusive. When num_tile_rows_minus1 is not present, the value of num_tile_rows_minus1 is inferred to be 0. The variable NumTilesInSubPic is set to (num_tile_columns_minus1+1)*(num_tile_rows_minus1+1). When single_tile_in_sub_pic_flag is equal to 0, NumTilesInSubPic should be greater than 1.

[0267] uniform_tile_spacing_flag is set to 1, indicating that tile column boundaries and tile row boundaries are uniformly distributed across sub-images. uniform_tile_spacing_flag is set to 0, indicating that tile column boundaries and tile row boundaries are not uniformly distributed across sub-images, but are explicitly indicated using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. When uniform_tile_spacing_flag is not present, the value of uniform_tile_spacing_flag is inferred to be 1. tile_column_width_minus1[i]+1 represents the width of the i-th tile column, in CTBs. tile_row_height_minus1[i]+1 represents the height of the i-th tile row, in CTBs. single_tile_per_tile_group is set to 1, indicating that each tile group of the reference PPS consists of one tile. single_tile_per_tile_group is set to 0, indicating that the tile group referencing the PPS may include more than one tile.

[0268] rect_tile_group_flag is set to 0, indicating that the tiles within each tile group in the sub-image are in raster scan order and the tile group information is not indicated in the PPS. rect_tile_group_flag is set to 1, indicating that the tiles within each tile group cover a rectangular area or a square area in the sub-image, and the tile group information is indicated in the PPS. When single_tile_per_tile_group_flag is set to 1, rect_tile_group_flag is inferred to be 1. num_tile_groups_in_sub_pic_minus1+1 indicates the number of tile groups in each sub-image referenced to the PPS. The value range of num_tile_groups_in_sub_pic_minus1 should be 0 to NumTilesInSubPic–1, inclusive. When num_tile_groups_in_sub_pic_minus1 is not present and single_tile_per_tile_group_flag is equal to 1, then the value of num_tile_groups_in_sub_pic_minus1 is inferred to be NumTilesInSubPic–1.

[0269] top_left_tile_idx[i] represents the tile index of the tile located at the top left corner of the i-th tile group in the sub-image. For any i not equal to j, the value of top_left_tile_idx[i] shall not be equal to the value of top_left_tile_idx[j]. When top_left_til_idx[i] does not exist, the value of top_left_til_idx[i] is inferred to be i. The length of the syntax element top_left_til_idx[i] is Ceil(Log2(NumTilesInSubPic)) bits. bottom_right_tile_idx[i] represents the tile index of the tile located at the bottom right corner of the i-th tile group in the sub-image. When single_tile_per_tile_group_flag is set to 1, bottom_right_tile_idx[i] is inferred to be top_left_tile_idx[i]. The length of the syntax element bottom_right_tile_idx[i] is Ceil(Log2(NumTilesInSubPic)) bits.

[0270] The requirement for codestream consistency is that any particular tile should only be included in one tile group. The variable NumTilesInTileGroup[i] represents the number of tiles in the i-th tile group in the sub-image. This variable and related variables are derived as follows:

[0271] deltaTileIdx=bottom_right_tile_idx[i]–top_left_tile_idx[i]

[0272] NumTileRowsInTileGroupMinus1[i]=deltaTileIdx / (num_tile_columns_minus1+1)

[0273] NumTileColumnsInTileGroupMinus1[ i ] = deltaTileIdx % ( num_tile_columns_minus1 + 1 ) (7-33)

[0274] NumTilesInTileGroup[i]=(NumTileRowsInTileGroupMinus1[i]+1)*

[0275] (NumTileColumnsInTileGroupMinus1[i]+1)

[0276] loop_filter_across_tiles_enabled_flag is set to 1, indicating that in-loop filtering operations can be performed across tile boundaries in sub-images that reference PPS. loop_filter_across_tiles_enabled_flag is set to 0, indicating that in-loop filtering operations are not performed across tile boundaries in sub-images that reference PPS. In-loop filtering operations include deblocking filter operations, sample adaptive offset filter operations, and adaptive loop filter operations. When loop_filter_across_tiles_enabled_flag is not present, the value of loop_filter_across_tiles_enabled_flag is inferred to be 1. num_ref_idx_default_active_minus1[i]+1, when i is equal to 0, represents the inferred value of the variable NumRefIdxActive[0] for the P slice group or B slice group with num_ref_idx_active_override_flag equal to 0; when i is equal to 1, represents the inferred value of NumRefIdxActive[1] for the B slice group with num_ref_idx_active_override_flag equal to 0. The value range of num_ref_idx_default_active_minus1[i] should be 0 to 14, inclusive.

[0277] Exemplary general tile group header semantics are described as follows. When the tile group header syntax elements tile_group_pic_order_cnt_lsb and tile_group_temporal_mvp_enabled_flag are present, the value of each of the syntax elements shall be the same in all tile group headers of a decoded image. When tile_group_pic_parameter_set_id is present, the value of tile_group_pic_parameter_set_id shall be the same in all tile group headers of a decoded sub-image. tile_group_pic_parameter_set_id represents the value of pps_pic_parameter_set_id of the currently used PPS. The value range of tile_group_pic_parameter_set_id shall be 0 to 63, inclusive. The requirement for codestream consistency is that the value of TemporalId of the current image shall be greater than or equal to the value of TemporalId of each PPS referenced by the tile group in the current image. tile_group_sub_pic_id identifies the sub-image to which the tile group header belongs. The length of tile_group_sub_pic_id is (sub_pic_id_len_minus1+1) bits. The value of tile_group_sub_pic_id should be the same for all tile group headers of a coded sub-picture.

[0278] The variables SubPicWidthInCtbsY, SubPicHeightInCtbsY, and SubPicSizeInCtbsY are derived as follows:

[0279] i=SubPicIdx[tile_group_subpic_id]

[0280] SubPicWidthInCtbsY=

[0281] Ceil(sub_pic_width_in_luma_samples[i]÷CtbSizeY)(7-34)

[0282] SubPicHeightInCtbsY=Ceil(sub_pic_height_in_luma_samples[i]÷CtbSizeY)

[0283] SubPicSizeInCtbsY=SubPicWidthInCtbsY*SubPicHeightInCtbsY

[0284] The following variables are derived by calling the CTB raster and block scan conversion process: the list ColWidth[i] (i ranges from 0 to num_tile_columns_minus1, inclusive), which represents the width of the i-th block column in CTBs; the list RowHeight[j] (j ranges from 0 to num_tile_rows_minus1, inclusive), which represents the height of the j-th block row in CTBs; the list ColBd[i] (i ranges from 0 to num_tile_columns_minus1+1, inclusive), which represents the position of the i-th block column boundary in CTBs; the list R owBd[j] (the value range of j is 0 to num_tile_rows_minus1+1, including the end value), indicating the position of the j-th tile row boundary in CTB units; the list CtbAddrRsToTs[ctbAddrRs] (the value range of ctbAddrRs is 0 to SubPicSizeInCtbsY–1, including the end value), indicating the conversion from the CTB address under the CTB raster scan in the sub-image to the CTB address under the block scan in the sub-image; the list CtbAddrTsToRs[ctbAddrTs] (the value range of ctbAddrTs is 0 to SubPicSizeInCtbsY–1, including the end value) Value), indicating the conversion of the CTB address under the block scanning in the sub-image to the CTB address under the CTB raster scanning in the sub-image; List TileId[ctbAddrTs] (the value range of ctbAddrTs is 0 to SubPicSizeInCtbsY–1, including the end value), indicating the conversion of the CTB address under the block scanning in the sub-image to the block ID; List NumCtusInTile[tileIdx] (the value range of tileIdx is 0 to NumTilesInSubPic–1, including the end value), indicating the conversion from the block index to the number of CTUs in the block; List FirstCtbAddrTs[tileIdx](t The value range of ileIdx is 0 to NumTilesInSubPic–1, including the end value), which represents the conversion from the tile ID to the CTB address under the tile scan of the first CTB in the tile; the list ColumnWidthInLumaSamples[i] (the value range of i is 0 to num_tile_columns_minus1, including the end value) represents the width of the i-th tile column in units of brightness samples; the list RowHeightInLumaSamples[j] (the value range of j is 0 to num_tile_rows_minus1, including the end value) represents the height of the j-th tile row in units of brightness samples.

[0285] The values ​​of ColumnWidthInLumaSamples[i] (i ranges from 0 to num_tile_columns_minus1, inclusive) and RowHeightInLumaSamples[j] (j ranges from 0 to num_tile_rows_minus1, inclusive) should both be greater than 0. The variables SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:

[0286]

[0287] For each tile in the current sub-image, index i = 0 ... NumTilesInSubPic - 1, the variables TileLeftBoundaryPos[i], TileTopBoundaryPos[i], TileRightBoundaryPos[i], and TileBotBoundaryPos[i] are derived as follows:

[0288]

[0289] tile_group_address represents the tile address of the first tile in the tile group. When tile_group_address is not present, the value of tile_group_address is inferred to be 0. If rect_tile_group_flag is equal to 0, the following applies: the tile address is the tile ID; the length of tile_group_address is Ceil(Log2(NumTilesInSubPic)) bits; the value of tile_group_address should range from 0 to NumTilesInSubPic–1, inclusive. Otherwise (rect_tile_group_flag is equal to 1), the following applies: the tile address is the tile group index of a tile group among multiple tile groups in the sub-picture; the length of tile_group_address is Ceil(Log2(num_tile_groups_in_sub_pic_minus1+1)) bits; the value of tile_group_address should range from 0 to num_tile_groups_in_sub_pic_minus1, inclusive.

[0290] The codestream conformance requirement is that the following constraints apply. The value of tile_group_address shall not be equal to the value of tile_group_address of any other decoded tile group NAL unit of the same decoded picture. The tile groups in a sub-picture shall be arranged in increasing order of their tile_group_address values. The shape of the tile groups in a sub-picture shall ensure that each tile, when decoded, has its entire left and entire top borders consisting of either the sub-picture boundary or the boundary of one or more previously decoded tiles.

[0291] num_tiles_in_tile_group_minus1, when present, represents the number of tiles in the tile group minus 1. The value of num_tiles_in_tile_group_minus1 should range from 0 to NumTilesInSubPic–1, inclusive. When num_tiles_in_tile_group_minus1 does not exist, the value of num_tiles_in_tile_group_minus1 is inferred to be 0. The variable NumTilesInCurrTileGroup represents the number of tiles in the current tile group, and TgTileIdx[i] represents the tile index of the i-th tile in the current tile group, both of which are derived as follows:

[0292]

[0293]

[0294] tile_group_type indicates the decoding type of the tile group.

[0295] An exemplary derivation of temporal luminance motion vector prediction is as follows. The variables mvLXCol and availableFlagLXCol are derived as follows: If tile_group_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set to 0 and availableFlagLXCol is set to 0. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), the following steps in order apply: The lower right corner concatenated motion vector and the lower right boundary sample position are derived as follows:

[0296] xColBr = xCb + cbWidth (8-414)

[0297] yColBr = yCb + cbHeight (8-415)

[0298] rightBoundaryPos=sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]]?

[0299] SubPicRightBoundaryPos:pic_width_in_luma_samples–1(8-415)

[0300] botBoundaryPos=sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]]?

[0301] SubPicBotBoundaryPos:pic_height_in_luma_samples–1(8-415)

[0302] If yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos, and xColBr is less than or equal to rightBoundaryPos, then the following applies: The variable colCb represents the luma coding block that overwrites the modified position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the collocated image represented by ColPic. The luma position (xColCb, yColCb) is set to the position of the top-left sample of the collocated luma coding block represented by colCb relative to the top-left luma sample of the collocated image represented by ColPic. The derivation of the collocated motion vector is called with as input currCb, colCb, (xColCb, yColCb), refIdxLX, and sbFlag set to 0, and the output is assigned to mvLXCol and availableFlagLXCol. Otherwise, both components of mvLXCol are set to 0 and availableFlagLXCol is set to 0.

[0303] An exemplary luma sample bilinear interpolation process is as follows. For i = 0..1, the chroma position (xInti, yInti) with an integer number of sample units is derived as follows: If sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 1, then the following applies:

[0304] xInti = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL+ i ) (8-458)

[0305] yInti = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i )(8-458)

[0306] Otherwise (sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 0), the following applies:

[0307] xInti=sps_ref_wraparound_enabled_flag?

[0308] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,(xIntL+i)):

[0309] (8-459)

[0310] Clip3(0,picW–1,xIntL+i)

[0311] yInti = Clip3( 0, picH – 1, yIntL + i ) (8-460)

[0312] The 8-tap interpolation filtering process for luma samples is as follows. For i = 0..7, the chroma positions (xInti, yInti) with integer number of sample units are derived as follows: If sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 1, then the following applies:

[0313] xInti = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xIntL+ i – 3 ) (8-830)

[0314] yInti = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yIntL + i– 3 ) (8-830)

[0315] Otherwise (sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 0), the following applies:

[0316] xInti=sps_ref_wraparound_enabled_flag?

[0317] ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xIntL+i–3):

[0318] (8-831)

[0319] Clip3(0,picW–1,xIntL+i–3)

[0320] yInti = Clip3( 0, picH – 1, yIntL + i – 3 ) (8-832)

[0321] An exemplary chroma sample interpolation process is as follows. The variable xOffset is set to (sps_ref_wraparound_offset_minus1+1)*MinCbSizeY) / SubWidthC. For i=0..3, the chroma position (xInti, yInti) with an integer number of sample units is derived as follows:

[0322] If sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 1, the following applies:

[0323] xInti=Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xIntL+i)

[0324] (8-844)

[0325] yInti=Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yIntL+i)

[0326] (8-844)

[0327] Otherwise (sub_pic_treated_as_pic_flag[SubPicIdx[tile_group_subpic_id]] is equal to 0), the following applies:

[0328] xInti = sps_ref_wraparound_enabled_flag? ClipH(xOffset, picWC,xIntC + i – 1): (8-845)

[0329] Clip3(0,picWC–1,xIntC+i–1)

[0330] yInti = Clip3( 0, picHC – 1, yIntC + i – 1 ) (8-846)

[0331] An exemplary deblocking filter process is described as follows. The deblocking filter process is applied to all coded sub-block edges and transform block edges in the image, except for the following types of edges: edges located at picture boundaries; edges coinciding with boundaries of sub-pictures with loop_filter_across_sub_pic_enabled_flag equal to 0; edges coinciding with boundaries of tiles with loop_filter_across_tiles_enabled_flag equal to 0; edges coinciding with top or left boundaries in or within tile groups with tile_group_deblocking_filter_disabled_flag equal to 1; edges that do not correspond to boundaries of the 8×8 sample grid of the considered component; edges within chroma components that use inter-frame prediction on both sides of the edge; edges of chroma transform blocks that are not edges of related transform units; and edges of luma transform blocks that span coding units with IntraSubPartitionsSplit value not equal to ISP_NO_SPLIT.

[0332] An exemplary deblocking filter process in one direction is described below. For each decoding unit with a decoded block width of log2CbW, a decoded block height of log2CbH, and the position of the top-left sample of the decoded block being (xCb, yCb), when edgeType is equal to EDGE_VER and xCb % 8 is equal to 0, or when edgeType is equal to EDGE_HOR and yCb % 8 is equal to 0, the edge is filtered through the following steps executed in sequence. The decoded block width nCbW is set to 1 << log2CbW, and the decoded block height nCbH is set to 1 << log2CbH. The variable filterEdgeFlag is derived as follows: If edgeType is equal to EDGE_VER and one or more of the following conditions are true, then filterEdgeFlag is set to 0: The left boundary of the current decoded block is the left boundary of the image. The left boundary of the current decoded block is the left or right boundary of a sub-image and loop_filter_across_sub_pic_enabled_flag is equal to 0. The left boundary of the current decoded block is the left boundary of a tile and loop_filter_across_tiles_enabled_flag is equal to 0. If edgeType is equal to EDGE_HOR and one or more of the following conditions are true, then the variable filterEdgeFlag is set to 0. The top boundary of the current luma decoded block is the top boundary of the image. The top boundary of the current decoded block is the top or bottom boundary of a sub-image and loop_filter_across_sub_pic_enabled_flag is equal to 0. The top boundary of the current decoded block is the top boundary of a tile and loop_filter_across_tiles_enabled_flag is equal to 0. Otherwise, filterEdgeFlag is set to 1.

[0333] An exemplary CTB modification process is as follows. For all sample positions (xSi, ySj) and (xYi, yYj), where i = 0..nCtbSw-1, j = 0..nCtbSh-1, the following applies depending on the values ​​of pcm_loop_filter_disabled_flag, pcm_flag[xYi][yYj], and cu_transquant_bypass_flag for the coding unit including the coding block covering recPicture[xSi][ySj]. edgeIdx is set to 0 if one or more of the following conditions are true for all sample positions (xSik', ySjk') and (xYik', yYjk') (k = 0..1). The sample at position (xSik', ySjk') is outside the picture boundary. The sample at position (xSik', ySjk') belongs to a different sub-image, and the loop_filter_across_sub_pic_enabled_flag in the tile group to which the sample recPicture[xSi][ySj] belongs is equal to 0. loop_filter_across_tiles_enabled_flag is equal to 0, and the sample at position (xSik', ySjk') belongs to a different tile.

[0334] An exemplary coding tree block filtering process for luma samples is described as follows. To derive the filtered reconstructed luma sample alfPictureL[x][y], each reconstructed luma sample recPictureL[x][y] within the current luma coding tree block is filtered as follows, where x, y = 0..CtbSizeY–1. The position (hx, vy) of each corresponding luma sample (x, y) within a given array of luma samples recPicture is derived as follows. If the loop_filter_across_tiles_enabled_flag of the tile tileA including the luma sample at position (hx, vy) is equal to 0, then assuming that the variable tileIdx is the tile index of tileA, the following applies:

[0335] hx=Clip3(TileLeftBoundaryPos[tileIdx],TileRightBoundaryPos[tileIdx],xCtb+x)

[0336] (8-1140)

[0337] vy=Clip3(TileTopBoundaryPos[tileIdx],TileBotBoundaryPos[tileIdx],yCtb+y)

[0338] (8-1141)

[0339] If loop_filter_across_sub_pic_enabled_flag is equal to 0 in the sub-picture that includes the luma sample at position (hx,vy), the following applies:

[0340] hx = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, xCtb + x )(8-1140)

[0341] vy = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, yCtb + y )(8-1141)

[0342] Otherwise, the following applies:

[0343] hx = Clip3( 0, pic_width_in_luma_samples – 1, xCtb + x ) (8-1140)

[0344] vy = Clip3( 0, pic_height_in_luma_samples – 1, yCtb + y ) (8-1141)

[0345] An exemplary derivation of the ALF transpose and filter index for luma samples is as follows. The position (hx, vy) of each corresponding luma sample (x, y) within a given array of luma samples recPicture is derived as follows. If loop_filter_across_tiles_enabled_flag is equal to 0 for tileA including the luma sample at position (hx, vy), then tileIdx is assumed to be the tile index of tileA, and the following applies:

[0346] hx = Clip3(TileLeftBoundaryPos[tileIdx], TileRightBoundaryPos[tileIdx], x) (8-1140)

[0347] vy = Clip3(TileTopBoundaryPos[tileIdx], TileBotBoundaryPos[tileIdx], y) (8-1141)

[0348] Otherwise, if loop_filter_across_sub_pic_enabled_flag is equal to 0 for the sub-picture including the luma sample at position (hx,vy), the following applies:

[0349] hx = Clip3( SubPicLeftBoundaryPos, SubPicRightBoundaryPos, x ) (8-1140)

[0350] vy = Clip3( SubPicTopBoundaryPos, SubPicBotBoundaryPos, y ) (8-1141)

[0351] Otherwise, the following applies:

[0352] hx = Clip3( 0, pic_width_in_luma_samples – 1, x ) (8-1145)

[0353] vy = Clip3( 0, pic_height_in_luma_samples – 1, y ) (8-1146)

[0354] An exemplary coding tree block filtering process for chroma samples is described as follows. To derive the filtered reconstructed chroma samples alfPicture[x][y], each reconstructed chroma sample recPicture[x][y] within the current chroma coding tree block is filtered as follows, where x, y = 0..ctbSizeC–1. The position (hx, vy) of each corresponding chroma sample (x, y) within a given array of chroma samples recPicture is derived as follows. If the loop_filter_across_tiles_enabled_flag of the tile tileA including the chroma sample at position (hx, vy) is equal to 0, then tileIdx is assumed to be the tile index of tileA, and the following applies:

[0355] hx=Clip3(TileLeftBoundaryPos[tileIdx] / SubWidthC,

[0356] TileRightBoundaryPos[tileIdx] / SubWidthC,xCtb+x)(8-1140)

[0357] vy=Clip3(TileTopBoundaryPos[tileIdx] / SubWidthC,

[0358] TileBotBoundaryPos[tileIdx] / SubWidthC,yCtb+y)(8-1141)

[0359] Otherwise, if loop_filter_across_sub_pic_enabled_flag is equal to 0 for the sub-picture including the chroma sample at position (hx,vy), the following applies:

[0360] hx=Clip3(SubPicLeftBoundaryPos / SubWidthC,

[0361] SubPicRightBoundaryPos / SubWidthC,xCtb+x)(8-1140)

[0362] vy=Clip3(SubPicTopBoundaryPos / SubWidthC,

[0363] SubPicBotBoundaryPos / SubWidthC,yCtb+y)(8-1141)

[0364] Otherwise, the following applies:

[0365] hx = Clip3( 0, pic_width_in_luma_samples / SubWidthC – 1, xCtbC + x )(8-1177)

[0366] vy = Clip3( 0, pic_height_in_luma_samples / SubHeightC - 1, yCtbC +y ) (8-1178)

[0367] The variable sum is derived as follows:

[0368]

[0369] sum = ( sum + 64 ) >> 7 (8-1180)

[0370] The modified filtered reconstructed chrominance image samples alfPicture[xCtbC+x][yCtbC+y] are derived as follows:

[0371] alfPicture[ xCtbC + x ][ yCtbC + y ] = Clip3( 0, ( 1 << BitDepthC )– 1, sum ) (8-1181)

[0372] Fig.12 1 is a schematic diagram of an exemplary video decoding device 1200. The video decoding device 1200 is suitable for implementing the disclosed examples / embodiments described herein. The video decoding device 1200 includes a downstream port 1220, an upstream port 1250 and / or a transceiver unit (Tx / Rx) 1210. The transceiver unit 1210 includes a transmitter and / or a receiver for communicating data in the upstream and / or downstream through a network. The video decoding device 1200 also includes a processor 1230 and a memory 1232. The processor 1230 includes a logic unit and / or a central processing unit (CPU) to process data. The memory 1232 is used to store data. The video decoding device 1200 may also include an electrical component, an optical-to-electrical (OE) component, an electrical-to-optical (EO) component, and / or a wireless communication component coupled to the upstream port 1250 and / or the downstream port 1220 for data communication via an electrical communication network, an optical communication network, or a wireless communication network. The video decoding device 1200 may also include an input and / or output (I / O) device 1260 for data communication with a user. The I / O device 1260 may include an output device, such as a display that displays video data, a speaker that outputs audio data, etc. The I / O device 1260 may also include input devices such as a keyboard, a mouse, a trackball, and / or corresponding interfaces for interacting with these output devices.

[0373] The processor 1230 is implemented by hardware and software. The processor 1230 can be implemented as one or more CPU chips, one or more cores (for example, implemented as a multi-core processor), one or more field-programmable gate arrays (FPGA), one or more application-specific integrated circuits (ASIC), and one or more digital signal processors (DSP). The processor 1230 communicates with the downstream port 1220, Tx / Rx 1210, the upstream port 1250, and the memory 1232. The processor 1230 includes a decoding module 1214. The encoding module 1214 implements the disclosed embodiments described herein, such as method 100, method 1300, and method 1400, which can use the in-loop filter 1000, the code stream 1100, the image 500, and / or the current block 801 and / or 901, and the current block 801 and / or 901 can be decoded according to the candidate list generated according to the mode 900 according to the unidirectional inter-frame prediction 600 and / or the bidirectional inter-frame prediction 700. The decoding module 1214 can also implement any other method / mechanism described herein. In addition, the decoding module 1214 can implement the encoding and decoding system 200, the encoder 300 and / or the decoder 400. For example, the decoding module 1214 can implement the first, second, third, fourth, fifth and / or sixth exemplary implementations as described above. Therefore, the decoding module 1214 enables the video decoding device 1200 to provide additional functions and / or improve decoding efficiency when decoding video data. Therefore, the decoding module 1214 improves the function of the video decoding device 1200 and solves the problems unique to the field of video decoding. In addition, the decoding module 1214 affects the conversion of the video decoding device 1200 to different states. Optionally, the decoding module 1214 can be implemented as an instruction stored in the memory 1232 and executed by the processor 1230 (for example, implemented as a computer program product stored in a non-transient medium).

[0374] The memory 1232 includes one or more memory types, such as a disk, a tape drive, a solid-state hard disk, a read only memory (ROM), a random access memory (RAM), a flash memory, a ternary content-addressable memory (TCAM), a static random-access memory (SRAM), etc. The memory 1232 can be used as an overflow data storage device to store programs when they are selected for execution and to store instructions and data read during the execution of the program.

[0375] Fig.13 Flow chart of an exemplary method 1300 for encoding a video sequence into a bitstream (e.g., bitstream 1100) when a clipping function is applied to an interpolation filter (e.g., interpolation filter 913) when a sub-image (e.g., sub-image 510) is treated as an image (e.g., image 500). The method 1300 may be adopted by the codec system 200, the encoder 300, and / or the video decoding device 1200 when performing the method 100 to encode the current block 801 and / or 901 according to the unidirectional inter-frame prediction 600 and / or the bidirectional inter-frame prediction 700 using the in-loop filter 1000 and / or the candidate list generated according to the mode 900.

[0376] Method 1300 may begin with an encoder receiving a video sequence including a plurality of images and determining to encode the video sequence into a bitstream according to user input or the like. In step 1301, the encoder divides a current image into sub-images. The encoder also divides the sub-images into blocks. In step 1303, the encoder determines to encode the blocks according to inter-frame prediction. Accordingly, the encoder selects a motion vector to encode the blocks.

[0377] In step 1305, the encoder applies a clipping function to the sample positions in the reference block to which the motion vector points. The clipping function is applied to support the application of an interpolation filter when the motion vector points outside the sub-image. This process occurs when a flag is set to indicate that the sub-image is to be treated as an image, so that the sub-image can be decoded to support extraction independently of other sub-images in the image. In this context, a sub-image is treated as a sub-image when it is decoded without reference to data in other sub-images, so that it can be extracted separately.

[0378] In step 1307, the encoder applies the interpolation filter to the result of the clipping function to obtain predicted sample values. In one example, the interpolation filter includes a luma sample bilinear interpolation process. In this case, the block includes a luma sample block. In addition, the predicted sample values ​​include predicted luma sample values. In this case, the luma sample bilinear interpolation process receives an input including a luma position (xIntL, yIntL) having an integer number of sample units, and the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). In this case, the clipping function in step 1305 is applied to the sample position as described below.

[0379] When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0380] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i),

[0381] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i),

[0382] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula:

[0383]

[0384] Where x, y, and z are numeric input values.

[0385] In another example, the interpolation filter comprises a luma sample 8-tap interpolation filter process. In this case, the block comprises a luma sample block and the predicted sample values ​​comprise predicted luma sample values. In addition, the luma sample 8-tap interpolation filter process receives an input comprising a luma position (xIntL, yIntL) having an integer number of sample units, and the luma sample 8-tap interpolation filter process outputs a predicted luma sample value (predSampleLXL). In addition, the clipping function in step 1305 is applied to the sample position as described below. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0386] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i–3),

[0387] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i–3),

[0388] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula:

[0389]

[0390] Where x, y, and z are numeric input values.

[0391] In another example, the interpolation filter comprises a chroma sample interpolation process. In this case, the block comprises a chroma sample block and the predicted sample values ​​comprise predicted chroma sample values. In this case, the chroma sample interpolation process receives an input comprising a chroma position (xIntC, yIntC) having an integer number of sample units. In addition, the chroma sample interpolation process outputs predicted chroma sample values ​​(predSampleLXC). In this case, the clipping function in step 1305 is applied to the sample position as described below. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0392] xInti=Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xIntC+i),

[0393] yInti=Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yIntC+i),

[0394] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as an image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, SubWidthC and SubHeightC represent the horizontal sampling rate ratio and the vertical sampling rate ratio between the luminance sample and the chrominance sample, and Clip3 represents the clipping function according to the following formula:

[0395]

[0396] Where x, y, and z are numeric input values.

[0397] In step 1309, the encoder encodes the block into a bitstream according to the predicted sample value and the motion vector. The bitstream may be stored for transmission to a decoder.

[0398] Fig.14A flowchart of an exemplary method 1400 for decoding a video sequence from a bitstream (e.g., bitstream 1100) while applying a clipping function to an interpolation filter (e.g., interpolation filter 913) when a sub-image (e.g., sub-image 510) is treated as an image (e.g., image 500). The method 1400 may be employed by the codec system 200, the decoder 400, and / or the video decoding device 1200 when performing the method 100 to decode the current block 801 and / or 901 according to the unidirectional inter-frame prediction 600 and / or the bidirectional inter-frame prediction 700 using the in-loop filter 1000 and / or the candidate list generated according to the mode 900.

[0399] Method 1400 may begin when a decoder begins receiving a codestream representing decoded data of a video sequence (e.g., the result of method 1300). In step 1401, the decoder receives a codestream including a current picture, wherein the current picture includes a sub-picture. Pictures, sub-pictures, slices, tiles, CTUs, and / or other sub-regions thereof are decoded according to inter-frame prediction. In step 1402, the decoder determines motion vectors for blocks in the sub-picture.

[0400] In step 1403, the decoder applies a clipping function to the sample positions in the reference block pointed to by the motion vector. The clipping function is applied to support the application of an interpolation filter when the motion vector points outside the sub-image. This process occurs when a flag is set to indicate that the sub-image is to be treated as an image, so that the sub-image can be decoded to support extraction independent of other sub-images in the image.

[0401] In step 1405, the decoder applies the interpolation filter to the result of the clipping function to obtain predicted sample values. In one example, the interpolation filter includes a luma sample bilinear interpolation process. In this case, the block includes a luma sample block. In addition, the predicted sample values ​​include predicted luma sample values. In this case, the luma sample bilinear interpolation process receives an input including a luma position (xIntL, yIntL) having an integer number of sample units, and the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL). In this case, the clipping function in step 1403 is applied to the sample position as described below.

[0402] When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0403] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i),

[0404] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i),

[0405] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula:

[0406]

[0407] Where x, y, and z are numeric input values.

[0408] In another example, the interpolation filter comprises a luma sample 8-tap interpolation filter process. In this case, the block comprises a luma sample block and the predicted sample values ​​comprise predicted luma sample values. In addition, the luma sample 8-tap interpolation filter process receives an input comprising a luma position (xIntL, yIntL) having an integer number of sample units, and the luma sample 8-tap interpolation filter process outputs a predicted luma sample value (predSampleLXL). In addition, the clipping function in step 1403 is applied to the sample position as described below. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0409] xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xIntL+i–3),

[0410] yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yIntL+i–3),

[0411] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as a sub-image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula:

[0412]

[0413] Where x, y, and z are numeric input values.

[0414] In another example, the interpolation filter comprises a chroma sample interpolation process. In this case, the block comprises a chroma sample block and the predicted sample values ​​comprise predicted chroma sample values. In this case, the chroma sample interpolation process receives an input comprising a chroma position (xIntC, yIntC) having an integer number of sample units. In addition, the chroma sample interpolation process outputs predicted chroma sample values ​​(predSampleLXC). In this case, the clipping function in step 1403 is applied to the sample position as described below. When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies:

[0415] xInti=Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xIntC+i),

[0416] yInti=Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yIntC+i),

[0417] Wherein, subpic_treated_as_pic_flag represents the flag set to indicate that the sub-image is treated as an image, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, SubWidthC and SubHeightC represent the horizontal sampling rate ratio and the vertical sampling rate ratio between the luminance sample and the chrominance sample, and Clip3 represents the clipping function according to the following formula:

[0418]

[0419] Where x, y, and z are numeric input values.

[0420] The decoder decodes the block based on the predicted sample values ​​in step 1407. The decoder may then forward the block for display as part of a decoded video sequence.

[0421] Fig.15 A schematic diagram of an exemplary system 1500 for decoding a video sequence composed of images in a bitstream (e.g., bitstream 1100) while applying a clipping function to an interpolation filter (e.g., interpolation filter 913) when a sub-image (e.g., sub-image 510) is treated as an image (e.g., image 500). The system 1500 may be implemented by an encoder and a decoder such as the encoding and decoding system 200, the encoder 300, the decoder 400, and / or the video decoding device 1200. In addition, the system 1500 may be adopted when implementing the method 100, the method 1300, and / or the method 1400 to decode the current block 801 and / or 901 according to the unidirectional inter-frame prediction 600 and / or the bidirectional inter-frame prediction 700 using the in-loop filter 1000 and / or the candidate list generated according to the mode 900.

[0422] The system 1500 includes a video encoder 1502. The video encoder 1502 includes a segmentation module 1503 for segmenting a current image into sub-images and segmenting the sub-images into blocks. The video encoder 1502 also includes a determination module 1504 for determining to encode the block according to inter-frame prediction. The video encoder 1502 also includes a selection module 1505 for selecting a motion vector for encoding the block. The video encoder 1502 also includes an application module 1506 for: when the motion vector points outside the sub-image and when a flag is set to indicate that the sub-image is treated as an image, applying a clipping function to a sample position in a reference block to support application of an interpolation filter. The application module 1506 is also used to apply the interpolation filter to the result of the clipping function to obtain a predicted sample value. The video encoder 1502 also includes an encoding module 1507 for encoding the block into a bitstream according to the predicted sample value and the motion vector. The video encoder 1502 also includes a storage module 1508 for storing the bitstream for sending to a decoder. The video encoder 1502 further includes a sending module 1509, which is used to send the code stream to the video decoder 1510. The video encoder 1502 can also be used to execute any step of the method 1300.

[0423] The system 1500 also includes a video decoder 1510. The video decoder 1510 includes a receiving module 1511 for receiving a code stream including a current image, wherein the current image includes a sub-image decoded according to inter-frame prediction. The video decoder 1510 also includes a determining module 1512 for determining a motion vector of a block in the sub-image. The video decoder 1510 also includes an applying module 1513 for applying a clipping function to a sample position in a reference block to support application of an interpolation filter when the motion vector points outside the sub-image and when a flag is set to indicate that the sub-image is treated as an image. The applying module 1513 is also used to apply the interpolation filter to the result of the clipping function to obtain a predicted sample value. The video decoder 1510 also includes a decoding module 1514 for decoding the block according to the predicted sample value. The video decoder 1510 also includes a forwarding module 1515 for forwarding the block for display as part of a decoded video sequence. The video decoder 1510 can also be used to perform any step of the method 1400.

[0424] A first component is directly coupled to a second component when there are no intermediate components between the first component and the second component other than a line, trace, or other medium. A first component is indirectly coupled to a second component when there are intermediate components between the first component and the second component other than a line, trace, or other medium. The term "coupled" and its variations include direct coupling and indirect coupling. Unless otherwise indicated, the use of the term "about" means a range including ±10% of the subsequent number.

[0425] It should also be understood that the steps of the exemplary methods set forth herein do not necessarily need to be performed in the order described, and the order of the steps of these methods should be understood to be merely exemplary. Similarly, in methods consistent with various embodiments of the present invention, these methods may include other steps, and certain steps may be omitted or combined.

[0426] Although the present invention provides several embodiments, it should be understood that the disclosed systems and methods may be embodied in a variety of other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention are to be considered illustrative rather than restrictive, and the present invention is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or some features may be omitted or not implemented.

[0427] In addition, without departing from the scope of the present invention, the techniques, systems, subsystems and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, components, techniques or methods. Other changes, substitutions, and modified examples can be determined by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.

Claims

1. A method implemented in a decoder, characterized in that: The method comprises: Receiving a code stream including a current image, wherein the current image includes a sub-image decoded according to inter-frame prediction, the sub-image includes a plurality of blocks, and the sub-image is not the block; Parsing the bitstream to obtain quantized transform coefficients; Inverse quantization and inverse transformation are performed on the quantized transformation coefficients to obtain a residual block; determining a motion vector for a block in the sub-image; applying a clipping function to sample positions in a reference block when the motion vector points outside the sub-image and when a flag indicates that the sub-image is to be processed as an image; applying an interpolation filter to the result of the clipping function to obtain a predicted sample value; Obtaining a reconstructed image block according to the residual block and the prediction block, wherein the prediction block includes the prediction sample value; wherein the interpolation filter comprises a luma sample bilinear interpolation process, the block comprises a luma sample block, and the predicted sample value comprises a predicted luma sample value; The luma sample bilinear interpolation process receives an input comprising a luma position (xIntL, yIntL) having an integer number of sample units, the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i), yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i), Wherein, subpic_treated_as_pic_flag represents the flag, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula: Where x, y, and z are numeric input values.

2. The method according to claim 1, characterized in that The interpolation filter comprises a luma sample 8-tap interpolation filtering process, the block comprises a luma sample block, and the predicted sample values ​​comprise predicted luma sample values.

3. The method according to claim 1, characterized in that The interpolation filter comprises a chroma sample interpolation process, the block comprises a chroma sample block, and the predicted sample values ​​comprise predicted chroma sample values.

4. A video decoding device, characterized in that: The video decoding device comprises: A processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to execute the method according to any one of claims 1 to 3.

5. A non-transitory computer-readable medium, characterized in that The non-transitory computer-readable medium includes a computer program product for use by a video decoding device, the computer program product including computer-executable instructions stored in the non-transitory computer-readable medium, and when a processor executes the computer-executable instructions, the video decoding device performs the method according to any one of claims 1 to 3.

6. A decoder, characterized in that: The decoder comprises: A receiving module, configured to receive a code stream including a current image, wherein the current image includes a sub-image decoded according to inter-frame prediction, the sub-image includes a plurality of blocks, and the sub-image is not the block; A parsing module, used for parsing the bit stream to obtain quantized transform coefficients; A processing module, configured to inverse quantize and inverse transform the quantized transform coefficients to obtain a residual block; A determination module, configured to determine a motion vector of a block in the sub-image; Application modules for: applying a clipping function to sample positions in a reference block when the motion vector points outside the sub-image and when a flag indicates that the sub-image is to be processed as an image; applying an interpolation filter to the result of the clipping function to obtain a predicted sample value; A reconstruction module, used for obtaining a reconstructed image block according to the residual block and the prediction block, wherein the prediction block includes the prediction sample value; wherein the interpolation filter comprises a luma sample bilinear interpolation process, the block comprises a luma sample block, and the predicted sample value comprises a predicted luma sample value; The luma sample bilinear interpolation process receives an input comprising a luma position (xIntL, yIntL) having an integer number of sample units, the luma sample bilinear interpolation process outputs a predicted luma sample value (predSampleLXL), the clipping function is applied to the sample position according to the following: When subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInti=Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i), yInti=Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i), Wherein, subpic_treated_as_pic_flag represents the flag, SubPicIdx represents the index of the sub-image, xInti and yInti represent the position of the clipped sample at index i, SubPicRightBoundaryPos represents the position of the right boundary of the sub-image, SubPicLeftBoundaryPos represents the position of the left boundary of the sub-image, SubPicTopBoundaryPos represents the position of the upper boundary of the sub-image, SubPicBotBoundaryPos represents the position of the lower boundary of the sub-image, and Clip3 represents the clipping function according to the following formula: Where x, y, and z are numeric input values.

7. The decoder according to claim 6, characterized in that The decoder is further configured to perform the method according to claim 2 or 3.

8. A computer program product, characterized in that The computer program product comprises computer executable instructions, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 3.