Selective use of coding tool in video processing

JP2024026117A5Pending Publication Date: 2025-07-03DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023196343
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-12
Filing Date
2023-11-20
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing video coding standards struggle with efficiently changing video resolution without introducing IDR or IRAP pictures, which can disrupt decoding and cause latency issues, especially in adaptive streaming and conferencing applications.

Method used

Adaptive Resolution Conversion (ARC) technology allows for seamless changes in video resolution by resampling reference pictures, enabling efficient switching between representations with different spatial resolutions without the need for IDR or IRAP pictures.

Benefits of technology

ARC reduces latency and improves user experience by allowing dynamic resolution adjustments in video streaming and conferencing, enhancing bandwidth efficiency and maintaining smooth playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a video processing method.SOLUTION: A video processing method includes the steps of determining that use of a coding tool is disabled for the current video block due to the use of a reference picture having dimensions different from the dimensions of a current picture for coding to the coded representation of the current video block in a case of conversion between the current video block of the current picture of the video and the coded representation of the video, and performing the transformation on the basis of the determination.SELECTED DRAWING: Figure 25C
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] Pursuant to the applicable patent laws and / or regulations pursuant to the Paris Convention, this application is timely filed claiming priority to and the benefit of International Patent Application No. PCT / CN2019 / 086487, filed May 11, 2019, and International Patent Application No. PCT / CN2019 / 110905, filed October 12, 2019. For all purposes under law, the entire disclosures of the above two applications are hereby incorporated by reference as part of the disclosure of this application.

[0002] [Technical field] This patent document relates to video processing techniques, devices and systems.

[0003] Despite advances in video compression, digital video still accounts for the largest bandwidth usage in the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands for digital video applications are expected to continue to increase. Summary of the Invention

[0004] Devices, systems, and methods related to digital video processing, such as, for example, adaptive loop filtering for video processing, are described. The described methods may be applicable to both existing video coding standards (e.g., High Efficiency Video Coding (HEVC)) and future video coding standards or codecs (e.g., Versatile Video Coding (VVC)).

[0005] Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes transform coding in addition to temporal prediction. In order to explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and are used in a reference software named the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was launched to work on the VVC standard, which aims for a 50% bitrate reduction compared to HEVC.

[0006] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes deriving one or more motion vector offsets for a conversion between a current video block of a current picture of a video and a coded representation of the video based on one or more resolutions of reference pictures associated with the current video block and a resolution of the current picture, and performing the conversion using the one or more motion vector offsets.

[0007] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes constructing a motion candidate list, in which motion candidates are included in a priority order, for conversion between a current video block of a current picture of a video and a coded representation of the video, whereby a priority of a motion candidate is based on a resolution of a reference picture associated with the motion candidate, and performing the conversion using the motion candidate list.

[0008] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining adaptive loop filter parameters for a current video picture that includes one or more video units based on dimensions of the current video picture, and performing a conversion between the current video picture and a coded representation of the current video picture by filtering the one or more video units according to the parameters of the adaptive loop filter.

[0009] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes applying a luma mapping with chroma scaling (LMCS) process to a current video block of a current picture of a video, the luma mapping with chroma scaling (LMCS) process reshaping luma samples of the current video block between a first region and a second region and scaling chroma residuals in a luma-dependent manner by using luma mapping with chroma scaling (LMCS) parameters associated with corresponding dimensions, and performing a conversion between the current video block and a coded representation of the video.

[0010] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining whether and / or how to enable a coding tool to distribute a current video block of a video to a plurality of sub-partitions according to a rule based on reference picture information of the plurality of sub-partitions for conversion between the current video block and a coded representation of the video, and performing the conversion based on the determination.

[0011] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes, in the case of a conversion between a current video block of a current picture of a video and a coded representation of the video, determining that use of a coding tool is disabled for the current video block due to use of a reference picture having dimensions different from dimensions of the current picture for coding of the current video block into the coded representation, and performing the conversion based on the determination.

[0012] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes generating a predictive block by applying coding tools to a current video block of a current picture of a video based on rules that determine whether and / or how to use reference pictures having dimensions different from dimensions of the current picture, and performing a conversion between the current video block of the video and a coded representation using the predictive block.

[0013] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining, for a current video block of a current picture of a video and a coded representation of the video, whether a coding tool is disabled based on a first resolution of a reference picture associated with one or more reference picture lists and / or a second resolution of a current reference picture used to derive a predictive block for the current video block, and performing the conversion based on the determination.

[0014] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes performing a conversion between a video picture that includes one or more video blocks and a coded representation of the video, at least a portion of the one or more video blocks being coded by referencing a reference picture list for the video picture according to rules, the rules specifying that the reference picture list includes reference pictures having up to K different resolutions, where K is an integer.

[0015] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method, comprising: performing a conversion between N successive video pictures of a video and a coded representation of the video, the N successive video pictures including one or more video blocks that have been coded at a number of different resolutions according to rules, the rules specifying that for the N successive video pictures, at most K different resolutions are allowed, where N and K are integers.

[0016] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method, comprising: performing a conversion between a video including a plurality of pictures and a coded representation of the video, at least some of the plurality of pictures being coded into the coded representation using a plurality of different coded video resolutions, the coded representation conforming to a format rule in which a first coding resolution of a previous frame is changed to a second coding resolution of a subsequent frame sequentially after the previous frame only if the subsequent frame is coded as an intra-coded frame.

[0017] In one representative aspect, the disclosed techniques may be used to provide a video processing method that includes analyzing a coded representation of video to determine that a current video block of a current picture of the video references a reference picture associated with a resolution different from a resolution of the current picture, generating a predictive block for the current video block by converting a bidirectional prediction mode to a unidirectional prediction mode applied to the current video block, and generating the video from the coded representation using the predictive block.

[0018] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes generating a predictive block for a current video block of a current picture of a video by enabling or disabling inter-frame prediction from reference pictures having different resolutions depending on a precision and / or a resolution ratio of a motion vector, and performing a conversion between the current video block and a coded representation of the video using the predictive block.

[0019] In one exemplary aspect, the disclosed techniques may be used to provide a video processing method that includes determining, based on coding characteristics of a current video block of a current picture of a video, whether reference pictures having dimensions different from dimensions of the current picture are allowed for generating a predictive block for the current video block during conversion between the current video block and a coded representation of the video, and performing the conversion in accordance with the determination.

[0020] In yet another representative aspect, the methods described above are embodied in the form of processor executable code and stored in a computer readable program medium.

[0021] In yet another exemplary aspect, a device configured or operable to perform the method described above is disclosed. The device may include a processor programmed to implement the method.

[0022] In yet another representative aspect, a video decoder device is capable of implementing the methods described herein.

[0023] These and other aspects and features of the disclosed technology are explained in detail in the drawings, specification, and claims. [Brief description of the drawings]

[0024] [Figure 1] 1 shows an example of adaptive streaming of two representations of the same content coded at different resolutions. [Diagram 2] 1 shows an example of adaptive streaming of two representations of the same content coded at different resolutions. [Diagram 3] 1 shows several examples of open GOP prediction structures in two representations. [Figure 4]An example of a representation switch at an open GOP position is shown. [Diagram 5] We present an example of the decoding process for a RASL picture by using a resampled reference picture from another bitstream as a reference. [Figure 6A] Here we present an example of 360° streaming that relies on the MCTS-based RWMR viewport. [Figure 6B] Here we present an example of 360° streaming that relies on the MCTS-based RWMR viewport. [Figure 6C] Here we present an example of 360° streaming that relies on the MCTS-based RWMR viewport. [Figure 7] 1 shows an example of collocated subpicture representations with different IRAP intervals and different sizes. [Figure 8] 1 shows an example of a segment received when a change in viewing orientation causes a change in resolution. [Figure 9] Compared to FIG. 6, an example of a viewing orientation change towards a slight upwards angle towards the right cube face is illustrated. [Figure 10] One example of an implementation that presents subpicture representations for two subpicture positions is shown. [Figure 11] 1 illustrates the implementation of the ARC encoder. [Figure 12] Illustrates the implementation of an ARC decoder. [Figure 13] Here is one example of tile group-based resampling for ARC. [Figure 14] Here is one example of adaptive resolution change. [Figure 15] An example of ATMVP motion prediction for a CU is given below. [Figure 16A] An example of a simplified four-parameter affine motion model is given. [Figure 16B] An example of a simplified six-parameter affine motion model is given. [Figure 17] An example of sub-block-wise affine MVF is shown below. [Figure 18A] Here is an example of a four-parameter affine model. [Figure 18B] Here is an example of a six-parameter affine model. [Figure 19] 4 shows the MVP (motion vector difference) for AF_INTER for the hereditary affine candidates. [Figure 20] We show the MVP for the AF_INTER case for the constructed affine candidates. [Figure 21A] Five adjacent blocks are shown. [Figure 21B] The derivation of the CPMV prediction value is shown. [Figure 22] 1 shows one example of candidate positions for the affine merge mode. [Figure 23A] FIG. 2 is a block diagram of an example hardware platform for implementing the visual media decoding or encoding techniques described herein. [Figure 23B] FIG. 2 is a block diagram of an example hardware platform for implementing the visual media decoding or encoding techniques described herein. [Figure 24A] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 24B] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 24C]1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 24D] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 24E] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25A] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25B] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25C] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25D] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25E] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25F] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25G] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Fig. 25H] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. [Figure 25I] 1 shows a flowchart of an example method for video processing according to some implementations of the disclosed technology. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] The techniques and devices disclosed in this document provide coding tools that use adaptive resolution conversion. AVC and HEVC do not have the ability to change resolution without the need to introduce IDR or intra random access point (IRAP) pictures, and such a capability may be referred to as adaptive resolution change (ARC). There are multiple use cases or application scenarios that would benefit from the ARC feature, including the following:

[0026] - Rate Adaptation in Video Telephony and Conferencing: To adapt the video being coded to changing network conditions, when the network conditions deteriorate and thus the available bandwidth becomes smaller, the encoder can adapt to the changing network conditions by encoding pictures with smaller resolution. Currently, picture resolution changes can only be made after IRAP pictures, which entails several problems. IRAP pictures with reasonable quality are much larger than inter-coded pictures and are correspondingly more complex to decode, which takes time and resources. This is problematic when a resolution change is required by the decoder for load reasons. It may also destroy low-latency buffer conditions, force audio resynchronization and increase the end-to-end delay of the stream at least temporarily. This may lead to a poor user experience.

[0027] - Active speaker change in multi-party video conference: In case of multi-party video conference, it is common that the active speaker is shown with a larger video size than the videos of the remaining conference participants. When the active speaker changes, it is also necessary to adjust the picture resolution for each participant. The need for ARC function becomes more important when such changes of the active speaker occur frequently.

[0028] - Faster start of streaming: In the case of streaming applications, it is common that the application buffers up to a certain length of the decoded pictures before starting to display them. Starting the bitstream at a smaller resolution will allow the application to have enough pictures in the buffer to start displaying sooner.

[0029] Adaptive stream switching in streaming: Dynamic Adaptive Streaming in HTTP (DASH) standard includes a feature called @mediaStreamStructureId. This feature allows switching between different representations at open GOP random access points with non-decodable preceding pictures, such as a CRA picture with an associated RASL picture in HEVC. When two different representations of the same video have different bitrates but the same spatial resolution and the two different representations have the same value of @mediaStreamStructureId, it is possible to perform switching between the two representations of a RASL picture and an associated CRA picture, and to decode the switching at the CRA picture and the associated RASL picture with acceptable quality, thus allowing seamless switching. Using ARC, the @mediaStreamStructureId feature will also be available for switching between DASH representations with different spatial resolutions.

[0030] ARC is also known as dynamic resolution conversion.

[0031] ARC can also be considered as a special case of Reference Picture Resampling (RPR) such as in H.263 Annex P.

[0032] 1.1 Reference Picture Resampling in H.263 Annex P

[0033] This mode describes an algorithm to distort a reference picture before using it for prediction. The algorithm may be useful to resample a reference picture that has a different source format than the picture to be predicted. The algorithm may also be used for global or rotational motion estimation by distorting the shape, size, and position of the reference picture. The syntax includes the resampling algorithm as well as the distortion parameters to be used. For the upsampling and downsampling processes, the simplest level of operation for the reference picture resampling mode is an implicit factor of 4 resampling, since only FIR filters need to be applied. In this case, no additional signaling overhead is required, since its use is understood when the size of the new picture (indicated in the picture header) is different from the size of the previous picture.

[0034] 1.2 Contribution of ARC to VVC

[0035] 1.2.1. JVET-M0135

[0036] The preliminary design of ARC, described below with portions taken from JCTVC-F158, is suggested as a substitute term only to start a discussion. Double brackets will be placed before or after the deleted text.

[0037] 2.2.1.1 Description of the Basic Tools

[0038] The basic tool constraints for supporting ARC are:

[0039] - The spatial resolution may differ from the nominal resolution by a factor of 0.5 when applied to both dimensions. The spatial resolution may be increased or decreased, resulting in scaling ratios of 0.5 and 2.0.

[0040] - The chroma format and aspect ratio of the video format remain unchanged.

[0041] -The cropping area is scaled proportionally to the spatial resolution.

[0042] -The reference pictures are simply rescaled if necessary and inter-frame prediction is applied as usual.

[0043] 2.2.1.2 Scaling Operations

[0044] It is proposed to use simple zero-phase separable downscaling and upscaling filters. It should be noted that these filters are for prediction only, and the decoder may use more advanced scaling for output purposes.

[0045] The following 1:2 downscaling filter is used, which has zero phase and 5 taps:

[0046] That is, (-1,9,16,9,-1) / 32.

[0047] The downsampling points are at even sample positions and are co-located. The same filter is used for luma and chroma.

[0048] In the case of 2:1 upsampling, the additional samples at odd grid positions are generated using the half-pixel motion compensation interpolation filter coefficients in the latest VVC WD.

[0049] Combining upsampling and downsampling does not change the location of the phase or chroma sampling points.

[0050] 2.2.1.2 Resolution Description in Parameter Sets

[0051] The signaling of the picture resolution in the SPS is modified as shown below.

[0052] [Table 1]

[0053] "[[pic_width_in_luma_samples" specifies the width of each decoded picture in units of luma samples. "pic_width_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of "MinCbSizeY".

[0054] "pic_height_in_luma_samples" specifies the height of each decoded picture in units of luma samples. "pic_height_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of "MinCbSizeY.]]".

[0055] "num_pic_size_in_luma_samples_minus1"+1 specifies the number of picture sizes (e.g., width and height) in units of luma samples that may be present in the video sequence being coded.

[0056] "pic_width_in_luma_samples[i]" specifies the width of the i-th picture being decoded in units of luma samples that may be present in the video sequence being coded. "pic_width_in_luma_samples[i]" shall not be equal to 0 and shall be an integer multiple of "MinCbSizeY".

[0057] "pic_height_in_luma_samples[i]" specifies the height of the i-th picture being decoded in units of luma samples that may be present in the video sequence being coded. "pic_height_in_luma_samples[i]" shall not be equal to 0 and shall be an integer multiple of "MinCbSizeY".

[0058] [Table 2]

[0059] "pic_size_idx" specifies the index to the ith picture size in the sequence parameter set. The width of the picture referencing that picture parameter set is "pic_width_in_luma_samples[pic_size_idx]" in luma samples. Similarly, the height of the picture referencing that picture parameter set is "pic_height_in_luma_samples[pic_size_idx]" in luma samples.

[0060] 1.2.2. JVET-M0259

[0061] 1.2.2.1. Background: Subpicture

[0062] The term "sub-picture track" is defined in the Omnidirectional Media Format (OMAF) as follows: A track has a spatial relationship with one or more other tracks and represents a spatial subset of an original video content, which is divided into multiple spatial subsets at the content creator before video encoding. A sub-picture track for HEVC may be constructed by rewriting the parameter set and slice segment headers for a motion-constrained tile set, so that the sub-picture track becomes a self-contained HEVC bitstream. A sub-picture representation may be defined as a DASH representation that carries a sub-picture track.

[0063] JVET-M0261 uses the term sub-picture as the unit of spatial partitioning for VVC, and is summarized as follows:

[0064] 1. A picture is divided into subpictures, tile groups, and tiles.

[0065] 2. A subpicture is a set of rectangles in a tile group, starting from the tile group with "tile_group_address" equal to 0.

[0066] 3. Each sub-picture may refer to its own PPS and therefore may have its own tile partitioning.

[0067] 4. Subpictures are treated like pictures in the decoding process.

[0068] 5. The reference picture for decoding the sub-picture is generated by extracting a region from the reference picture that is located alongside the current sub-picture in the buffer of the picture being decoded. The extracted region shall become the sub-picture being decoded, i.e., inter prediction is performed between sub-pictures of the same size and position in the picture.

[0069] 6. A tile group is a sequence of tiles during a tile raster scan of a subpicture.

[0070] In this contribution it is possible to understand the term sub-picture as defined in JVET-M0261. However, the tracks encapsulating sub-picture sequences defined in JVET-M0261 have very similar characteristics to the sub-picture tracks defined in OMAF, and the examples given below apply to both cases.

[0071] Usage

[0072] 1.2.2.2.1. Adaptive resolution change when streaming

[0073] Requirements for supporting adaptive streaming

[0074] Section 5.13 of MPEG N17074 ("Support for Adaptive Streaming") includes the following requirements for VVC: The standard shall support fast representation switching in the case of adaptive streaming services offering multiple representations of the same content, each with different characteristics (such as spatial resolution or sample bit depth). The standard shall enable the use of efficient prediction structures (such as so-called Open Picture Groups) without compromising the ability to fast and seamlessly switch representations between representations of different characteristics, such as different spatial resolutions.

[0075] Example of an open GOP prediction structure with representation switching

[0076] Content generation for adaptive bitrate streaming involves the generation of different representations, which may have different spatial resolutions. The client requests segments from those representations and may therefore decide at what resolution and at what bitrate to receive the content. At the client, the segments of the different representations are concatenated, decoded and played. The client needs to be able to achieve seamless playback with one decoder instance. A closed GOP structure (starting from the IDR picture) is conventionally used as illustrated in Figure 1, which shows adaptive streaming of two representations of the same content coded at different resolutions.

[0077] The open GOP prediction structure (starting from the CRA picture) allows better compression performance than the respective closed GOP prediction structure. For example, with an IRAP picture interval of 24 pictures, an average bitrate reduction of 5.6% is achieved for the luma Bjontegaard delta bitrate. For convenience, the simulation conditions and results are summarized in section YY.

[0078] An open GOP prediction structure has also been reported to reduce subjectively visible quality pumping.

[0079] A challenge with the use of open GOPs when streaming is that it is not possible to decode RASL pictures with the correct reference pictures after switching representations. This challenge with these representations is illustrated in Figure 2, which shows adaptive streaming of two representations of the same content coded at different resolutions. In Figure 2, segments use either a closed GOP prediction structure or an open GOP prediction structure.

[0080] A segment starting with a CRA picture contains an RASL picture for which at least one reference picture exists in the previous segment. This is illustrated in Figure 3, which shows the open GOP prediction structure of the two representations. In Figure 3, picture 0 in both bitstreams exists in the previous segment and is used as a reference for predicting the RASL picture.

[0081] The representation switch, marked in Fig. 2 using a dashed rectangular shape, is illustrated below in Fig. 4, which shows a representation switch at an open GOP position. It can be observed that the reference picture for the RASL picture ("Picture 0") has not been decoded. As a result, the RASL picture will not be decodable and a gap will occur in the playback of the video.

[0082] However, referring to Section 4, it has been found that it is subjectively acceptable to decode RASL pictures using resampled reference pictures. The procedure of resampling "picture 0" and using the resampled reference pictures as reference pictures for decoding those RASL pictures is illustrated in Figure 5. Figure 5 shows the decoding process of RASL pictures by using resampled reference pictures from the other bitstream as a reference.

[0083] 2.2.2.2.2. Viewport changes in Region-by-Region Mixed Resolution (RWMR) 360° video streaming

[0084] Background: HEVC-based RWMR streaming

[0085] RWMR360° streaming provides increased effective spatial resolution for the viewport. The scheme where tiles covering the viewport originate from 6K (6144x3072) ERP pictures or equivalent CMP resolutions as illustrated in Figure 6, with "4K" decoding capability (HEVC Level 5.1), is included in OMAF Sections D.6.3 and D.6.4 and adopted in the VR Industry Forum guidelines. Such resolution is claimed to be suitable for head-mounted displays using Quad HD (2560x1440) display panels.

[0086] encoding The content is encoded at two spatial resolutions with cubic surface sizes of 1536x1536 and 768x768, respectively. In both bitstreams, a 6x4 tile grid is used, and for each tile position a motion-constrained tile set (MCTS) is coded.

[0087] EncapsulationEach MCTS sequence is encapsulated as a subpicture track and made available as a subpicture representation in DASH.

[0088] Selection of streamed MCTS : Select 12 MCTS from the high-resolution bitstream and extract the complementary 12 MCTS from the low-resolution bitstream. In this way, a hemisphere (180° × 180°) of the streamed content originates from the high-resolution bitstream.

[0089] Merging MCTS into the bitstream to be decoded : The received MCTS of a single time instance is merged into a 1920x4608 coded picture that complies with HEVC Level 5.1. Another option for the merged picture is to have four tile columns of width 768, two tile columns of width 384, and three tile rows of luma samples of height 768, resulting in a picture of 3840x2304 luma samples.

[0090] Figure 6 shows an example of MCTS-based RWMR viewport-dependent 360° streaming: Figure 6a shows an example of a bitstream being coded, Figure 6b shows an example of an MCTS selected for streaming, and Figure 6c shows an example of a picture merged from the MCTS.

[0091] Background: Several representations of different IRAP intervals for viewport-dependent 360° streaming

[0092] When the viewing direction changes in HEVC-based viewport-dependent 360° streaming, a new selection of sub-picture representations may take effect at the next IRAP-aligned segment boundary: the sub-picture representations are merged into the picture being coded for decoding, and thus the VCL NAL unit types are aligned among all selected sub-picture representations.

[0093] Multiple versions of the content may be coded at different IRAP intervals to provide a tradeoff between response time to changes in viewing direction and rate-distortion performance when the viewing direction is stable. This is illustrated in FIG. 7 for one set of collocated subpicture representations for encoding shown in FIG. 6 and discussed in detail in Section 3 of "Separate list for sub-block merge candidates," H. Chen, H. Yang, J. Chen, JVET-L0368, Oct. 2018.

[0094] FIG. 7 shows examples of collocated sub-picture representations with different IRAP intervals and different sizes.

[0095] Figure 8 shows an example where a sub-picture location is initially selected to be received at a lower resolution (384x384). Changing the viewing direction results in a new selection of a sub-picture location to be received at a higher resolution (768x768). In the example of Figure 8, the segments received when the viewing direction changes are located at the beginning of segment 4. In this example, the viewing direction changes, whereby segment 4 is received from a sub-picture representation with a short IRAP interval. Afterwards, the viewing direction is stable and the version with a long IRAP interval can be used from segment 5 onwards.

[0096] Drawbacks of updating all subpicture positions

[0097] Since the viewing direction moves gradually in a typical viewing situation, the resolution changes only among a subset of sub-picture positions in RWMR viewport-dependent streaming. FIG. 9 illustrates the change in viewing direction from FIG. 6 toward the right cube face and slightly upwards. The cube face segments with different resolutions as above are denoted by "C". It can be observed that the resolution has changed in 6 of the 24 cube face segments. However, as explained above, the segments starting from the IRAP picture need to be received for all of the 24 cube face segments in response to the viewing direction change. Updating all sub-picture positions in the segments starting from the IRAP picture is inefficient in terms of streaming rate-distortion performance.

[0098] In addition, the ability to use an open GOP prediction structure with sub-picture representation of RWMR 360° streaming is desirable to improve rate-distortion performance while avoiding the visible picture quality pumping that a closed GOP prediction structure causes.

[0099] Proposed design example

[0100] The following design goals are proposed:

[0101] 1. The VVC design needs to allow merging sub-pictures originating from random access pictures and other sub-pictures originating from non-random access pictures into the same coded picture in a VVC-compliant manner.

[0102] 2. The VVC design should allow the merging of multiple sub-picture representations into a single VVC bitstream while allowing the use of open GOP prediction structures in the sub-picture representations without compromising the ability to rapidly and seamlessly switch representations between sub-picture representations of different characteristics, such as different spatial resolutions.

[0103] An example of the design goal may be illustrated using FIG. 10, where subpicture representations for two subpicture locations are shown. For both subpicture locations, separate versions of the content are coded for each combination of the two resolutions and the two random access intervals. Some of the segments start with an open GOP prediction structure. A change in viewing direction causes a resolution switch for subpicture location 1 at the beginning of segment 4. Since segment 4 starts with a CRA picture associated with a RASL picture, it is necessary to resample those reference pictures of the RASL picture present in segment 3. It is noted that this resampling is applied to subpicture location 1, but the decoded subpictures of some other subpicture locations are not resampled. In this example, the change in viewing direction does not cause a change in resolution for subpicture location 2, and therefore the decoded subpictures of subpicture location 2 are not resampled. In the first picture of segment 4, the segment in sub-picture position 1 contains a sub-picture originating from a CRA picture, while the segment in sub-picture position 2 contains a sub-picture originating from a non-random access picture, implying that merging of these sub-pictures into the picture being coded is allowed by VVC.

[0104] 2.2.2.2.3. Adaptive Resolution Change in Videoconferencing

[0105] JCTVC-F158 primarily proposes adaptive resolution change for video conferencing. The following subsections are copied from JCTVC-F158 and present use cases in which adaptive resolution change is claimed to be useful:

[0106] Seamless network adaptation and error tolerance

[0107] Applications such as video conferencing and streaming over packet networks frequently require that the encoded stream adapts to changing network conditions, especially when the bit rate is too high and data is lost. Such applications typically have a return channel that allows the encoder to detect errors and perform adjustments. The encoder has two main tools at its disposal: bit rate reduction and resolution change, either temporal or spatial. Temporal resolution change may be effectively achieved by coding using hierarchical prediction structures. On the other hand, for best quality, spatial resolution change is necessary, as is part of a well-designed encoder for video communication.

[0108] To change spatial resolution in AVC, an IDR frame must be sent and the stream reset. This change in spatial resolution causes significant problems: a reasonable quality IDR frame is much larger than an inter-coded picture and is correspondingly more complex to decode, and more complex decoding requires time and resources. This becomes problematic when a resolution change is required by the decoder for loading reasons. Also, the resolution change would destroy low-latency buffer conditions, force audio resynchronization, and the end-to-end delay of the stream would increase, at least temporarily. This results in a poor user experience.

[0109] To minimize these problems, IDRs are typically sent at a lower quality, using a similar number of bits to P frames, and take a significant amount of time to restore full quality for a given resolution. To make the delay small enough, the quality can be very low, often resulting in a visible blur before the image is "refocused." In effect, inter-coded frames are of little use in compressed terms: they are just a way to restart the stream.

[0110] Therefore, there is a need for a method in HEVC that allows for varying resolution with minimal impact on the subjective experience, especially in difficult network conditions.

[0111] Fast start

[0112] To reduce delay and reach normal quality more quickly without initially introducing unacceptable image blurring, it would be useful to have a "fast start" mode in which the first frames are sent at a reduced resolution, and then the resolution is increased over the next few frames.

[0113] "Compose" a meeting

[0114] Videoconferencing also often has the feature that the speaker is shown full screen and other participants are shown in a window with a smaller resolution. To efficiently support this feature, a smaller picture is often transmitted at a lower resolution. Then, when that participant becomes the speaker and is shown full screen, the resolution is increased. Sending intraframe-coded frames at this point would cause unpleasant hiccups in the video stream. When speakers change rapidly, this effect can be quite noticeable and unpleasant.

[0115] 2.2.2.3. Proposed Design Goals

[0116] Listed below are the high-level design choices proposed for VVC Version 1.

[0117] 1. We propose to include a reference picture resampling process in VVC Version 1 for the following use cases:

[0118] The use of efficient prediction structures (such as so-called open picture groups) in adaptive streaming, without compromising the ability to quickly and seamlessly switch between representations of different characteristics, such as different spatial resolutions.

[0119] -Adaptation of interactive video content to network conditions with low latency and changing resolution resulting from the application without significant latency or latency changes.

[0120] 2. A VVC design is proposed that allows merging sub-pictures originating from random access pictures and other sub-pictures originating from non-random access pictures into the same coded picture that is VVC compatible. This VVC design is claimed to enable efficient handling of viewing direction changes in mixed-quality and mixed-resolution viewport-adaptive 360° streaming.

[0121] 3. We propose to include a per-subpicture resampling process in VVC Version 1. This resampling process is argued to enable an efficient prediction structure for more efficient handling of viewing direction changes in mixed-resolution viewport-adaptive 360° streaming.

[0122] 2.2.3. JVET-N0048

[0123] The use cases and design goals for Adaptive Resolution Change (ARC) are discussed in detail by JVET-M0259. The following description provides a summary.

[0124] 1. Real-time communication

[0125] JCTVC-F158 was originally designed with the following usage scenarios in mind for adaptive resolution change:

[0126] a. Seamless network adaptation and error resilience (with dynamic adaptive resolution changes)

[0127] b. Fast Start (a gradual increase in resolution at session start or session reset)

[0128] c. "Compose" the meeting (giving speakers greater resolution)

[0129] 2. Adaptive Streaming

[0130] MPEG N17074, clause 5.13 ("Support for adaptive streaming") includes the following requirements for VVC: The standard shall support fast representation switching in the case of adaptive streaming services offering multiple representations of the same content, each with different characteristics (such as spatial resolution or sample bit depth). The standard shall enable the use of efficient prediction structures (such as so-called open picture groups) without compromising the ability to fast and seamlessly switch representations between representations with different characteristics, such as different spatial resolutions.

[0131] JVET-M0259 describes how to meet this requirement by resampling the reference pictures of the preceding picture.

[0132] 3. 360-degree viewport-dependent streaming

[0133] JVET-M0259 describes a method to address this usage scenario by resampling certain independently coded picture regions of a reference picture of a preceding picture.

[0134] This contribution proposes an adaptive resolution coding approach that is claimed to satisfy all the above mentioned use cases and design goals. 360 degree viewport dependent streaming and conferencing "compositions" are addressed by this proposal together with JVET-N0045 (which proposes an independent sub-picture layer).

[0135] Proposed specification text

[0136] Double brackets are placed before and after the deleted text.

[0137] Signaling

[0138] [Table 3]

[0139] "sps_max_rpr" specifies the maximum number of active reference pictures in reference picture list 0 or 1 for any tile group in a CVS that has "pic_width_in_luma_samples" and "pic_height_in_luma_samples" not equal to the "pic_width_in_luma_samples" and "pic_height_in_luma_samples" of the current picture, respectively.

[0140] [Table 4] [Table 5]

[0141] "max_width_in_luma_samples" specifies that it is a bitstream conformance requirement that "pic_width_in_luma_samples" in any active PPS for any picture in a CVS in which this SPS is active is less than or equal to "max_width_in_luma_samples".

[0142] "max_height_in_luma_samples" specifies that it is a bitstream conformance requirement that "pic_height_in_luma_samples" in any active PPS for any picture in a CVS in which this SPS is active is less than or equal to "max_height_in_luma_samples".

[0143] High-Level Decoding Process

[0144] The decoding process operates as follows for the current picture, CurrPic:

[0145] 1. Decoding of NAL units is specified in section 8.2.

[0146] 2. The process in Section 8.3 specifies the following decoding process using the syntax elements in and above the tile group header layer.

[0147] - Picture order count related variables and functions are derived as specified in section 8.3.1. It needs to be called only for the first tile group of a picture.

[0148] At the start of the decoding process for each tile group of a non-IDR picture, the decoding process for reference picture list construction specified in Section 8.3.2 is invoked to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).

[0149] - The decoding process for reference picture marking in section 8.3.3 is invoked and reference pictures may be marked as "unused for reference" or "used for long-term reference". This needs to be invoked only for the first tile group of a picture.

[0150] -For each active reference picture in RefPicList[0] and RefPicList[1] that has "pic_width_in_luma_samples" or "pic_height_in_luma_samples" not equal to "pic_width_in_luma_samples" or "pic_height_in_luma_samples", respectively, of CurrPic, the following applies:

[0151] The resampling process of the -XYZ clause is invoked with an output that has the same reference picture markings and picture order counts as the input.

[0152] - Reference pictures used as input to the resampling process are marked as "unused for reference".

[0153] Furthermore, it is possible to describe the invocation of the decoding process for the coding tree units, scaling, transformation, in-loop filtering, etc.

[0154] After decoding all of the tile groups of the current picture, the currently decoded picture is marked as "used for short-term reference".

[0155] The resampling process

[0156] The SHVC resampling process (HEVC section H.8.1.4.2) is proposed with the following additions:

[0157] …

[0158] When “sps_ref_wraparound_enabled_flag” is equal to 0, the sample values ​​tempArray[n] (n=0,...,7) are derived as follows:

number

[0159] Otherwise, the sample value tempArray[n]( n=0,…,7) is derived as follows:

number

[0160] When “sps_ref_wraparound_enabled_flag” is equal to 0, the sample values ​​tempArray[n] (n=0,...,3) are derived as follows:

number

[0161] Otherwise, the sample values ​​tempArray[n] (n=0,...,3) are derived as follows:

number

[0162] 2.2.4. JVET-N0052

[0163] Adaptive resolution change has been a concept in video compression standards since at least 1996, particularly in H.263+ related proposals for reference picture resampling (RPR, Annex P) and reduced resolution updating (Annex Q). It has recently gained some attention, starting with a Cisco proposal in the JCT-VC era, in the context of VP9 (which is moderately widely deployed today), and more recently in the context of VVC. ARC reduces the number of samples that need to be coded for a given picture, allowing the resulting reference picture to be amplified to a higher resolution if desired.

[0164] ARCs of particular interest are considered in two scenarios.

[0165] (1) Intraframe coded pictures, such as IDR pictures, are often significantly larger than interframe coded pictures. Regardless of the reason, downsampling pictures intended for intraframe coding can provide better input for future prediction, and is clearly advantageous from a rate control point of view, at least in low-latency applications.

[0166] (2) ARC can be useful even for non-intracoded pictures, such as scene transitions that do not have hard transition points, when operating a codec near breaking points, as at least some cable and satellite operators routinely do.

[0167] (3) Is the notion of a fixed resolution generally justifiable, perhaps to look at bits that are too far ahead? With the departure of CRTs and the widespread use of scaling engines in rendering devices, the hard bind between rendering and coding resolution is a thing of the past. It should also be noted that there is available research suggesting that most people are unable to focus on fine details (presumably associated with high resolution) when there is a lot of activity going on in a video sequence, even when that activity is taking place elsewhere spatially. If that is correct and generally accepted, then fine grain resolution variation may be a better rate control mechanism than adaptive QP. This point is addressed here. Removing the notion of a fixed resolution bitstream has myriad system layer and implementation implications that are well known (at least at the level of their existence, if not their detailed nature).

[0168] Technically, ARC may be implemented as reference picture resampling. The implementation of reference picture resampling has two main aspects: the resampling filter and the signaling of resampling information in the bitstream. This document focuses on the latter and mentions the former to the extent that implementation experience exists. Further research into suitable filter designs is encouraged, and Tencent will carefully consider and, where appropriate, support any suggestions in this regard that substantially improve upon the provided Strawman design.

[0169] Overview of Tencent's ARC implementation

[0170] Figures 11 and 12 show the implementation of Tencent's ARC encoder and decoder, respectively. The implementation of the disclosed technology allows the width and height of a picture to be changed at each picture granularity, regardless of the type of picture. In the encoder, the input picture data is downsampled to the picture size selected for the encoding of the current picture. After the first input picture is encoded as an intra-frame coded picture, the decoded picture is stored in the decoded picture buffer. When the resulting picture is downsampled with a different sampling ratio and encoded as an inter-frame coded picture, the reference picture in the DPB is upscaled / downscaled according to the spatial ratio between the reference picture size and the current picture size. In the decoder, the decoded picture is stored in the DPB without resampling. Meanwhile, the reference picture in the DPB is scaled up / down in proportion to the spatial ratio between the currently decoded picture and the reference picture when used for motion compensation. When a picture being decoded is bumped out for display, it is upsampled to the original picture size or the desired output picture size. In the motion estimation / compensation process, the motion vectors are scaled relative to the picture size ratio and the picture order count difference.

[0171] Signaling ARC parameters

[0172] The term "ARC parameters" is used herein as any combination of parameters required for ARC to function. In the simplest case it may be a zoom factor or an index into a table with a defined zoom factor. It may also be a target resolution (e.g., at the granularity of samples or maximum CU size) or an index into a table providing the target resolution, as proposed in JVET-M0135. It may also include filter selectors or filter parameters (down to the filter coefficients) of the upsampling / downsampling filters in use.

[0173] From the beginning, the implementation proposed here is to allow, at least conceptually, different ARC parameters for different parts of a picture. It is proposed that the appropriate syntax structure according to the current VVC draft would be a rectangular Tile Group (TG). Those using scan order TG would be restricted to using only ARC, or to the extent that scan order TG is included in a rectangular TG. This can easily be specified by bitstream constraints.

[0174] Since different TGs may have different ARC parameters, the appropriate location of the ARC parameters is either in the TG header, or in a parameter set for the scope of the TG, i.e., the adaptation parameter set of the current VVC draft, and referenced by the TG header, or a more detailed reference (index) to a table of a higher parameter set. Of these three options, it is currently proposed to use the TG header to code a reference to a table entry containing the ARC parameters, the table being placed in the SPS, and the maximum table value being coded in the (to come) DPS. The zoom factor may be coded directly in the TG header without using a parameter set value. The use of the PPS for reference, as proposed in JVET-M0135, is shown in contrast to the case where per-tile-group signaling of the ARC parameters is a design criterion.

[0175] For the table entries themselves, the following options are available:

[0176] -When coding the downsampling factor, should we code a factor for both dimensions, or code it independently in the X and Y dimensions? This is mostly an (HW-)implementation debate, some will prefer results where the zoom factor in the X dimension is fairly flexible, but the Y dimension is fixed at 1 or has little choice. It is suggested that the syntax is the wrong place to express such constraints, and where they are desirable, prefer constraints expressed as requirements for conformance. In other words, keep the syntax flexible.

[0177] - The coding of the target resolution is as follows: There may possibly be more or less complex constraints on these resolutions relative to the current resolution, expressed in the bitstream conformance requirements.

[0178] - Downsampling per tile group is preferred to enable picture compositing / picture extraction, but is not important from a signaling point of view. If the group makes the unwise decision to allow ARC only at picture granularity, it MAY include a bitstream conformance requirement that all TGs use the same ARC parameters.

[0179] -Control information for ARC, and in the following design, the control information includes reference picture size.

[0180] -Do you need flexibility in your filter design? Are there more than a few code points? If there are more than a few code points, do you feed them into the APS? It has been suggested that in some implementations the bitstream needs to eat overhead if the downsample filter changes and the ALF remains the same.

[0181] At this time, in order to keep the proposed technology consistent and simple (as much as possible), the following is suggested:

[0182] -Fixed filter design

[0183] - The target resolution of the table in the SPS with the bitstream constraint TBD.

[0184] -Min target resolution / max target resolution in DPS to make cap exchange / negotiation easier.

[0185] The resulting syntax is:

[0186] [Table 6]

[0187] "max_pic_width_in_luma_samples" specifies the maximum width of a picture being decoded, in units of luma samples in the bitstream. "max_pic_width_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of "MinCbSizeY". The value of "dec_pic_width_in_luma_samples[i]" cannot be greater than the value of "max_pic_width_in_luma_samples".

[0188] "max_pic_height_in_luma_samples" specifies the maximum height of the picture being decoded, in units of luma samples. "max_pic_height_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of "MinCbSizeY". The value of "dec_pic_height_in_luma_samples[i]" cannot be greater than the value of "max_pic_height_in_luma_samples".

[0189] [Table 7] TIFF2024026117000013.tif14170

[0190] "adaptive_pic_resolution_change_flag" equal to 1 specifies the output picture sizes ("output_pic_width_in_luma_samples", "output_pic_height_in_luma_samples"), an index of the number of picture sizes being decoded ("num_dec_pic_size_in_luma_samples_minus1"), and that at least one decoded picture size ("dec_pic_width_in_luma_samples[i]", "dec_pic_height_in_luma_samples[i]") is present in the SPS. The reference picture sizes ("reference_pic_width_in_luma_samples", "reference_pic_height_in_luma_samples") are presented conditional on the value of "reference_pic_size_present_flag".

[0191] "output_pic_width_in_luma_samples" specifies the width of the output picture in units of luma samples. "output_pic_width_in_luma_samples" shall not be equal to 0.

[0192] "output_pic_height_in_luma_samples" specifies the height of the output picture in luma samples. "output_pic_height_in_luma_samples" shall not be equal to 0.

[0193] "reference_pic_size_present_flag" equal to 1 specifies that "reference_pic_width_in_luma_samples" and "reference_pic_height_in_luma_samples" are present.

[0194] "reference_pic_width_in_luma_samples" specifies the width of the reference picture in units of luma samples. "output_pic_width_in_luma_samples" shall not be equal to 0. If not present, the value of "reference_pic_width_in_luma_samples" is inferred to be equal to "dec_pic_width_in_luma_samples[i]".

[0195] "reference_pic_height_in_luma_samples" specifies the height of the reference picture in units of luma samples. "output_pic_width_in_luma_samples" shall not be equal to 0. If not present, the value of "reference_pic_width_in_luma_samples" is inferred to be equal to "dec_pic_width_in_luma_samples[i]".

[0196] NOTE 1 – The size of the output picture shall be equal to the values ​​of "output_pic_width_in_luma_samples" and "output_pic_height_in_luma_samples". The size of a reference picture shall be equal to the values ​​of "reference_pic_width_in_luma_samples" and "_pic_height_in_luma_samples" when using that reference picture for motion compensation.

[0197] "num_dec_pic_size_in_luma_samples_minus1"+1 specifies the number of picture sizes ("dec_pic_width_in_luma_samples[i]", "dec_pic_height_in_luma_samples[i]") being decoded, in units of luma samples in the video sequence being coded.

[0198] "dec_pic_width_in_luma_samples[i]" specifies the width of the i-th picture size being decoded, in units of luma samples of the video sequence being coded. "dec_pic_width_in_luma_samples[i]" is not equal to 0 and is an integer multiple of MinCbSizeY.

[0199] "dec_pic_height_in_luma_samples[i]" specifies the height of the i-th picture size being decoded, in units of luma samples of the video sequence being coded. "dec_pic_height_in_luma_samples[i]" shall not be equal to 0 and shall be an integer multiple of "MinCbSizeY".

[0200] NOTE 2 - The i-th decoded picture size ("dec_pic_width_in_luma_samples[i]", "dec_pic_height_in_luma_samples[i]") may be equal to the decoded picture size of the decoded picture in the video sequence being coded.

[0201] [Table 8]

[0202] "dec_pic_size_idx" specifies that the width of the picture being decoded shall be equal to "pic_width_in_luma_samples[dec_pic_size_idx]" and the height of the picture being decoded shall be equal to "pic_height _in_luma_samples[dec_pic_size_idx]".

[0203] filter

[0204] The proposed design conceptually includes four different sets of filters: a downsampling filter from the original picture to the input picture, an upsampling filter / downsampling filter that rescales the reference picture for motion estimation / compensation, and an upsampling filter from the picture being decoded to the output picture. The first and last ones may be left as non-standard matters. Within the scope of the standard, the upsampling filter / downsampling filter needs to be configured by explicit signaling with the appropriate parameter set or needs to be predefined.

[0205] For downsampling, our implementation uses the SHVC (SHM ver. 12.4) downsampling filter, which is a 12-tap, 2D separable filter, to resize the reference picture used for motion compensation. In the current implementation, only dyadic sampling is supported. Therefore, the phase of the downsampling filter is set to 0 by default. For upsampling, a 16-phase, 8-tap interpolation filter is used to shift the phase and align the luma and chroma pixel positions with respect to the original positions.

[0206] Tables 9 and 10 show the 8-tap filter coefficients fL[p,x] (p=0,...,15 and x=0,...,7) used in the luma upsampling process and the 4-tap filter coefficients fC[p,x] (p=0,...,15 and x=0,...,3) used in the chroma upsampling process.

[0207] Table 11 shows the 12-tap filter coefficients for the downsampling process. The same filter coefficients are used for both luma and chroma for downsampling.

[0208] [Table 9]

[0209] [Table 10]

[0210] [Table 11]

[0211] It is expected that (possibly) subjective as well as objective benefits can be expected when using filters that adapt to content and / or scaling factors.

[0212] Tile Group Boundaries Explanation

[0213] Perhaps, as is true for many studies related to tile groups, our implementation was not perfect with respect to tile group (TG) based ARC. Our preference is to reconsider the implementation if the description of spatial organization and extraction of multiple sub-pictures in the compressed domain at least gives rise to a working draft. On the other hand, this does not prevent us to extrapolate the results to some extent and adapt our signaling design accordingly.

[0214] For now, the tile group header is the correct place for tile group header syntax such as "dec_pic_size_idx" proposed above, for the reasons already stated. A single ue(v) codepoint, "dec_pic_size_idx", is conditionally present in the tile group header and is used to indicate the ARC parameters employed. In order to match implementations that are ARC only per picture, in the spec-space it is necessary to either code only a single tile group, or to condition bitstream compliance on all of the TG headers for a given coded picture having the same value of "dec_pic_size_idx" (when present).

[0215] It is possible to move the parameter "dec_pic_size_idx" to any header that starts a subpicture. That header can remain a tile group header.

[0216] Beyond these syntactic considerations, some additional work is needed to enable tile-group or subpicture based ARC. Perhaps the most challenging part is how to address the issue of unnecessary samples in pictures where the subpictures are resampled to a smaller size.

[0217] Figure 13 shows an example of tile-group based resampling for ARC. Consider the diagram on the right, which consists of four sub-pictures (probably expressed in the bitstream syntax as four rectangular tile-groups). To the left, the bottom right TG is subsampled to half the size. We need to explain what to do with the samples outside the relevant area marked "half".

[0218] Many (most? all?) previous video coding standards have in common that they did not support spatial extraction of parts of a picture in the compressed domain. This means that each sample of a picture is represented by one or more syntax elements, and each syntax element affects at least one sample. To maintain this, it may be necessary to fill in somehow the area around the samples covered by the downsampled TG, which is labeled "half". Annex P of H.263+ solves this problem by padding, and in fact (under certain strict restrictions) the sample values ​​of the padded samples may be signaled in the bitstream.

[0219] Although perhaps a significant departure from previous assumptions, an alternative needed in either case to support extraction (and construction) of sub-bitstreams based on rectangular portions of a picture would be to relax the current understanding that each sample of the reconstructed picture needs to be represented by something in the picture being coded (even if it is only a skipped block).

[0220] Implementation Considerations, System Impact and Profiles / Levels

[0221] Basic ARCs are proposed to be included in the "Baseline / Main" profile. It is possible to remove them using sub-profiles if they are not needed in a particular application scenario. Several specific limitations may be acceptable. In this respect, it should be noted that the specific H.263+ profile and the "Preferred Mode" (the profile previously set) contain the limitation on Annex P used only as an "implicit factor of 4", i.e. dyadic downsampling in both dimensions. This was enough to get a fast start in videoconferencing (quickly get past the I-frames).

[0222] The design is such that it is possible to do all of the filtering "on the fly" with no or negligible memory bandwidth increase, so there seems to be no need to move ARC to an exotic profile.

[0223] Complex tables etc. may not be used in a meaningful way in capability exchange, as discussed in Marrakesh in the context of JVET-M0135. The number of options is simply too large to allow meaningful interoperability between vendors, assuming offers and answers and similar limited depth handshakes. So far, in reality, to support ARC in a meaningful way in capability exchange scenarios, most interoperability points will need to fall back to a small number of points, e.g. no ARC, ARC with an implicit modulus of 4, full ARC etc. Alternatively, it is possible to specify the required support for all of ARC and leave the bitstream complexity limit to higher level SDOs. This is a strategic discussion that should be made at some point anyway (beyond what has already been discussed in the context of sub-profiling and flags).

[0224] With respect to multiple levels, the basic design principle should be that, as a condition for bitstream conformance, no matter how much upsampling is signaled in the bitstream, the sample count of the picture being upsampled should conform to the level of the bitstream, and all samples should conform to the picture being upsampled and coded. Note that this is not the case in H263+, and certain samples may not exist.

[0225] 2.2.5. JVET-N0118

[0226] The following aspects have been proposed:

[0227] 1. A list of picture resolutions is signaled in the SPS, and an index into the list is signaled in the PPS, specifying the size of each picture.

[0228] 2. For any picture to be output, the decoded picture before resampling is cropped (if necessary) and output, i.e., the resampled picture is not for output but is only a reference for inter-frame prediction.

[0229] 3. Support resampling ratios of 1.5x and 2x. Do not support arbitrary resampling ratios. Additionally, consider the need for more than one or two other resampling ratios.

[0230] 4. Between picture-level resampling and block-level resampling, proponents prefer block-level resampling.

[0231] On the other hand, if picture-level resampling is chosen, the following aspects are proposed:

[0232] i. When resampling a reference picture, both the resampled version of the reference picture and the original resampled version are stored in the DPB, and therefore both will affect the fullness of the DPB.

[0233] ii. A resampled reference picture is marked as “not used for reference” when the corresponding non-resampled reference picture is marked as “not used for reference”.

[0234] iii. The RPL signaling syntax remains unchanged, while the RPL construction process is modified as follows: When a reference picture needs to be included in an RPL entry and a version of the reference picture with the same resolution as the current picture does not exist in the DPB, a picture resampling process is triggered and a resampled version of the reference picture is included in the RPL entry.

[0235] iv. The number of resampled reference pictures that can be present in a DPB should be limited, for example, to 2 or less.

[0236] b. In other cases (when choosing block-level resampling), the following is suggested:

[0237] i. In order to limit the worst case decoder complexity, it is proposed to not allow bidirectional prediction of blocks from reference pictures with a different resolution than the current picture.

[0238] ii. Another option is to combine the two filters and apply the operations simultaneously when resampling and quarter pixel interpolation need to be done.

[0239] 5. Regardless of whether a picture-based or block-based resampling approach is chosen, it is proposed that temporal motion vector scaling be applied where necessary.

[0240] Implementation

[0241] The ARC software is implemented on top of VTM-4.0.1 with the following changes:

[0242] - The list of supported resolutions is signaled in the SPS.

[0243] -Move spatial resolution signaling from SPS to PPS.

[0244] - Implements a picture-based resampling scheme to resample reference pictures. After decoding a picture, the reconstructed picture may be resampled to a different spatial resolution. Both the original reconstructed picture and the resampled reconstructed picture are stored in the DPB and are available for reference by future pictures in decoding order.

[0245] -The implemented resampling filters are based on the filters tested in JCTVC-H0234, as follows:

[0246] -Upsampling filter: 4-tap + / - 1 / 4-phase DCTIF with taps (-4,54,16,-2) / 64

[0247] --Downsampling filter: h11 filter with taps (1,0,-3,0,10,16,10,0,-3,0,1) / 32

[0248] When building the reference picture list for the current picture (i.e. L0 and L1), only use reference pictures that have the same resolution as the current picture. Note that these reference pictures are available in both the original size or the resampled size.

[0249] - It is possible to enable TMVP and ATVVP, but when the original coding resolutions of the current picture and the reference picture are different, disable TMVP and ATMVP for the reference picture.

[0250] - For convenience and ease of starting point software implementation, when outputting pictures the decoder outputs the highest resolution available.

[0251] Picture size and picture output signaling

[0252] 1. List of spatial resolutions of coded pictures in the bitstream

[0253] Currently, all coded pictures in the CVS have the same resolution. Therefore, it is easy to signal only one resolution (i.e., picture width and height) in the SPS. For ARC support, a list of picture resolutions needs to be signaled instead of one resolution. This list is proposed to be signaled in the SPS and an index into the list is signaled in the PPS to specify the size of each individual picture.

[0254] 2. Picture output

[0255] It is proposed that for any picture to be output, the decoded picture before resampling is cropped (if necessary) and then output, i.e. the resampled picture is not a picture for output but a picture for inter-frame prediction reference only. The ARC resampling filter needs to be designed to optimize the use of the resampled picture for inter-frame prediction, and such a filter may not be optimal for picture output / display purposes, while video terminal devices usually optimize the output zoom / scaling functions already implemented.

[0256] 2.2.5.3. About resampling

[0257] The resampling of the picture being decoded may be either picture-based resampling or block-based resampling. For the final ARC design in VVC, block-based resampling is preferred over picture-based resampling. It is recommended that these two approaches be discussed and that JVET decide which of these two approaches should be specified for ARC support in VVC.

[0258] Picture-Based Resampling

[0259] In the case of picture-based resampling for ARC, a picture is resampled only once for a particular resolution, and that particular resolution is stored in the DPB, and at the same time, a non-resampled version of the same picture is also stored in the DPB.

[0260] Adopting picture-based resampling for ARC entails two problems: (1) it requires an additional DPB buffer to store the reference pictures being resampled, and (2) it requires additional memory bandwidth due to the increased operations of reading reference picture data from the DPB and writing reference picture data to the DPB.

[0261] Keeping only one version of a reference picture in the DPB would not be a good idea for picture-based resampling. If only the non-resampled version is stored, a reference picture may need to be resampled multiple times since multiple pictures may refer to the same reference picture. On the other hand, if a reference picture is resampled and only the resampled version is kept, it is better to output the non-resampled picture as explained above, so inverse resampling needs to be applied when the reference picture needs to be output. This is problematic because the resampling process is not a lossless operation. If we take a picture A, then downsample it, and then upsample it to get a picture A' with the same resolution as A, picture A and picture A' will not be the same, picture A' will contain less information than picture A because picture A' has lost some high frequency information during the downsampling and upsampling processes.

[0262] To address the issue of additional DPB buffer and memory bandwidth, it is proposed that the following applies if the ARC design for VVC uses picture-based resampling.

[0263] 1. When a reference picture is resampled, both the resampled version of the reference picture and the original resampled version are stored in the DPB, and therefore both will affect the fullness of the DPB.

[0264] 2. A resampled reference picture is marked as “unused for reference” when the corresponding non-resampled reference picture is marked as “unused for reference”.

[0265] 3. The Reference Picture List (RPL) of each tile group contains a reference picture with the same resolution as the current picture. The RPL signaling syntax does not need to change, but the RPL construction process is modified to ensure what is stated in the previous sentence: when a version of that reference picture with the same resolution as the current picture is not yet available but the reference picture needs to be included in an RPL entry, invoke the picture resampling process and include the resampled version of that reference picture.

[0266] 4. The number of resampled reference pictures that can be present in a DPB needs to be limited, for example to 2 or less.

[0267] Furthermore, to enable the use of temporal MVs (e.g., in merge mode and ATMVP) where the temporal MVs originate from a reference frame with a different resolution than the current frame, we propose to scale the temporal MVs to the current resolution, if necessary.

[0268] Block-Based ARC Resampling

[0269] In the case of block-based resampling for ARC, the reference blocks are resampled if necessary and the picture being resampled is not stored in the DPB.

[0270] The main problem here is the additional decoder complexity, since a block in a reference picture may be referenced multiple times by multiple blocks in other pictures and by multiple blocks in multiple pictures.

[0271] When a block in a reference picture is referenced by a block in a current picture and the resolutions of the reference picture and the current picture are different, the reference block is resampled by invoking an interpolation filter, so that the reference block has an integer pixel resolution. When the motion vector is within a quarter pixel, the interpolation process is again invoked to obtain a resampled reference block with a quarter pixel resolution. Thus, for each motion compensation operation for the current block from a reference block containing multiple different resolutions, a maximum of two interpolation filtering operations are required instead of one. Without ARC support, a maximum of one interpolation filtering operation (i.e., generating a reference block with a quarter pixel resolution) is only required.

[0272] To limit the worst case complexity, it is proposed that the following applies when the ARC design for VVC uses block-based resampling.

[0273] Bidirectional prediction of blocks from reference pictures with a different resolution than the current picture is not allowed.

[0274] More precisely, the constraint is: if a current block blkA in a current picture picA references a reference block blkB in a reference picture picB, then block blkA shall be a unidirectionally predicted block when picA and picB have different resolutions.

[0275] This constraint limits the worst case number of interpolation operations required to decode a block to 2. If a block references blocks from pictures of different resolution, the number of interpolation operations required is 2 as explained above. Since the number of interpolation operations is also two (i.e., one for each quarter-pixel resolution acquisition for each reference block), this constraint is the same as when a block references reference blocks from pictures of the same resolution and is coded as a bidirectionally predicted block.

[0276] To simplify the implementation, if the ARC design of VVC uses block-based resampling, another variant is proposed as follows:

[0277] - When the reference frame and the current frame have different resolutions, the corresponding position of each pixel of the predicted value is calculated first, and then the interpolation is applied only once, i.e., two interpolation operations (one for resampling and one for quarter-pixel interpolation) are combined into only one interpolation operation. The sub-pixel interpolation filter in the current VVC can be reused, but in this case the granularity of the interpolation needs to be enlarged, but the number of interpolation operations is reduced from two to one.

[0278] - To enable the use of temporal MVs (e.g. in merge mode and ATMVP), we propose to scale the temporal MVs to the current resolution if necessary when the temporal MVs originate from a reference frame at a different resolution than the current frame.

[0279] Resampling Ratio

[0280] In JVET-M0135, to initiate the discussion of ARC, it is proposed to consider only resampling ratios of 2x (meaning 2x2 for upsampling and 1 / 2x1 / 2 for downsampling) for the starting point of ARC. Further discussion on this topic after the Marrakech conference has found that supporting only a 2x resampling ratio is very limited, since in some cases it is more beneficial to have a smaller difference between the resampled and non-resampled resolution.

[0281] Although it would be desirable to support arbitrary resampling ratios, this appears to be difficult, since the number of resampling filters that would have to be defined and implemented to support arbitrary resampling ratios would appear to be too large, imposing a large burden on decoder implementations.

[0282] One or more small resampling ratios need to be supported, but it is suggested that at least 1.5x and 2x resampling ratios, and no arbitrary resampling ratios be supported.

[0283] 2.2.5.4. Maximum DPB Buffer Size and Buffer Fullness

[0284] In ARC, a DPB may contain multiple decoded pictures of different spatial resolutions in the same CVS, and for DPB management and related aspects, counting DPB size and fullness in units of decoded pictures no longer works.

[0285] Below is a discussion of some specific aspects that need to be addressed if ARC is supported and possible solutions in the final VVC standard.

[0286] 1. Instead of using the value of PicSizeInSamplesY (i.e., PicSizeInSamplesY=pic_width_in_luma_samples*pic_height_in_luma_samples) to derive MaxDpbSize (i.e., the maximum number of reference pictures that may be present in the DPB), the derivation of MaxDpbSize is based on the value of MinPicSizeInSamplesY, where MinPicSizeInSampleY is defined as follows:

number

[0287] 2. Each decoded picture is associated with a value called PictureSizeUnit, which is an integer value that specifies how large the decoded picture size is relative to MinPicSizeInSampleY. The definition of PictureSizeUnit depends on the resampling ratios supported by VVC's ARC.

[0288] For example, if the ARC only supports a resampling ratio of 2, then PictureSizeUnit is defined as follows:

[0289] The picture being decoded that has the smallest resolution in the bitstream is associated with a PictureSizeUnit of 1.

[0290] - A picture being decoded with a resolution of 2x2, the minimum resolution of the bitstream, is associated with a PictureSizeUnit of 4 (i.e., 1x4).

[0291] In another example, if the ARC supports both resampling ratios 1.5 and 2, then PictureSizeUnit is defined as follows:

[0292] The picture being decoded with the smallest resolution in the bitstream is associated with a PictureSizeUnit of 4.

[0293] A picture being decoded, having a resolution that is 1.5x1.5 of the minimum resolution of the bitstream, is associated with a PictureSizeUnit of 9 (i.e. 2.25x4).

[0294] A picture being decoded, with a resolution that is 2x2 of the minimum resolution of the bitstream, is associated with a PictureSizeUnit of 16 (i.e., 4x4).

[0295] For other resampling ratios supported by ARC, you will need to determine the value of PictureSizeUnit for the respective picture size using the same principles as in the example above.

[0296] 3. Let the variable MinPictureSizeUnit be the smallest possible value of PictureSizeUnit, i.e. if the ARC only supports a resampling ratio of 2 then MinPictureSizeUnit is 1, if the ARC supports resampling ratios of 1.5 and 2 then MinPictureSizeUnit is 4, and similarly, the same principle is used to determine the value of MinPictureSizeUnit.

[0297] 4. The value range of sps_max_dec_pic_buffering_minus1[i] is specified to be in the range of 0 to (MinPictureSizeUnit*(MaxDpbSize-1)). The variable MinPictureSizeUnit is the smallest possible value of PictureSizeUnit.

[0298] 5. The DPB fullness operation is specified based on PictureSizeUnit as follows:

[0299] - HRD is initialized in decoding unit 0, and both CPB and DPB are set to empty (DPB fullness is set to 0).

[0300] - When a DPB is flushed (i.e. all of the pictures are removed from the DPB), the fullness of the DPB is set to 0.

[0301] When a picture is removed from a DPB, the fullness of the DPB is reduced by the value of the PictureSizeUnit associated with the removed picture.

[0302] When a picture is inserted into a DPB, the fullness of the DPB is increased by the value of the PictureSizeUnit associated with the inserted picture.

[0303] 2.2.5.5. Resampling filters

[0304] In the case of the software implementation, the implemented resampling filters were simply taken from previously available filters described in JCTVC-H0234. If other resampling filters offer better performance and / or exhibit lower complexity, then these other resampling filters should be tested and used. Various resampling filters have been proposed that are tested to strike a trade-off between complexity and performance. Such tests may be performed in the CE.

[0305] 2.2.5.6. Other necessary modifications to existing tools

[0306] To support ARC, some modifications and / or additional operations may be required for some of the existing coding tools. For example, in the ARC software-implemented picture-based resampling, for simplicity, TMVP and ATMVP are disabled when the original coding resolutions of the current picture and the reference picture are different.

[0307] 2.2.6. JVET-N0279

[0308] According to "Requirements for Future Video Coding Standards", "For adaptive streaming services offering multiple representations of the same content, the standard shall support fast representation switching, each of the multiple representations having different characteristics (such as spatial resolution or sample bit depth)". In the case of real-time video communication, allowing resolution changes within a coded video sequence without inserting I-pictures not only makes it possible to seamlessly adapt the video data to dynamic channel conditions or user preferences, but also eliminates the beating effect caused by I-pictures. A hypothetical example of adaptive resolution change is shown in Figure 14, where the current picture is predicted from multiple reference pictures of different sizes.

[0309] This contribution proposes a high-level syntax for signaling adaptive resolution changes as well as modifications to the current motion compensated prediction process in the VTM. These modifications are limited to motion vector scaling and sub-pixel position derivation, without any changes to the existing motion compensated interpolator. This allows the existing motion compensated interpolator to be reused and does not require new processing blocks to support adaptive resolution changes, which would incur additional costs.

[0310] 2.2.6.1. Adaptive Resolution Change Signaling

[0311] Double brackets are placed before and after the deleted text.

[0312] [Table 12]

[0313] [["pic_width_in_luma_samples" specifies the width of each decoded picture, in units of luma samples. "pic_width_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0314] "pic_height_in_luma_samples" specifies the height of each decoded picture in units of luma samples. "pic_height_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0315] "max_pic_width_in_luma_samples" specifies the maximum width of the picture being decoded that refers to the SPS in units of luma samples. "max_pic_width_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0316] "max_pic_height_in_luma_samples" specifies the maximum height of a picture being decoded that references an SPS, in units of luma samples. "max_pic_height_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0317] [Table 13]

[0318] "pic_size_different_from_max_flag" equal to 1 specifies that the PPS signals a picture width or height different from the "max_pic_width_in_luma_samples" and "max_pic_height_in_luma_sample" of the referenced SPS. Specifies that "pic_width_in_luma_samples" and "pic_height_in_luma_sample" are the same as "max_pic_width_in_luma_samples" and "max_pic_height_in_luma_sample" of the referenced SPS.

[0319] "pic_width_in_luma_samples" specifies the width of each decoded picture in units of luma samples. "pic_width_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. When "pic_width_in_luma_samples" is not present, it is inferred to be equal to "max_pic_width_in_luma_samples".

[0320] "pic_height_in_luma_samples" specifies the height of each decoded picture, in units of luma samples. "pic_height_in_luma_samples" shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. When "pic_height_in_luma_samples" is not present, it is inferred to be equal to "max_pic_height_in_luma_samples".

[0321] It is a bitstream conformance requirement that the horizontal and vertical scaling ratios for any active reference picture are in the range of 1 / 8 to 2. The scaling ratios are defined as follows:

[0322] horizontal_scaling_ratio=((reference_pic_width_in_luma_samples<<14)+(pic_width_in_luma_samples / 2)) / pic_width_in_luma_samples.

[0323] vertical_scaling_ratio=((reference_pic_height_in_luma_samples<<14)+(pic_height_in_luma_samples / 2)) / pic_height_in_luma_samples.

[0324] [Table 14]

[0325] Reference Picture Scaling Process

[0326] When changing resolution in CVS, a picture may have a size different from that of one or more of its reference pictures. This proposal normalizes all of the motion vectors to the current picture grid, not to the corresponding reference picture grids of those reference pictures. It is argued that this normalization is beneficial to maintain design consistency and make resolution change transparent to the motion vector prediction process. Otherwise, due to different scales, it is impossible to directly use adjacent motion vectors that point to multiple reference pictures with different sizes to perform spatial motion vector prediction.

[0327] When a change in resolution occurs, both the motion vectors and the reference blocks need to be scaled during motion compensation prediction. The scaling range is limited to [1 / 8,2], upscaling is limited to 1:8, and downscaling is limited to 2:1. It should be noted that upscaling refers to the case where the reference picture is smaller than the current picture, and downscaling refers to the case where the reference picture is larger than the current picture. The following section will explain the scaling process in more detail.

[0328] Luminance Block

[0329] The scaling coefficients and their fixed-point representations are defined as follows:

number

[0330] The scaling process involves two parts.

[0331] 1. Map the top-left pixel of the current block to the reference picture.

[0332] 2. Using the horizontal step size and the vertical step size, address the reference position of other pixels in the current block.

[0333] If the coordinates of the upper left corner pixel of the current block are (x, y), then the sub-pixel position (x^', y^') in the reference picture pointed to by the motion vector (mvX, mvY) in 1 / 16 pixel units is defined as follows:

[0334] The horizontal position of the reference picture is x'=((x<<4)+mvX)·hori_scale_fp [Formula 3] And x' is further defined as follows: x'=Sign(x')·((Abs(x')+(1<<7)>>8) [Formula 4] According to, it is reduced to retain only 10 fractional bits.

[0335] Similarly, the vertical position of the reference picture is y'=((y<<4)+mvY)·vert_scale_fp [Formula 5] And y' is further expressed as follows: y'=Sign(y')·((Abs(y')+(1<<7)>>8) [Formula 6] It is reduced accordingly.

[0336] At this point, the reference position of the top-left corner pixel of the current block is at (x^',y^'). Other reference sub-pixel / pixel positions are calculated relative to (x^',y^') using horizontal and vertical step sizes. These step sizes are derived with 1 / 1024 pixel accuracy from the horizontal and vertical scaling factors above as follows: x_step=(hori_scale_fp+8)>>4 [Formula 7] y_step=(vert_scale_fp+8)>>4 [Formula 8]

[0337] As one example, if a pixel in the current block is i columns and j rows away from the upper-left pixel, the horizontal and vertical coordinates of its corresponding reference pixel are derived by the following formulas: x' i =x'+i*x_step [Equation 9] y' j =y'+j*y_step [Formula 10]

[0338] In the case of sub-pixel interpolation, x' i and y' j needs to be divided into a full-pel part and a fractional-pel part.

[0339] The full pixel portion for addressing the reference block is (x' i +32)>>10 [Formula 11] (y' j +32)>>10 [Formula 12] is equal to.

[0340] The fractional pixel portion used to select the interpolation filter is: Δx=((x' i +32)>>6)&15 [Formula 13] Δy=((y' j +32) >>6)&15 [Formula 14]

[0341] Once the full-pel and fractional-pel locations in the reference picture are determined, it is possible to use the existing motion compensated interpolator without any additional modifications: the full-pel locations are used to obtain the reference block partitions from the reference picture, and the fractional-pel locations are used to select the appropriate interpolation filter.

[0342] Saturation Block

[0343] When the chroma format is 4:2:0, the chroma motion vectors have 1 / 32 pixel accuracy. The scaling process of chroma motion vectors and chroma reference blocks is almost the same as that of luma blocks, except for the chroma format-related adjustments.

[0344] The coordinates of the upper-left pixel of the current chroma block are (x c ,y c ), the initial horizontal and vertical positions in the reference chroma picture are x c '=((x c <<5)+mvX)·hori_scale_fp [Formula 15] y c '=((y c<<5)+mvY)·vert_scale_fp [Formula 16] where mvX and mvY are the original luma motion vectors, but now need to be checked with 1 / 32 pixel accuracy.

[0345] x c ' and y c 'teeth, x c '=Sign(x c ')·(Abs(x c ')+(1<<8)>>9) [Formula 17] y c '=Sign(y c ')·(Abs(y c ')+(1<<8)>>9) [Formula 18] It is then further reduced accordingly to maintain 1 / 1024 pixel accuracy.

[0346] Compared to the related luminance equation, the above right shift increases by one extra bit.

[0347] The step size used is the same as for luma. For a chroma pixel at (i,j) relative to the top-left pixel, the horizontal and vertical coordinates of its reference pixel are x c ' i =x c '+i*x_step [Equation 19] y c ' j =y c '+j*y_step [Formula 20] It is derived by:

[0348] In the case of subpixel interpolation, x' i and y' j into a full-pel part and a fractional-pel part.

[0349] The full pixel portion for addressing the reference block is (xc ' i +16)>>10 [Formula 21] (y c ' j +16)>>10 [Formula 22] is equal to.

[0350] The fractional pixel portion used to select the interpolation filter is Δx=((x c ' i +16)>>5)&31 [Formula 23] Δy=((y c ' j +16)>>5)&31 [Formula 24] is equal to.

[0351] Interacting with other coding tools

[0352] Due to the excessive complexity and memory bandwidth associated with the interaction between some coding tools and the scaling of reference pictures, it is recommended to add the following constraints to the VVC standard:

[0353] When "tile_group_temporal_mvp_enabled_flag" is equal to 1, the current picture and its collocated picture must have the same size.

[0354] When resolution changes are allowed within a sequence, the decoder's motion vector refinement needs to be turned off.

[0355] When resolution changes are allowed within a sequence, "sps_bdof_enabled_flag" should be equal to 0.

[0356] 2.3. Coding Tree Block (CTB)-based Adaptive Loop Filter (ALF) in JVET-N0415

[0357] Slice-level time filters

[0358] VTM4 employs Adaptive Parameter Sets (APS). Each APS contains ALF filters sent by a set of signaling, and up to 32 APSs are supported. In the proposal, slice-level temporal filters are tested. Tile groups can reuse ALF information from APSs to reduce overhead. APSs are updated as a first-in-first-out (FIFO) buffer.

[0359] Coding Tree Block (CTB) based Adaptive Loop Filter (ALF)

[0360] For the luma component, when applying the ALF to the luma CTB, it is indicated whether to select 16 fixed filter sets, 5 temporal filter sets, or one filter set sent by signaling. Only the filter set index is sent by signaling. For one slice, only one new set of 25 filters may be sent by signaling. When a new set is sent by signaling for a slice, all of the luma CTBs in the same slice share the set. The fixed filter set may be used to predict a new slice-level filter set and may also be used as a candidate filter set for the luma CTB. The number of filters is 64 in total.

[0361] For the chroma component, when applying the ALF to the chroma CTB, if a new filter is signaled for a slice, the new filter is used for that CTB, otherwise the latest temporal chroma filter that satisfies the temporal scalability constraints is applied.

[0362] As a slice-level temporal filter, the APS is updated as a first-in-first-out (FIFO) buffer.

[0363] 2.4 Alternative Temporal Motion Vector Prediction (also known as Sub-block-based Temporal Merging Candidates in VVC)

[0364] In the alternative temporal motion vector prediction (ATMVP) method, the motion vector temporal motion vector prediction (TMVP) is modified by taking multiple sets of motion information (including motion vectors and reference indexes) from blocks smaller than the current CU. As shown in Figure 14, a sub-CU is a square NxN block (N is set to 8 by default).

[0365] ATMVP predicts motion vectors of sub-CUs in a CU in two steps. The first step is to identify the corresponding blocks in a reference picture using a so-called time vector. This reference picture is called a motion source picture. The second step is to split the current CU into multiple sub-CUs and obtain the motion vector and reference index of each sub-CU from the blocks corresponding to each sub-CU, as shown in Figure 15, which shows an example of ATMVP motion prediction for a CU.

[0366] In the first step, the reference picture and corresponding block are determined by the motion information of the spatial neighboring blocks of the current CU. To avoid the repetition of the scanning process of the neighboring blocks, a merge candidate from block A0 (the left block) in the merge candidate list of the current CU is used. The first available motion vector from block A0 that refers to the co-located reference picture is set to be the temporal vector. Thus, compared with TMVP, ATMVP can more accurately identify the corresponding block, which is always located at the bottom right or the center with respect to the current CU (sometimes called the co-located block).

[0367] In the second step, the time vector in the motion source picture is used to identify the corresponding block of the sub-CU by adding the time vector to the coordinates of the current CU. For each sub-CU, the motion information of the corresponding block (the smallest motion grid covering the center sample) is used to derive the motion information of the sub-CU. After identifying the motion information of the corresponding N×N block, the motion information is converted to the motion vector and reference index of the current sub-CU, and motion scaling and other procedures are applied, in a manner similar to TMVP in HEVC.

[0368] 2.5. Affine Motion Estimation

[0369] In HEVC, only the translation motion model is applied to motion compensation prediction (MCP). In the real world, many kinds of motions exist, such as zoom in / zoom out, rotation, squinting movement, and other irregular motions. In VVC, simplified affine transformation motion compensation prediction is applied to the four-parameter affine model and the six-parameter affine model. Figures 16a and 16b show the simplified four-parameter affine motion model and the simplified six-parameter affine motion model, respectively. As shown in Figures 16a and 16b, the affine motion field of the block is described by two control point motion vectors (CPMVs) for the four-parameter affine model and three CPMVs for the six-parameter affine model.

[0370] The motion vector field (MVF) of a block is described by the following equation, using the four-parameter affine model in equation (1) (the four parameters are defined as the variables a, b, e, and f) and the six-parameter affine model in equation (2) (the six parameters are defined as the variables a, b, c, d, e, and f):

number

[0371] In the above formula, (mvh0,mvh0) is the motion vector of the control point of the upper left corner, (mvh1,mvh1) is the motion vector of the control point of the upper right corner, (mvh1,mvh1) is the motion vector of the control point of the lower left corner, all three motion vectors are called control point motion vectors (CPMVs), (x,y) represent the coordinates of the representative point relative to the upper left sample in the current block, and (mvh(x,y),mvv(x,y)) is the motion vector derived for the sample located at (x,y). The CP motion vector may be sent by signaling (as in the case of affine AMVP mode) or derived on-the-fly (as in the case of affine merge mode). w and h are the width and height of the current block. In practice, the division is implemented by a right shift using a rounding operation. In VTM, the representative point is defined as the center position of a subblock, for example, when the coordinates of the upper left corner of the subblock relative to the upper left sample in the current block are (xs, ys), the coordinates of the representative point are defined as (xs+2, ys+2). For each subblock (i.e., 4×4 in VTM), the representative point is used to derive the motion vector of the entire subblock.

[0372] To further simplify the motion compensation prediction, we apply subblock-based affine transformation prediction. To derive the motion vector of each M×N subblock (in the current VVC case, both M and N are set to 4), the motion vector of the center sample of each subblock is calculated according to Equation 25 and Equation 26, as shown in Figure 17, and then rounded to 1 / 16 decimal precision. Then, we apply the motion compensation filter for 1 / 16 pixels to generate the prediction of each subblock with the derived motion vector. We introduce the interpolation filter for 1 / 16 pixels in affine mode.

[0373] After the MCP, the high-precision motion vectors of each subblock undergo a rounding process and are stored with the same precision as the regular motion vectors.

[0374] 2.5.1. Signaling Affine Predictions

[0375] Similar to the translational motion model, there are also two modes for signaling the side information resulting from the affine prediction: "AFFINE_INTER" mode and "AFFINE_MERGE" mode.

[0376] AF_INTER mode

[0377] The AF_INTER mode may be applied for CUs whose width and height are both greater than 8. A CU-level affine flag is signaled in the bitstream to indicate whether to use the AF_INTER mode or not.

[0378] In this mode, for each reference picture list (such as List0 or List1), three types of affine motion predictors are used in the following order to build an affine AMVP candidate list, where each candidate contains an estimated CPMV of the current block: The difference between the best CPMV found at the encoder side, such as mv0, mv1, mv2 in Figures 18a and 18b, and the estimated CPMV is signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is also signaled.

[0379] (1) Inherited affine motion predictors

[0380] The test order is similar to that of spatial MVP in HEVC AMVP list construction. First, the left hereditary affine motion predictor is derived from the first block in {A1, A0}, which is affine coded and has the same reference picture as the reference picture of the current block. Next, the hereditary affine motion predictor is derived from the first block in {B1, B0, B2}, which is affine coded and has the same reference picture as the reference picture of the current block. Five blocks A1, A0, B1, B0, B2 are shown in Figure 19.

[0381] When a neighboring block is found to be coded in affine mode, the CPMV of the coding unit that covers the neighboring block is used to derive a predictor of the CPMV of the current block. For example, if A1 is coded in non-affine mode and A0 is coded in 4-parameter affine mode, the left inherited affine MV predictor will be derived from A0. In this case,

number

number

number

[0382] (2) Constructed affine motion estimate

[0383] The constructed affine motion estimates are composed of control-point motion vectors (CPMVs), which are derived from adjacent inter coded blocks with the same reference picture, as shown in Figure 20. If the current affine motion model is 4-parameter affine, the number of CPMVs is 2, otherwise, if the current affine motion model is 6-parameter affine, the number of CPMVs is 3. The top left CPMV

number

number

number

[0384] If the current affine motion model is a four-parameter affine, the constructed affine motion estimate is

number

number

[0385] If the current affine motion model is a 6-parameter affine, the constructed affine motion estimate is

number

number

[0386] No pruning process is applied when inserting the constructed affine motion predictors into the candidate list.

[0387] (3) Normal AMVP motion prediction

[0388] The following procedure is applied until the number of affine motion predictors reaches a maximum.

[0389] (i) Where available,

number

[0390] (ii) Where available,

number

[0391] (iii) Where available,

number

[0392] (iv) Derive affine motion estimates by setting all of the CPMVs equal to HEVC TMVP, if available.

[0393] (v) Derive the affine motion estimates by setting all of the CPMVs equal to zero MV.

[0394]

number

[0395] As shown in Figures 18a and 18b, in the case of AF_INTER mode, when using the 4 / 6 parameter affine mode, 2 / 3 control points are required, and therefore 2 / 3 MVDs (Motion Vector Differences) need to be coded for these control points. In JVET-K0337, the formula:

number

number

[0396] AF_MERGE mode

[0397] When applying CU in AF_MERGE mode, the first block coded by affine mode is obtained from the valid neighboring reconstructed blocks. Figure 21 shows candidates for AF_MERGE. The selection order of candidate blocks is from left, top, top right, bottom left to top left as shown in Figure 21a (indicated by A, B, C, D, E in order). For example, as shown by A0 in Figure 21b, when coding the neighboring bottom left block by affine mode, the control point (CP) motion vector mv0 of the top left corner, top right corner, and bottom left corner of the neighboring CU / PU containing block A is selected. N , mv1 N , and mv2 N Also, take out mv0 N , mv1 N , and mv2 N Based on the current CU / PU's top left corner / top right corner / bottom left corner (only used for 6-parameter affine model), the motion vector mv0 C , mv1 C , and mv2 CIn VTM-2.0, if the current block is affine coded, the sub-block located in the upper left corner (e.g., 4x4 block in VTM) stores mv0, and the sub-block located in the upper right corner stores mv1. If the current block is coded by a 6-parameter affine model, the sub-block located in the lower left corner stores mv2, otherwise (when using a 4-parameter affine model), the LB stores mv2'. The other sub-blocks store the MV used for MC.

[0398] After deriving the CPMVs of current CUs mv0C, mv1C, and mv2C, generate the MVF of the current CU according to the simplified affine motion models (Equation 25) and (Equation 26). To identify whether the current CU is coded by AF_MERGE mode, an affine flag is signaled in the bitstream when at least one neighboring block is coded by affine mode.

[0399] For JVET-L0142 and JVET-L0632, we construct the affine merge candidate list by the following steps.

[0400] (1) Insert genetic affine candidates

[0401] An inherited affine candidate means that the candidate is derived from the affine motion model of its valid neighbor affine coded block. Up to two inherited affine candidates are derived from the affine motion models of neighboring blocks and inserted into the candidate list. For the left predictor, the scan order is {A0,A1}, and for the above predictor, the scan order is {B0,B1,B2}.

[0402] (2) Insert the constructed affine candidates

[0403] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (e.g., 5), the constructed affine candidate is inserted into the candidate list. A constructed affine candidate means that the candidate is constructed by combining the neighboring motion information of each control point.

[0404] (a) Motion information of a control point is first derived from a particular spatial and temporal neighborhood shown in Fig. 22, which shows examples of candidate positions for affine merge mode. CPk (k=1,2,3,4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are spatial positions for predicting CPk (k=1,2,3). T is the temporal position for predicting CP4.

[0405] The coordinates of CP1, CP2, CP3, and CP4 are (0,0), (W,0), (H,0), and (W,H), respectively, where W and H are the width and height of the current block.

[0406] The motion information of each control point is obtained according to the following priority:

[0407] - For CP1, the inspection priority is B2 → B3 → A2. If B2 is available, use B2. Otherwise, if B2 is available, use B3. If both B2 and B3 are not available, use A2. If all three candidates are not available, it is not possible to obtain the motion information of CP1.

[0408] For -CP2, the inspection priority is B1 → B0.

[0409] For -CP3, the inspection priority is A1 → A0.

[0410] For -CP4, use T.

[0411] (b) Second, a combination of the control points is used to construct an affine merge candidate.

[0412] - Motion information of three control points is needed to construct a 6-parameter affine candidate. The three control points may be selected from one of four combinations: ({CP1,CP2,CP4},{CP1,CP2,CP3},{CP2,CP3,CP4},{CP1,CP3,CP4}). The combinations {CP1,CP2,CP3}, {CP2,CP3,CP4}, {CP1,CP3,CP4} will be transformed into a 6-parameter motion model represented by the top-left control point, the top-right control point, and the bottom-left control point.

[0413] - Motion information of two control points is needed to construct a 4-parameter affine candidate. Those two control points may be selected from one of two pairs ({CP1,CP2},{CP1,CP3}). Those two pairs will be transformed into a 4-parameter motion model represented by the top-left control point and the top-right control point.

[0414] -The constructed affine candidate combinations are inserted into the candidate list in the following order: {CP1,CP2,CP3},{CP1,CP2,CP4},{CP1,CP3,CP4},{CP2,CP3,CP4},{CP1,CP2},{CP1,CP3}.

[0415] (i) For each combination, check the reference indexes of list X for each CP, and if the reference indexes are all the same, then the combination has a valid CPMV for list X. If the combination does not have a valid CPMV for both list 0 and list 1, then the combination is marked as invalid. Otherwise, the combination is valid and the CPMV is put into the subblock merge list.

[0416] (3) Padding with zero motion vectors

[0417] If the number of candidates in the affine merge candidate list is less than five, insert a zero motion vector with a zero reference index into the candidate list until the list is full.

[0418] More specifically, for the subblock merging candidate list, it is a four-parameter merging candidate with MV set to (0,0), prediction direction set to unidirectional prediction from list0 (for P slices) and prediction direction set to bidirectional prediction (for B slices).

[0419] Shortcomings of existing implementations

[0420] When applied to VVC, ARC may have the following problems:

[0421] In the case of VVC, it is not clear to apply coding tools such as ALF, Luminance Mapping Chroma Scaling (LMCS), Decoder-Side Motion Vector Refinement (DMVR), Bidirectional Optical Flow (BDOF), Affine Prediction, Triangular Prediction Mode (TPM), Symmetric Motion Vector Difference (SMVD), Merged Motion Vector Difference (MMVD), Inter-frame Intra-frame Prediction (known in VVC as Combined Inter-Picture Merged and Intra-Picture Prediction (CIIP)), Localized Illumination Compensation (LIC), and History-Based Motion Vector Prediction (HMVP).

[0422] Exemplary Method for a Coding Tool Using Adaptive Resolution Conversion - Patent application

[0423] The embodiments of the disclosed technology obviate the shortcomings of existing implementations. The examples of the disclosed technology shown below are described to facilitate understanding of the disclosed technology and should not be construed as limiting the disclosed technology. Unless expressly indicated to the contrary, various features described in the examples can be combined.

[0424] In the following description, SatShift(x,n) is

number

[0425] Shift(x,n) is defined as Shift(x,n)=(x+offset0)>>n.

[0426] In one example, offset0 and / or offset1 are (1<<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.

[0427] Another example is offset0=offset1=((1<<n)> >1)-1 or ((1<<(n-1)))-1.

[0428] Clip3(min,max,x) is

number

[0429] Floor(x) is defined to be the largest integer less than or equal to x.

[0430] Ceil(x) is defined to be the smallest integer greater than or equal to x.

[0431] Log2(x) is defined to be the base 2 logarithm of x.

[0432] Some of the multiple aspects of the implementation of the disclosed technology are described below. 1. The derivation of the MV offset in MMVD / SMVD and / or the refined motion vector in the decoder-side derivation process may depend on the resolution of the reference picture associated with the current block and the resolution of the current picture. a. For example, the second MV offset referring to the second reference picture may be scaled from the first MV offset referring to the first reference picture. The scaling factor may depend on the resolutions of the first reference picture and the second reference picture. 2. The motion candidate list construction process may be constructed according to the reference picture resolution associated with the spatial / temporal / history motion candidates. a. In one example, a merge candidate referring to a reference picture with a higher resolution may have a higher priority than a merge candidate referring to a reference picture with a lower resolution. In the description, when W0 < W1 and H0 < H1, the resolution W0*H0 is lower than the resolution W1*H1. b. For example, a merge candidate referring to a reference picture with a higher resolution may be placed before a merge candidate referring to a reference picture with a lower resolution in the merge candidate list. c. For example, a motion vector referring to a reference picture with a resolution lower than the resolution of the current picture cannot exist in the merge candidate list. b. In one example, whether to update the history buffer (lookup table) and / or how to update the history buffer may depend on the reference picture resolution associated with the decoded motion candidate. i. In one example, when a reference picture associated with a decoded motion candidate is associated with multiple different resolutions, such a motion candidate is not permitted to update its history buffer. 3. It is proposed that it is necessary to fill in the picture using the ALF parameter associated with the corresponding dimension. In one example, ALF parameters signaled in a video unit such as an APS may relate to dimensions of one or more pictures. b. In one example, a video unit such as an APS that signals ALF parameters may be associated with one or more picture dimensions. c. For example, a picture may only apply ALF parameters that are signaled in a video unit, such as an APS, that is associated with the same dimensions. d. The PPS resolution / index / resolution indicator may be signaled in the ALF APS. e. ALF parameters are restricted in that they may only be inherited / predicted from ALF parameters used for pictures with the same resolution. 4. It is proposed that the ALF parameters associated with a first corresponding dimension may be inherited from or predicted from the ALF parameters associated with a second corresponding dimension. In one example, a first corresponding dimension must be the same as a second corresponding dimension. b. In one example, the first corresponding dimension may be different from the second corresponding dimension. 5. It is proposed that samples in a picture need to be reshaped using the corresponding dimensions and associated LMCS parameters. In one example, LMCS parameters signaled in a video unit such as an APS may relate to one or more picture dimensions. b. In one example, a video unit such as an APS that signals LMCS parameters may be associated with one or more picture dimensions. c. For example, a picture may only apply LMCS parameters that are signaled in a video unit such as an APS that is associated with the same dimensions. d. The PPS resolution / index / resolution indicator may be sent by signaling within the LMCS APS. e. LMCS parameters are restricted in that they may only be inherited / predicted from LMCS parameters used for pictures with the same resolution. 6. It is proposed that the LMCS parameters associated with a first corresponding dimension may be inherited from or predicted from the LMCS parameters associated with a second corresponding dimension. In one example, a first corresponding dimension must be the same as a second corresponding dimension. b. In one example, the first corresponding dimension may be different from the second corresponding dimension. 7. Whether and / or how to enable TPM (triangular prediction mode) / GEO (inter-frame prediction using geometric partitioning) or other coding tools capable of distributing a block into two or more sub-partitions may depend on the associated reference picture information of the two or more sub-partitions. In one example, whether and / or how to activate may depend on the resolution of one of the two reference pictures and the resolution of the current picture. i. In one example, such coding tool is disabled if at least one of the two reference pictures is associated with a different resolution compared to the current picture. b. In one example, whether and / or how to enable may depend on whether the resolution of the two reference pictures is the same. i. In one example, if two reference pictures are associated with different resolutions, then such coding tools may be disabled. ii. In one example, such coding tools are disabled if both of the two reference pictures are associated with a different resolution compared to the current picture. iii. Alternatively, such coding tools may still be disabled when two reference pictures are both associated with a different resolution compared to the current picture, but the two reference pictures are associated with the same resolution. iv. Alternatively, coding tool X may be disabled if at least one of the reference pictures has a different resolution than the resolution of the current picture and the reference pictures are associated with a different resolution. c. Alternatively, whether and / or how to enable may also depend on whether the two reference pictures are the same reference picture. d. Alternatively, further, whether and / or how to enable may depend on whether the two reference pictures are in the same reference list. e. Alternatively, such coding tools may always be disabled when RPR (Reference Picture Resampling) is enabled in the slice / picture header / sequence parameter set. 8. It is proposed that if a block references at least one reference picture that has different dimensions compared to the current picture, coding tool X may be disabled for that block. In one example, the information related to coding tool X may not be sent by signaling. b. In one example, the motion information of such blocks may not be inserted into the HMVP table. c. Alternatively, when coding tool X is applied to a block, the block is not allowed to reference a reference picture that has different dimensions compared to the current picture. i. In one example, merging candidates that refer to reference pictures that have different dimensions compared to the current picture may be omitted or not included in the merging candidate list. ii. In one example, reference indices corresponding to reference pictures that have different dimensions compared to the current picture may be skipped or are not allowed to be sent by signaling. d. Alternatively, coding tool X may be applied after scaling the two reference blocks or two reference pictures according to the resolution of the current picture and the resolution of the reference picture. e. Alternatively, coding tool X may be applied after scaling the two MVs or the two MVDs according to the resolution of the current picture and the resolution of the reference picture. f. In one example, for a block (e.g., a bidirectionally predictively coded block or a block with multiple hypotheses from the same reference picture list with multiple different reference pictures or multiple different MVs, or a block with multiple hypotheses from multiple different reference picture lists), whether to disable or enable coding tool X may depend on the resolution of the reference picture list and / or the reference picture associated with the current reference picture. i. In one example, coding tool X may be disabled for one reference picture list but enabled for the other reference picture list. ii. In one example, coding tool X may be disabled for one reference picture but enabled for another reference picture, where the two reference pictures may be from different reference picture lists or the same reference picture list. iii. In one example, for each reference picture list L, the enabling / disabling of the coding tool is determined without regard to reference pictures in other reference picture lists different from list L. 1. In one example, the enabling / disabling of a coding tool may be determined by the reference pictures in list L and the current picture. 2. In one example, if the associated reference picture of list L is different from the current picture, the tool may be disabled for list L. iv. Alternatively, the enabling / disabling of a coding tool is determined by all of the reference pictures and / or the resolution of the current picture. 1. In one example, coding tool X may be disabled if at least one of the reference pictures has a different resolution than the current picture. 2. In one example, coding tool X may still be enabled if at least one of the reference pictures has a different resolution than the current picture, but the reference pictures are associated with the same resolution. 3. In one example, coding tool X may be disabled if at least one of the reference pictures has a different resolution than the current picture and the reference pictures are associated with a different resolution. g. Coding tool X may be one of the following: iii. DMVR iv. BDOF v. Affine prediction vi. Triangular prediction mode vii. SMVD viii. MMVD ix. Inter-frame Intra Prediction in VVC x.LIC xi. HMVP xii. Multiple Transformation Sets (MTS) xiii. Sub-Block Transform (SBT) xiv. PROF and / or other decoder-side movement / prediction refinement methods xv. LFNST (Low Frequency Non-Squaring Transform) xvi. Filtering methods (e.g. deblocking filters / SAO / ALF etc.) xvii. ALF between GEO / TPM / components 9. The reference picture list of a picture cannot contain more than K different resolutions. In one example, K is equal to 2. 10. For N consecutive pictures (in decoding order or display order), it is not possible to allow more than K different resolutions. In one example, N=3 and K=3. b. In one example, N=10 and K=3. c. In one example, it is not possible to allow more than K different resolutions in a GOP. d. In one example, it is not possible to allow more than K different resolutions between two pictures with the same temporal layer identifier (denoted as tid). i. For example, K=3, tid=0. 11. It is possible to allow resolution changes within the picture only. 12. When one or two reference pictures of a block are associated with a different resolution than the current picture, bidirectional prediction may be converted to unidirectional prediction in the decoding process. In one example, predictions from list X that have corresponding reference pictures with a different resolution than the current picture may be discarded. 13. Enabling or disabling inter-frame prediction from multiple reference pictures of different resolutions may depend on the motion vector accuracy and / or the resolution ratio. In one example, if the motion vectors that have been scaled according to the resolution ratio point to integer positions, inter-frame prediction may still be applied. b. In one example, if a motion vector that has been scaled according to the resolution ratio points to a sub-pixel location (e.g., 1 / 4 pixel) that is allowed when not using ARC, inter-frame prediction may still be applied. c. Alternatively, bi-prediction may not be allowed when both of the reference pictures are associated with a different resolution than the current picture. d. Alternatively, bi-directional prediction may be enabled when one reference picture is associated with a different resolution than the current picture and the other reference picture is associated with the same resolution as the current picture. e. Alternatively, unidirectional prediction may not be allowed for a block when the reference picture is associated with a different resolution than the current picture and the dimensions of the block meet certain conditions. 14. A first flag (eg, pic_disable_X_flag) indicating whether coding tool X is disabled or not may be signaled in the picture header. a. Whether the coding tools are enabled for slices / tiles / bricks / subpictures / other video units smaller than pictures may be controlled by this flag in the picture header and / or slice type. b. In one example, when the first flag is true, coding tool X is disabled. i. Alternatively, when the first flag is false, coding tool X is enabled. ii. In one example, coding tool X is enabled / disabled for all of the samples in a picture. c. In one example, the signaling of the first flag may further depend on an SPS / VPS / DPS / PPS syntax element or multiple syntax elements. i. In one example, the signaling of the flag may depend on an override flag for coding tool X in the SPS. ii. Alternatively, a second flag (eg, sps_X_slice_present_flag, etc.) indicating the presence of the first flag in the picture header may also be sent by signaling in the SPS. (1) Alternatively, the second flag may also be conditionally signaled when coding tool X is enabled for a sequence (e.g., when sps_X_enabled_flag is true). (2) Alternatively, further, only the second flag may indicate the presence of the first flag, and the first flag may be sent by signaling in the picture header. d. In one example, the first flag and / or the second flag are coded using one bit. e. Coding tool X may be the following coding tool: i. In one example, coding tool X is PROF. ii. In one example, coding tool X is a DMVR. iii. In one example, coding tool X is BDOF. iv. In one example, coding tool X is an inter-component ALF. v. In one example, coding tool X is GEO. vi. In one example, coding tool X is a TPM. vii. In one example, coding tool X is MTS. 15. Whether a block can refer to a reference picture having different dimensions than the current picture may depend on the width (WB) and / or height (HB) of the block and / or the block prediction mode (i.e., bidirectional prediction or unidirectional prediction). In one example, if WB≧T1 and HB≧T2, eg, T1=T2=8, then the block may refer to a reference picture that has different dimensions than the current picture. b. In one example, if WB*HB≧T, eg, T=64, then the block may refer to a reference picture having different dimensions than the current picture. c. In one example, if Min(WB,HB)≧T, eg, T=8, then the block may refer to a reference picture that has different dimensions than the current picture. d. In one example, if Max(WB,HB)≧T, eg, T=8, then the block may refer to a reference picture having different dimensions than the current picture. e. In one example, if WB≦T1 and HB≦T2, eg, T1=T2=64, then the block may refer to a reference picture having different dimensions than the current picture. f. In one example, if WB*HB≦T, eg, T=4096, then the block may refer to a reference picture that has different dimensions than the current picture. g. In one example, if Min(WB,HB)≦T, eg, T=64, then a block may reference reference pictures with different dimensions relative to the current picture. h. In one example, if Max(WB,HB)≦T, eg, T=64, then the block may refer to a reference picture that has different dimensions than the current picture. i. Alternatively, if WB≦T1 and / or HB≦T2, eg, T1=T2=8, then the block is not allowed to refer to a reference picture that has different dimensions than the current picture. j. Alternatively, if WB≦T1 and / or HB≦T2, eg, T1=T2=8, then the block is not allowed to refer to a reference picture that has different dimensions than the current picture.

[0433] FIG. 23A is a block diagram of a video processing device 2300. The device 2300 may be used to implement one or more of the methods described herein. The device 2300 may be embodied by a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 2300 may include one or more processors 2302, one or more memories 2304, and video processing hardware 2306. The one or more processors 2302 may be configured to implement one or more methods described herein, including but not limited to the methods illustrated in FIG. 24A-FIG. 25I. The memory(s) 2304 may be used to store data and codes used to implement the methods and techniques described herein. The video processing hardware 2306 may be used to implement some of the techniques described herein by hardware circuits.

[0434] FIG. 23B is another example of a block diagram of a video processing system capable of implementing the disclosed techniques. FIG. 23B is a block diagram illustrating an example video processing system 2400 capable of implementing various techniques disclosed herein. Various implementations may include some or all of the components of the system 2400. The system 2400 may include an input 2402 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or coded format. The input 2402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces, such as Ethernet, passive optical network (PON), and wireless interfaces, such as a Wi-Fi interface or a cellular interface.

[0435] The system 2400 may include a coding component 2404, which may implement various coding or encoding methods described herein. The coding component 2404 may reduce the average bit rate of the video from the input 2402 to the output of the coding component 2404 to generate a coded representation of the video. Thus, the coding techniques may be referred to as video compression techniques or video transcoding techniques. The output of the coding component 2404 may be stored or transmitted over a connected communication, as represented by component 2406. The stored or communicated bitstream (or coded) representation of the video received at the input 2402 may be used by component 2408 to generate pixel values ​​or displayable video that is sent to a display interface 2409. The process of generating a video that can be viewed by a user from a bitstream representation may be referred to as video decompression. Furthermore, although certain video processing operations are referred to as "coding" operations or tools, it will be understood that the coding tools or operations are used in an encoder, and that corresponding decoding tools or operations that reverse the results of the coding would be performed by a decoder.

[0436] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, and IDE interfaces, etc. The techniques described herein may be embodied by a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0437] 24A shows a flowchart of an example method for video processing. Referring to FIG. 24A, the method 2410 includes, at step 2412, deriving one or more motion vector offsets based on one or more resolutions of a reference picture associated with the current video block and a resolution of the current picture for conversion between a current video block of a current picture of the video and a coded representation of the video. The method 2410 further includes, at step 2414, performing the conversion using the one or more motion vector offsets.

[0438] 24B shows a flowchart of an example method for video processing. Referring to FIG. 24B, the method 2420 includes, at step 2422, constructing a motion candidate list in which motion candidates are included in a priority order for conversion between a current video block of a current picture of the video and a coded representation of the video, whereby the priority of the motion candidate is based on the resolution of a reference picture associated with the motion candidate. The method 2420 further includes, at step 2424, performing the conversion using the motion candidate list.

[0439] 24C shows a flowchart of an example method for video processing. Referring to FIG. 24C, the method 2430 includes determining adaptive loop filter parameters for a current video picture including one or more video units based on dimensions of the current video picture at step 2432. The method 2430 further includes performing a conversion between the current video picture and a coded representation of the current video picture by filtering the one or more video units according to the adaptive loop filter parameters at step 2434.

[0440] FIG. 24D shows a flowchart of an example method for video processing. Referring to FIG. 24D, the method 2440 includes applying a luma mapping with chroma scaling (LMCS) process to a current video block of a current picture of a video, where the LMCS process reshapes luma samples of the current video block between a first region and a second region, and scales the chroma residual in a luma-dependent manner by using LMCS parameters associated with corresponding dimensions. The method 2440 further includes performing a conversion between the current video block and a coded representation of the video, at step 2444.

[0441] 24E shows a flowchart of an example method for video processing. Referring to FIG. 24E, the method 2450 includes an operation for determining, at step 2452, whether and / or how to enable a coding tool to distribute a current video block to a plurality of sub-partitions according to a rule based on reference picture information of the plurality of sub-partitions for conversion between a current video block of a video and a coded representation of the video. The method 2450 further includes an operation for performing the conversion based on the determination, at step 2454.

[0442] 25A shows a flowchart of an example method for video processing. Referring to FIG. 25A, the method 2510 includes, at step 2512, determining that use of a coding tool is disabled for a current video block due to use of a reference picture having dimensions different from dimensions of the current picture for coding into the coded representation of the current video block for conversion between a current video block of a current picture of a video and a coded representation of the video. The method 2510 further includes, at step 2514, performing the conversion based on the determination.

[0443] FIG. 25B shows a flowchart of an example method for video processing. Referring to FIG. 25B, the method 2520 includes, at step 2522, generating a predictive block by applying coding tools to a current video block of a current picture of the video based on rules that determine whether and / or how to use a reference picture having dimensions different from dimensions of the current picture. The method 2520 further includes, at step 2524, performing a conversion between the current video block of the video and the representation being coded using the predictive block.

[0444] FIG. 25C shows a flowchart of an example method for video processing. Referring to FIG. 25C, the method 2530 includes, at step 2532, determining whether a coding tool is disabled for a current video block based on a first resolution of a reference picture associated with one or more reference picture lists and / or a second resolution of a current reference picture used to derive a predictive block for the current video block for conversion between a current video block of a current picture of a video and a coded representation of the video. The method 2530 further includes, at step 2534, performing the conversion based on the determination.

[0445] 25D shows a flowchart of an example method for video processing. Referring to FIG. 25D, the method 2540 includes, at step 2542, performing a conversion between a video picture including one or more video blocks and a coded representation of the video. In some of the implementations, at least a portion of the one or more video blocks are coded by referencing a reference picture list for the video picture according to a rule that specifies that the reference picture list includes reference pictures having at most K different resolutions, where K is an integer.

[0446] 25E shows a flowchart of an example method for video processing. Referring to FIG. 25E, the method 2550 includes, at step 2552, performing a conversion between N consecutive video pictures of a video and a coded representation of the video. In some implementations, the N consecutive video pictures include one or more video blocks that are coded with different resolutions according to a rule that specifies that up to K different resolutions are allowed for the N consecutive pictures, where N and K are integers.

[0447] FIG. 25F shows a flowchart of an example method for video processing. Referring to FIG. 25F, the method 2560 includes, at step 2562, performing a conversion between a video including a plurality of pictures and a coded representation of the video. In some of the implementations, at least some of the plurality of pictures are coded into a coded representation using a plurality of different coded video resolutions, and the coded representation changes a first coded resolution of a previous frame to a second coded resolution of a next frame that is in order after the previous frame only if the coded representation complies with a format rule that causes the next frame to be coded as an intraframe-coded frame.

[0448] FIG. 25G shows a flowchart of an example method for video processing. Referring to FIG. 25G, the method 2570 includes, at step 2572, analyzing a coded representation of video to determine that a current video block of a current picture of the video references a reference picture associated with a resolution different from that of the current picture. The method 2570 further includes, at step 2574, generating a predictive block for the current video block by converting a bidirectional prediction mode to a unidirectional prediction mode applied to the current video block. The method 2570 further includes, at step 2570, generating video from the coded representation using the predictive block.

[0449] FIG. 25H shows a flowchart of an example method for video processing. Referring to FIG. 25H, the method 2580 includes generating a predictive block for a current video block of a current picture of the video by enabling or disabling inter-frame prediction from reference pictures having different resolutions depending on a precision and / or a resolution ratio of the motion vector at step 2582. The method 2580 further includes performing a conversion between the current video block and a coded representation of the video using the predictive block at step 2584.

[0450] FIG. 25I shows a flowchart of an example method for video processing. Referring to FIG. 25I, the method 2590 includes an operation for determining, at step 2592, based on coding characteristics of a current video block of a current picture of the video, whether a reference picture having dimensions different from those of the current picture is allowed to generate a predictive block for the current video block when converting between the current video block and a coded representation of the video. The method 2590 further includes an operation for performing the conversion according to the determination, at step 2594.

[0451] In some of the embodiments, the video coding method may be implemented using an apparatus implemented by a hardware platform as described with respect to Figure 23A or Figure 23B. It will be appreciated that the disclosed methods and techniques are beneficial for embodiments of video encoders and / or video decoders that are integrated into video processing devices such as smartphones, laptops, desktops, and similar devices by enabling the use of the techniques disclosed herein.

[0452] Some of the embodiments of the disclosed technology include making a judgment or decision to enable a video processing tool or a video processing mode. In one example, when a video processing tool or a video processing mode is enabled, an encoder will use or implement the tool or mode in processing blocks of video, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or a video processing mode is enabled based on a judgment or decision, conversion of blocks of video to a bitstream representation of video will use the video processing tool or video processing mode. In another example, when a video processing tool or a video processing mode is enabled, a decoder will process the bitstream using knowledge that the bitstream has been modified based on the video processing tool or video processing mode. That is, conversion of a bitstream representation of video to blocks of video will be performed using the video processing tool or video processing mode that has already been enabled based on the judgment or decision.

[0453] Some of the embodiments of the disclosed techniques include making a judgment or decision to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder will not use the tool or mode in converting blocks of video into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, a decoder will process the bitstream using the knowledge that the bitstream has not been modified using the video processing tool or mode that has already been disabled based on the judgment or decision.

[0454] In this specification, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during conversion from a pixel representation of a video to a corresponding bitstream representation or vice versa. The bitstream representation of a current video block may correspond to bits that are either located at the same location or distributed at multiple different locations in the bitstream, e.g., as defined by a syntax. For example, a macroblock may be coded using bits in a header and other fields in the bitstream, for error residual values ​​that are transformed and coded. In this specification, a video block may be a logical unit that corresponds to processing operations that are performed, such as, for example, a coding unit, a transform unit, and a prediction unit.

[0455] The first set of sections below describes certain features and aspects of the disclosed techniques described in the previous sections.

[0456] 1. A method for video processing, comprising: deriving a motion vector offset based on a resolution of a reference picture associated with the current video block and a resolution of the current picture upon conversion between the current video block and a bitstream representation of the current video block; and performing a transformation using the motion vector offsets.

[0457] 2. The method according to claim 1, wherein the step of deriving a motion vector offset comprises: deriving a first motion vector offset referring to a first reference picture; and deriving a second motion vector offset that references a second reference picture based on the first motion vector offset.

[0458] 3. The method of claim 1, further comprising the step of performing a motion candidate list construction process for the current video block based on the resolution of a reference picture associated with a spatial motion candidate, a temporal motion candidate, or a historical motion candidate.

[0459] 4. A method according to clause 1, wherein whether or how to update the look-up table depends on the reference picture resolution associated with the motion candidate being decoded.

[0460] 5. The method of claim 1, further comprising the step of performing a filtering operation on the current picture using adaptive loop filter (ALF) parameters associated with the corresponding dimensions.

[0461] 6. The method of claim 5, wherein the ALF parameters include a first ALF parameter associated with a first corresponding dimension and a second ALF parameter associated with a second corresponding dimension, the second ALF parameter being inherited or predicted from the first ALF parameter.

[0462] 7. The method of claim 1, further comprising the step of reshaping samples in the current picture using LMCS (Luminance Mapping with Chroma Scaling) parameters associated with corresponding dimensions.

[0463] 8. The method of claim 7, wherein the LMCS parameters include a first LMCS parameter associated with a first corresponding dimension and a second LMCS parameter associated with a second corresponding dimension, the second LMCS parameter being inherited or predicted from the first LMCS parameter.

[0464] 9. The method according to clause 5 or 7, wherein the ALF parameters or LMCS parameters sent by signaling in the video unit are associated with dimensions of one or more pictures.

[0465] 10. The method of claim 1, further comprising disabling a coding tool for a current video block when the current video block references at least one reference picture having different dimensions from the current picture.

[0466] 11. The method of claim 10, further comprising skipping or omitting merging candidates that reference reference pictures having multiple different dimensions from the dimensions of the current picture.

[0467] 12. The method according to clause 1, further comprising the step of applying a coding tool after scaling two reference blocks or two reference pictures based on the resolution of the reference picture and the resolution of the current picture, or after scaling two MVs or MVDs (motion vector differences) based on the resolution of the reference picture and the resolution of the current picture.

[0468] 13. The method according to clause 1, wherein the current picture does not include more than K different resolutions, where K is a natural number.

[0469] 14. The method according to clause 13, wherein K different resolutions are allowed for N consecutive pictures, where N is a natural number.

[0470] 15. The method of claim 1, further comprising applying a resolution change to a current picture, the current picture being an intraframe coded picture.

[0471] 16. Provided is a method as described in clause 1, further comprising a step of converting bidirectional prediction to unidirectional prediction when one or two reference pictures of the current video block have a resolution different from the resolution of the current picture.

[0472] 17. The method according to clause 1, further comprising a step of enabling or disabling inter-frame prediction from a reference picture at a plurality of different resolutions depending on at least one of motion vector accuracy or resolution ratio between a current block dimension and a reference block dimension.

[0473] 18. The method according to clause 1, further comprising applying bidirectional prediction depending on whether the reference pictures or one reference picture has a resolution different from the resolution of the current picture.

[0474] 19. The method of claim 1, wherein whether a current video block references a reference picture having dimensions different from the dimensions of the current picture depends on at least one of the size of the current video block or the block prediction mode.

[0475] 20. The method of claim 1, wherein performing the transformation includes generating the current video block from a bitstream representation.

[0476] 21. The method of claim 1, wherein performing the transformation includes generating a bitstream representation from the current video block.

[0477] 22. An apparatus in a video system including a non-transitory memory having instructions and a processor, the instructions, when executed by the processor, cause the processor to implement a method as described in any one of clauses 1 to 21.

[0478] 23. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any one of clauses 1 to 21.

[0479] The second set of sections below includes, for example, example implementations 1 through 7, and describes certain features and aspects of the disclosed techniques described in the previous sections.

[0480] 1. A video processing method comprising: deriving one or more motion vector offsets based on one or more resolutions of reference pictures associated with the current video block and a resolution of the current picture for a conversion between a current video block of a current picture of the video and a coded representation of the video; and performing a transformation using the one or more motion vector offsets.

[0481] 2. A method according to clause 1, wherein the one or more motion vector offsets correspond to motion vector offsets in a merge with a motion vector difference (MMVD), the motion vector difference (MMVD) comprising a motion vector representation including a distance index defining the distance between two motion candidates or a symmetric motion vector difference (SMVD) that processes the motion vector difference symmetrically.

[0482] 3. The method of claim 1, wherein the one or more motion vector offsets correspond to a refined motion vector used in a decoder-side derivation process.

[0483] 4. The method according to claim 1, wherein the step of deriving one or more motion vector offsets comprises: deriving a first motion vector offset referring to a first reference picture; and deriving a second motion vector offset that references a second reference picture based on the first motion vector offset.

[0484] 5. The method according to any one of clauses 1 to 4, wherein the one or more motion vector offsets include a first offset (offset0) and a second offset (offset1), and performing the conversion includes:

number

[0485] 6. A method according to any one of clauses 1 to 5, wherein performing the conversion includes generating a coded representation from a video or generating a video from the coded representation.

[0486] 7. A video processing method comprising: building a motion candidate list, in which motion candidates are included in a priority order for conversion between a current video block of a current picture of the video and a coded representation of the video, whereby a priority of the motion candidates is based on a resolution of a reference picture associated with the motion candidate; and performing a transformation using the motion candidate list.

[0487] 8. The method according to clause 7, wherein motion candidates that refer to reference pictures with higher resolution have a higher priority than other motion candidates that refer to other reference pictures with lower resolution.

[0488] 9. The method according to clause 7, wherein a motion candidate that references a higher resolution reference picture is placed in a motion candidate list before other merge candidates that reference other lower resolution reference pictures.

[0489] 10. The method according to clause 7, wherein the step of constructing a motion candidate list is performed so as to not include motion candidates that refer to reference pictures having a lower resolution than the resolution of the current picture containing the current video block.

[0490] 11. The method of claim 7, wherein whether and / or how to update the lookup table depends on the resolution of the reference picture associated with the motion candidate being decoded.

[0491] 12. In the method described in section 11, a method is provided in which, when a reference picture is associated with a motion candidate being decoded and has a resolution different from the resolution of a current picture that includes a current video block, motion candidates from the reference picture are not allowed to update the lookup table.

[0492] 13. The method according to any one of clauses 7 to 12, wherein the motion candidates are temporal motion candidates.

[0493] 14. The method according to any one of clauses 7 to 12, wherein the motion candidates are spatial motion candidates.

[0494] 15. The method according to any one of clauses 7 to 12, wherein the motion candidates are history-based motion candidates.

[0495] 16. Provided is a method according to any one of clauses 7 to 15, wherein performing the conversion includes generating a coded representation from the video or generating a video from the coded representation.

[0496] 17. A video processing method comprising: determining adaptive loop filter parameters for a current video picture comprising one or more video units based on dimensions of the current video picture; and performing a conversion between the current video picture and a coded representation of the current video picture by filtering one or more video units according to parameters of an adaptive loop filter.

[0497] 18. The method according to clause 17, wherein the parameters are signaled within a video unit and relate to dimensions of one or more pictures.

[0498] 19. The method according to clause 17, wherein the video unit used for transmitting the parameters by signaling is associated with one or more picture dimensions.

[0499] 20. The method according to clause 17, wherein the same dimensions and related parameters are applied to the current picture as are sent by signaling in the video unit.

[0500] 21. The method of claim 17, wherein the coded representation includes a data structure that transmits at least one of the following by signaling: a resolution, a picture parameter set (PPS) index, and an indication of the resolution.

[0501] 22. The method according to clause 17, wherein parameters are inherited or predicted from parameters used for other pictures having the same resolution as the resolution of the current picture, is provided.

[0502] 23. The method of claim 17, wherein the parameters include a first set of parameters associated with a first corresponding dimension and a second set of parameters associated with a second corresponding dimension, the second set of parameters being inherited or predicted from the first set of parameters.

[0503] 24. The method of claim 23, wherein the first corresponding dimension is the same as the second corresponding dimension.

[0504] 25. The method of claim 23, wherein the first corresponding dimension is different from the second corresponding dimension.

[0505] 26. Provided is a method according to any one of clauses 17 to 25, wherein performing the conversion includes generating a coded representation from the video or generating a video from the coded representation.

[0506] 27. A video processing method comprising: applying a luma mapping with chroma scaling (LMCS) process to a current video block of a current picture of the video, in which the luma mapping with chroma scaling (LMCS) process reshapes luma samples of the current video block between a first region and a second region, and scales chroma residuals in a luma-dependent manner by using luma mapping with chroma scaling (LMCS) parameters associated with corresponding dimensions; and performing a conversion between the current video block and a coded representation of the video.

[0507] 28. The method according to clause 27, wherein the LMCS parameters are signaled within a video unit and are associated with dimensions of one or more pictures.

[0508] 29. The method according to clause 27, wherein the video unit used to transmit the LMCS parameters by signaling is associated with one or more picture dimensions.

[0509] 30. The method according to clause 27, wherein LMCS parameters associated with the same dimensions as those signaled in a video unit are applied to the current picture.

[0510] 31. The method of claim 27, wherein the coded representation includes a data structure that transmits at least one of the following by signaling: a resolution, a picture parameter set (PPS) index, and an indication of the resolution.

[0511] 32. The method according to clause 27, wherein the LMCS parameters are inherited or predicted from parameters used for other pictures having the same resolution as the resolution of the current picture, is provided.

[0512] 33. The method of claim 27, wherein the LMCS parameters include a first LMCS parameter associated with a first corresponding dimension and a second LMCS parameter associated with a second corresponding dimension, the second LMCS parameter being inherited or predicted from the first LMCS parameter.

[0513] 34. The method of claim 33, wherein the first corresponding dimension is the same as the second corresponding dimension.

[0514] 35. The method of claim 33, wherein the first corresponding dimension is different from the second corresponding dimension.

[0515] 36. Provided is a method according to any one of clauses 27 to 35, wherein performing the conversion includes generating a coded representation from the video or generating a video from the coded representation.

[0516] 37. A video processing method comprising: determining whether and / or how to enable a coding tool to distribute, according to a rule, a current video block to a plurality of sub-partitions based on reference picture information of the plurality of sub-partitions for conversion between a current video block of the video and a coded representation of the video; and performing a transformation based on the determination.

[0517] 38. In the method described in section 37, the coding tool provides a method that supports inter-frame prediction using a triangular prediction mode (TPM) in which at least one of the multiple subpartitions is a non-rectangular partition, or a geometric partition (GEO) in which it is possible to distribute video blocks using non-horizontal or non-vertical lines.

[0518] 39. In the method according to clause 37, the rules specify whether and / or how to enable a coding tool based on whether the resolution of one of two reference pictures corresponding to two sub-partitions is the same as or different from the resolution of the current picture that contains the current video block.

[0519] 40. The method according to clause 39, wherein the rules specify that the coding tool is not enabled if at least one of the two reference pictures is associated with a resolution different from the resolution of the current picture.

[0520] 41. The method according to clause 39, wherein the rules specify that the coding tool is not enabled if the two reference pictures are associated with different resolutions.

[0521] 42. The method according to clause 39, wherein the rules specify that the coding tool is not enabled if the two reference pictures are associated with a resolution that is different from the resolution of the current picture.

[0522] 43. The method according to clause 39, wherein the rules specify that the coding tool is not enabled if two reference pictures are associated with the same resolution and that resolution is different from the resolution of the current picture.

[0523] 44. The method according to clause 39, wherein the rules specify that the coding tool is not enabled if at least one of the two reference pictures is associated with a resolution different from the resolution of the current picture.

[0524] 45. In the method according to clause 37, a method is provided in which the rules specify whether and / or how to enable a coding tool based on whether two reference pictures corresponding to two sub-partitions are the same reference picture.

[0525] 46. ​​In the method described in clause 37, a method is provided in which the rules specify whether and / or how to enable a coding tool based on whether two reference pictures corresponding to two sub-partitions are in the same reference list.

[0526] 47. The method according to clause 37, wherein the rules specify that the coding tool is always disabled if reference picture resampling is enabled in a video unit of the video.

[0527] 48. Provided is a method according to any one of clauses 37 to 47, wherein performing the conversion includes generating a coded representation from the video or generating a video from the coded representation.

[0528] 49. An apparatus in a video system including a non-transitory memory having instructions and a processor, the instructions, when executed by the processor, cause the processor to implement a method as described in any one of clauses 1 to 48.

[0529] 50. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any one of clauses 1 to 48.

[0530] The third set of sections below, including, for example, exemplary implementations 8-13 and 15, describe certain features and aspects of the disclosed techniques described in the previous sections.

[0531] 1. A video processing method comprising: In the case of a conversion between a current video block of a current picture of the video and a coded representation of the video, determining that use of a coding tool is disabled for a current video block due to use of a reference picture having dimensions different from dimensions of the current picture for coding of the current video block into the coded representation; and performing a transformation based on the determination.

[0532] 2. The method according to clause 1, wherein information relating to the coding tool is not sent by signaling if use of the coding tool is disabled.

[0533] 3. A method as described in Section 1, wherein motion information of a current video block is not inserted into a history-based motion vector prediction (HMVP) table, and the HMVP table includes one or more entries corresponding to motion information of one or more previously processed blocks.

[0534] 4. In the method according to any one of clauses 1 to 3, the coding tool corresponds to decoder-side motion vector refinement (DMVR), bidirectional optical flow (BDOF), affine prediction, triangular prediction mode, symmetric motion vector difference (SMVD), merge mode using motion vector difference (MMVD), inter-frame prediction and intra-frame prediction, local illumination compensation (LIC), history-based motion vector prediction (HMVP), multiple transform set (MTS), sub-block transform (SBT), prediction refinement using optical flow (PROF), low-frequency non-squared transform (LFNST), or a filtering tool.

[0535] 5. The method according to any one of clauses 1 to 4, wherein the dimensions include at least one of a width and a height of the current picture.

[0536] 6. A method according to any one of clauses 1 to 5, wherein performing the conversion includes generating a coded representation from a video or generating a video from the coded representation.

[0537] 7. A video processing method comprising: generating a predictive block by applying coding tools to a current video block of a current picture of a video based on rules that determine whether and / or how to use a reference picture having dimensions different from dimensions of the current picture; and performing a transformation between a current video block and a coded representation of the video using the predictive block.

[0538] 8. The method of clause 7, wherein the rules specify that no reference pictures are used in generating a predictive block by a coding tool applied to a current video block.

[0539] 9. The method according to clause 8, wherein a merging candidate that references the reference picture is skipped or not included in the merging candidate list.

[0540] 10. The method according to clause 8, wherein a reference index corresponding to a reference picture is skipped or is not allowed to be sent by signaling.

[0541] 11. The method according to clause 7, wherein the rules specify scaling the reference picture according to the resolution of the current picture and the resolution of the reference picture before applying the coding tool.

[0542] 12. The method according to clause 7, wherein the rules specify scaling the motion vector or the motion vector difference according to the resolution of the current picture and the resolution of the reference picture before applying the coding tool.

[0543] 13. A method according to any one of clauses 7 to 12, wherein the coding tool corresponds to decoder-side motion vector refinement (DMVR), bidirectional optical flow (BDOF), affine prediction, triangular prediction mode, symmetric motion vector difference (SMVD), merge mode using motion vector difference (MMVD), inter-frame prediction and intra-frame prediction, local illumination compensation (LIC), history-based motion vector prediction (HMVP), multiple transform set (MTS), sub-block transform (SBT), prediction refinement using optical flow (PROF), low-frequency non-squared transform (LFNST), or a filtering tool.

[0544] 14. Provided is a method according to any one of clauses 7 to 12, wherein performing the conversion includes generating a coded representation from the video or generating a video from the coded representation.

[0545] 15. A video processing method comprising: in the case of a conversion between a current video block of a current picture of the video and a coded representation of the video, determining whether a coding tool is disabled for the current video block based on a first resolution of a reference picture associated with the one or more reference picture lists and / or a second resolution of the current reference picture used to derive a predictive block for the current video block; and performing a transformation based on the determination.

[0546] 16. A method according to clause 15, wherein the determining step determines that for one reference picture list, a coding tool is disabled and for another reference picture list, a coding tool is enabled.

[0547] 17. A method as described in clause 15, wherein the determining step determines that a coding tool is disabled for a first reference picture in the reference picture list, and that a coding tool is enabled for a second reference picture in the reference picture list or another reference picture list.

[0548] 18. A method according to clause 15, wherein the determining step determines whether the coding tool is disabled for a first reference picture list without considering a second reference picture list that is different from the first reference picture list.

[0549] 19. The method according to clause 18, wherein the determining step determines whether the coding tool is disabled based on a reference picture in the first reference picture list and the current picture.

[0550] 20. A method according to clause 18, wherein the determining step determines whether the coding tool is disabled for a first reference picture list when a reference picture associated with the first reference picture list is different from the current picture.

[0551] 21. In the method described in clause 15, the determining step further comprises determining whether the coding tool is disabled based on reference pictures associated with one or more reference picture lists and / or other resolutions of the current picture.

[0552] 22. The method according to clause 21, wherein the determining step determines that the coding tool is disabled if at least one of the reference pictures has a resolution different from the resolution of the current picture.

[0553] 23. The method according to clause 21, wherein the determining step determines that the coding tool is not disabled if at least one of the reference pictures has a resolution different from the resolution of the current picture and the reference pictures are associated with the same resolution.

[0554] 24. The method according to clause 21, wherein the determining step determines that the coding tool is disabled if at least one of the reference pictures has a resolution different from the resolution of the current picture and the reference pictures are associated with resolutions different from each other.

[0555] 25. A method according to any one of clauses 15 to 24, wherein the coding tool corresponds to decoder-side motion vector refinement (DMVR), bidirectional optical flow (BDOF), affine prediction, triangular prediction mode, symmetric motion vector difference (SMVD), merge mode using motion vector difference (MMVD), inter-frame prediction and intra-frame prediction, local illumination compensation (LIC), history-based motion vector prediction (HMVP), multiple transform set (MTS), sub-block transform (SBT), prediction refinement using optical flow (PROF), low-frequency non-squared transform (LFNST), or a filtering tool.

[0556] 26. A method according to any one of clauses 15 to 25, wherein the performing of the conversion includes generating the coded representation from the video or generating the video from the coded representation.

[0557] 27. A video processing method comprising: performing a conversion between a video picture including one or more video blocks and a coded representation of the video; At least a portion of the one or more video blocks are coded by referencing a reference picture list for the video picture according to a rule; The rules provide that a reference picture list contains at most K reference pictures with different resolutions, where K is an integer.

[0558] 28. The method according to paragraph 27, wherein K is equal to 2.

[0559] 29. The method of claim 27 or 28, wherein the performing of the conversion includes generating the coded representation from the video or generating the video from the coded representation.

[0560] 30. A video processing method comprising: performing a conversion between N successive video pictures of a video and a coded representation of the video, the N consecutive video pictures include one or more video blocks coded at a plurality of different resolutions according to a rule; The rules provide a video processing method that specifies that for N consecutive video pictures, at most K different resolutions are allowed, where N and K are integers.

[0561] 31. The method of claim 30, wherein N and K are equal to 3.

[0562] 32. The method of claim 30, wherein N is equal to 10 and K is equal to 3.

[0563] 33. The method of claim 30, wherein K different resolutions are allowed within a group of pictures (GOP) in the representation being coded.

[0564] 34. The method of claim 30, wherein K different resolutions are allowed between two pictures that have the same temporal layer identifier.

[0565] 35. Provided is a method according to any one of clauses 30 to 34, wherein performing the conversion includes generating a coded representation from the video or generating a video from the coded representation.

[0566] 36. A video processing method comprising: performing a conversion between a video comprising a plurality of pictures and a coded representation of the video; At least some of the pictures are coded into a coded representation using a plurality of different coded video resolutions; The coded representation complies with a format rule in which a first coded resolution of a previous frame is changed to a second coded resolution of a next frame sequentially after the previous frame only if the next frame is coded as an intraframe-coded frame.

[0567] 37. The method of claim 36, wherein the order corresponds to a coding order in which the multiple pictures are coded.

[0568] 38. The method of claim 36, wherein the order corresponds to a decoding order in which the multiple pictures are decoded.

[0569] 39. The method of claim 36, wherein the order corresponds to a display order in which the multiple pictures are displayed after decoding.

[0570] 40. The method of any one of clauses 36 to 39, wherein the intra-coded frame is an intra-coded random access point picture.

[0571] 41. The method of any one of clauses 36 to 39, wherein the intra-coded frame is an instantaneous decoding refresh (IDR) frame.

[0572] 42. Provided is a method according to any one of clauses 36 to 41, wherein performing the conversion includes generating a coded representation from the video or generating a video from the coded representation.

[0573] 43. A video processing method comprising: analyzing a coded representation of the video to determine that a current video block of a current picture of the video references a reference picture associated with a resolution that is different from a resolution of the current picture; generating a prediction block for the current video block by converting a bi-directional prediction mode to a uni-directional prediction mode applied to the current video block; generating video from the coded representation using the predictive blocks.

[0574] 44. The method according to clause 43, wherein generating the prediction block includes discarding predictions from the list for reference pictures associated with a resolution different from the resolution of the current picture.

[0575] 45. A video processing method comprising: generating a prediction block for a current video block of a current picture of the video by enabling or disabling inter-frame prediction from a plurality of reference pictures having different resolutions according to a precision and / or a resolution ratio of the motion vector; and performing a transformation between the current video block and a coded representation of the video using the predictive block.

[0576] 46. ​​The method of claim 45, wherein inter-frame prediction is enabled if motion vectors scaled according to the resolution ratio point to integer positions.

[0577] 47. The method of claim 45, wherein inter-frame prediction is enabled if motion vectors scaled according to a resolution ratio point to sub-pixel positions.

[0578] 48. The method of claim 45, wherein bidirectional prediction is disabled if the reference pictures are associated with a resolution different from the resolution of the current picture.

[0579] 49. A method as described in clause 45, wherein bidirectional prediction is enabled when one reference picture is associated with a resolution different from the resolution of the current picture and the other reference picture is associated with the same resolution as the resolution of the current picture.

[0580] 50. The method of claim 45, wherein unidirectional prediction is not allowed when the reference picture is associated with a resolution different from the resolution of the current picture and the block dimensions of the current video block meet certain conditions.

[0581] 51. A method according to any one of clauses 45 to 50, wherein the performing of the conversion includes generating the coded representation from the video or generating the video from the coded representation.

[0582] 52. A video processing method comprising: determining, based on coding characteristics of a current video block of a current picture of the video, whether a reference picture having dimensions different from dimensions of the current picture is allowed for generating a prediction block for the current video block during conversion between the current video block and a coded representation of the video; and performing a transformation according to the determination.

[0583] 53. The method of claim 52, wherein the coding characteristics include a dimension of the current video block and / or a prediction mode of the current video block.

[0584] 54. A method according to clause 52, wherein the determining step determines that a reference picture is allowed if WB≧T1 and HB≧T2, whereby WB and HB correspond to the width and height of the current video block, respectively, and T1 and T2 are positive integers.

[0585] 55. The method according to clause 52, wherein the determining step determines that a reference picture is allowed if WB*HB≧T, whereby WB and HB correspond to the width and height of the current video block, respectively, and T is a positive integer.

[0586] 56. The method according to clause 52, wherein the determining step determines that a reference picture is allowed if Min(WB,HB)≧T, whereby WB and HB correspond to the width and height of the current video block, respectively, and T is a positive integer.

[0587] 57. The method according to clause 52, wherein the determining step determines that a reference picture is allowed if Max(WB,HB)≧T, whereby WB and HB correspond to the width and height of the current video block, respectively, and T is a positive integer.

[0588] 58. Provided is a method according to clause 52, wherein the determining step determines that a reference picture is allowed if WB≦T1 and HB≦T2, whereby WB and HB correspond to the width and height of the current video block, respectively, and T1 and T2 are positive integers.

[0589] 59. The method according to clause 52, wherein the determining step determines that a reference picture having dimensions different from the dimensions of the current video block is allowed if WB*HB≦T, whereby WB and HB correspond to the width and height of the current video block, respectively, and T is a positive integer.

[0590] 60. The method according to clause 52, wherein the determining step determines that a reference picture is allowed if Min(WB,HB)≦T, whereby WB and HB correspond to the width and height of the current video block, respectively, and T is a positive integer.

[0591] 61. The method according to clause 52, wherein the determining step determines that a reference picture is allowed if Max(WB,HB)≦T, whereby WB and HB correspond to the width and height of the current video block, respectively, and T is a positive integer.

[0592] 62. The method according to clause 52, wherein the determining step determines that reference pictures are not allowed if WB≦T1 and / or HB≦T2, whereby WB and HB correspond to the width and height of the current video block, respectively, and T1 and T2 are positive integers.

[0593] 63. A method according to any one of clauses 52 to 62, wherein the performing of the conversion includes generating the coded representation from the video or generating the video from the coded representation.

[0594] 64. An apparatus in a video system, the apparatus including a processor and a non-transitory memory having instructions, the instructions, when executed by the processor, causing the processor to implement a method according to any one of clauses 1 to 63.

[0595] 65. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing a method according to any one of clauses 1 to 63.

[0596] From the foregoing, it will be appreciated that, although several specific embodiments of the presently disclosed technology have been described herein for purposes of illustration, various modifications may be made without departing from the scope of the presently disclosed technology. Accordingly, the presently disclosed technology is not to be limited except as by the appended claims.

[0597] Implementations of the subject matter and functional operations described herein may be realized by various systems, including structures disclosed herein and their structural equivalents, digital electronic circuits, or computer software, firmware, or hardware, or by one or more combinations thereof. Implementations of the subject matter described herein may be realized as one or more computer program products, i.e., as one or more modules of computer program instructions encoded in a tangible computer-readable medium and a non-transitory computer-readable medium for execution by or for controlling the operation of a data processing device. A computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of media affecting a machine-readable propagated signal, or one or more combinations thereof. The term "data processing unit" or "data processing device" encompasses all apparatuses, devices, and machines, including, by way of example, a programmable processor, a computer, or multiple processors or computers for processing data. An apparatus may include, in addition to hardware, code that creates an execution environment for a subject computer program, such as code that constitutes a processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.

[0598] A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed either as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file that contains one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communications network.

[0599] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC).

[0600] Processors suitable for executing computer programs include, by way of example, both general purpose and special purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices, such as, for example, magneto-optical, magneto-optical, or optical disks, for storing data, or is operatively coupled to receive data from or transfer data to such mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by or incorporated in special purpose logic circuitry.

[0601] The specification, together with the drawings, are considered to be illustrative only, with illustrative meaning one example. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. Additionally, the use of "or" is intended to include "and / or" unless the context clearly dictates otherwise.

[0602] Although the present specification contains numerous details, these should not be construed as limitations on the scope of any invention or the scope of the invention as described in the claims, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although multiple features may be described above as operating in a particular combination and initially claimed as such, one or more features of the combinations described in the claims may, in some cases, be carved out of the combination, and the combinations described in the claims may be subject to subcombinations or variations of the subcombinations.

[0603] Similarly, although acts are shown in the figures in a particular order, the acts in those figures should not be understood as requiring such acts to be performed in the particular order shown, or in the order that they occur, or to perform all of the acts shown, to achieve desirable results.Furthermore, the separation of various system components in the embodiments described herein should not be understood as requiring such separation in all of the embodiments.

[0604] Although only a few implementations and examples have been described, other implementations, extensions, and variations may be made based on what has been described and illustrated herein.

Claims

1. 1. A method for processing video data, the method comprising: determining, in the case of a conversion between a current video block of a current picture of a video and a bitstream of the video, that at least one coding tool is disabled for the current video block due to use of a reference picture to perform the conversion, the reference picture having a precision different from a precision of the current picture; and performing the conversion based on the determination. method.

2. the at least one coding tool comprises a decoder-side motion vector refinement (DMVR) tool, a bidirectional optical flow (BDOF) tool, a prediction refinement using optical flow (PROF) tool, a temporal motion vector prediction tool, or a sub-block based temporal motion vector prediction tool; 2. The method of claim 1, wherein in the sub-block based temporal motion vector prediction tool, a block is divided into at least one sub-block, and motion information of the at least one sub-block is derived based on a video region in a collocated picture, the location of the video region being derived based on available certain neighboring blocks.

3. The method of claim 1 , wherein the precision of the current picture is indicated by at least one of a width or a height of the current picture.

4. The step of performing the conversion based on the determination further comprises: calculating a prediction of the current video block; and applying a single interpolation process to the predicted value without performing a resampling process.

5. 2. The method of claim 1, wherein the at least one coding tool includes at least one of affine prediction, triangular prediction mode, symmetric motion vector difference (SMVD), merge mode using motion vector difference (MMVD), inter-frame prediction and intra-frame prediction, local illumination compensation (LIC), history-based motion vector prediction (HMVP), multiple transform set (MTS), sub-block transform (SBT), prediction refinement using optical flow (PROF), low frequency non-squared transform (LFNST), or a filtering tool.

6. The method of claim 1 , wherein information related to the coding tool is not included if the use of the coding tool is disabled.

7. determining whether the coding tool is disabled for the picture based on a first syntax element; The method of claim 1 , wherein the first syntax element indicating whether the coding tool is disabled for the picture is signaled in a picture header.

8. 8. The method of claim 7, wherein for video units smaller than a picture, whether the coding tool is disabled or enabled depends on the first syntax element and / or a slice type of the video.

9. 8. The method of claim 7, wherein if the first syntax element is true, the coding tool is disabled for the picture, and if the first syntax element is false, the coding tool is enabled for the picture.

10. The method of claim 7 , wherein the coding tool is disabled or enabled for all of the samples in the picture.

11. 8. The method of claim 7, wherein whether the first syntax element is signaled in the picture header depends on one or more syntax elements of a sequence parameter set (SPS) associated with a video region of a picture of the video.

12. 12. The method of claim 11 , wherein the one or more syntax elements include a second syntax element indicating a presence of the first syntax element in the picture header and a third syntax element indicating whether the coding tool is enabled for the sequence of video.

13. The method of claim 12, wherein the second syntax element is conditionally signaled in the SPS based on the third syntax element.

14. 13. The method of claim 12, wherein if the second syntax element is true, the first syntax element is sent by signaling in the picture header, and if the coding tool is enabled for the sequence of video, the second syntax element is sent by signaling in the SPS.

15. The method of claim 12 , wherein the first syntax element and / or the second syntax element are coded using one bit.

16. The coding tool includes a first coding tool, the first coding tool comprising: generating an initial prediction sample of a sub-block of the current video block to be coded using an affine mode; 8. The method of claim 7, wherein an optical flow operation is used to generate a final prediction sample for the sub-block by deriving a prediction refinement based on motion vector differences dMvH and / or dMvV, where dMvH and dMvV indicate motion vector differences along horizontal and vertical directions.

17. 8. The method of claim 7, wherein the coding tool includes a second coding tool, the second coding tool being used to refine the motion vector sent by the signaling based on at least one motion vector having an offset relative to the motion vector sent by the signaling.

18. 8. The method of claim 7, wherein the coding tool includes a third coding tool, the third coding tool being used to obtain a refinement of a motion vector based on at least one gradient value corresponding to a sample in a reference block of the current block.

19. 19. The method of claim 1, wherein the performing of the transformation comprises decoding the current video block from the bitstream.

20. 19. The method of claim 1, wherein the performing of the transformation comprises encoding the current video block into the bitstream.

21. 1. An apparatus for processing video data comprising a non-transitory memory having instructions and a processor, the instructions, when executed by the processor, cause the processor to: determining, in the case of a conversion between a current video block of a current picture of a video and a bitstream of the video, that at least one coding tool is disabled for the current video block due to use of a reference picture to perform the conversion, the reference picture having a different precision than a precision of the current picture; performing said conversion based on said determination. Device.

22. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to: determining, in the case of a conversion between a current video block of a current picture of a video and a bitstream of the video, that at least one coding tool is disabled for the current video block due to use of a reference picture to perform the conversion, the reference picture having a different precision than a precision of the current picture; performing said conversion based on said determination. A non-transitory computer-readable storage medium.

23. 1. A method for storing a bitstream of a video, the method comprising: determining, for a current video block of a current picture of the video, that at least one coding tool is disabled for the current video block due to use of a reference picture to generate the bitstream, the reference picture having a different precision than a precision of the current picture; generating the bitstream based on said determination; storing the bitstream in a non-transitory computer readable recording medium; method.