Motion prediction from a temporal block by reference picture resampling

By introducing Adaptive Resolution Conversion (ARC) technology and an open GOP prediction structure, the shortcomings of existing video codec standards in terms of flexible resolution changes are addressed, achieving more efficient coding and improved user experience, and making it suitable for adaptive resolution switching in video conferencing and streaming.

CN113875250BActive Publication Date: 2025-10-24DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080035724.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-12
Filing Date
2020-05-12
Publication Date
2025-10-24
Estimated Expiration
2040-05-12

AI Technical Summary

Technical Problem

Existing video codec standards such as HEVC and H.265 struggle to achieve flexible resolution changes without introducing IDR or intra-frame random access points, resulting in low coding efficiency when network conditions change, unstable quality during speaker switching in video conferences, and excessively long switching delays during streaming, failing to meet the requirements for adaptive resolution conversion.

Method used

Adaptive Resolution Conversion (ARC) technology is introduced, which signals changes in video segment resolution in the bitstream representation and allows for resampling of reference images and open GOP prediction structures, enabling seamless switching between different resolution representations. Combined with sub-picture tracks and sub-picture representation techniques, flexible encoding and decoding of video segments is achieved.

Benefits of technology

It improves the flexibility and efficiency of video encoding and decoding, reduces latency due to changes in network conditions and speaker switching, achieves seamless resolution switching and lower latency in streaming, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113875250B_ABST
    Figure CN113875250B_ABST
Patent Text Reader

Abstract

Devices, systems, and methods for digital video coding are described, including reference picture resampling. An example method of video processing includes performing a conversion between a video comprising one or more video segments comprising one or more video units and a bitstream representation of the video, wherein the bitstream representation conforms to a format rule and comprises information related to an adaptive resolution conversion (ARC) process, wherein the format rule specifies an applicability of the ARC process to a video segment, wherein an indication of coding the one or more video units of the video segment at a different resolution is included in the bitstream representation in a syntax structure that is different from a header syntax structure, a decoder parameter set, a video parameter set, a picture parameter set (PPS), a sequence parameter set, and an adaptation parameter set.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 086513, filed on May 12, 2019, in accordance with applicable patent law and / or the rules applicable to the Paris Convention. The entire disclosure of the above application is incorporated by reference into this disclosure for legal purposes. Technical Field

[0003] This patent document relates to video encoding and decoding technology, equipment and systems. Background Art

[0004] Despite advances in video compression, digital video still accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth required for digital video usage is expected to continue to grow. Summary of the Invention

[0005] Devices, systems, and methods related to digital video codecs are described, and in particular, reference picture resampling for video codecs is described. The described methods can be applied to existing video codec standards (e.g., High Efficiency Video Codec (HEVC)) and future video codec standards or video codecs.

[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing conversion between a video comprising one or more video segments and a bitstream representation of the video, the video segment comprising one or more video units, wherein the bitstream representation conforms to format rules and includes information related to adaptive resolution conversion (ARC) processing, wherein the format rules specify the applicability of the ARC processing to the video segments, and wherein an indication that the one or more video units of the video segments are to be encoded or decoded at different resolutions is included in the bitstream representation in a syntax structure different from a header syntax structure, a decoder parameter set (DPS), a video parameter set (VPS), a picture parameter set (PPS), a sequence parameter set (SPS), and an adaptation parameter set (APS).

[0007] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a video comprising one or more video segments comprising one or more video units and a bitstream representation of the video, wherein the bitstream representation conforms to a format rule and includes information related to an adaptive resolution conversion (ARC) process, wherein a dimension of the one or more video units coded with a K-th order exponential Golomb code is signaled in the bitstream representation, and K is a positive integer, and wherein the format rule specifies an applicability of the ARC process to a video segment and includes, in the bitstream representation, an indication of coding the one or more video units of the video segment at different resolutions in a syntax structure.

[0008] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a video comprising one or more video segments comprising one or more video units and a bitstream representation of the video, wherein the bitstream representation conforms to a format rule and includes information related to an adaptive resolution conversion (ARC) process, wherein a dimension of the one or more video units coded with a K-th order exponential Golomb code is signaled in the bitstream representation, and K is a positive integer, and wherein the format rule specifies an applicability of the ARC process to a video segment and includes, in the bitstream representation, an indication of coding the one or more video units of the video segment at different resolutions in a syntax structure.

[0009] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes determining (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is the same as a resolution of a current picture comprising the current video block, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is different from the resolution of the current picture, and as a result of the determination, performing a conversion between the current video block and a bitstream representation of the video by disabling motion information of the second temporal neighboring block in a prediction of the first temporal neighboring block.

[0010] In another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes determining (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is the same as a resolution of a current picture comprising the current video block, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is different from the resolution of the current picture, and as a result of the determination, performing a conversion between the current video block and a bitstream representation of the video by disabling motion information of the second temporal neighboring block in a prediction of the first temporal neighboring block.

[0011] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes determining, for a current video block of a video, that a resolution of a reference picture including a video block associated with the current video block is different from a resolution of a current picture including the current video block, and performing, due to the determination, a conversion between the current video block and a bitstream representation of the video by disabling a prediction process based on the video block in the reference picture.

[0012] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes making a decision, based on at least one dimension of a picture, as to whether the picture is allowed to be used as a collocated reference picture of a current video block of the picture, and performing, based on the decision, a conversion between the current video block of the video and a bitstream representation of the video.

[0013] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes identifying, based on a determination that a dimension of a collocated reference picture including the collocated block is the same as a dimension of a current picture including the current video block, a collocated block for a prediction of a current video block of a video, and performing, using the collocated block, a conversion between the current video block and a bitstream representation of the video.

[0014] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes determining, for a current video block of a video, that a reference picture associated with the current video block has a resolution that is different from a resolution of a current picture including the current video block, and performing, as part of a conversion between the current video block and a bitstream representation of the video, an upsampling operation or a downsampling operation on one or more reference samples of the reference picture, motion information of the current video block, or coding information of the current video block.

[0015] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes determining, for a conversion between a current video block of a video and a bitstream representation of the video, that a height or a width of a current picture including the current video block is different from a height or a width of a collocated reference picture associated with the current video block, and performing, based on the determination, an upsampling operation or a downsampling operation on a buffer storing one or more motion vectors of the collocated reference picture.

[0016] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes deriving, based on dimensions of a current picture including a current video block of a video and dimensions of a collocated picture associated with the current video block, information of an optional temporal motion vector prediction (ATMVP) process applied to the current video block, and performing a conversion between the current video block and a bitstream representation of the video using a temporal motion vector.

[0017] In yet another representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes configuring a bitstream representation of a video for applying an adaptive resolution conversion (ARC) process to a current video block of the video, wherein information related to the ARC process is signaled in the bitstream representation, wherein a current picture including the current video block has a first resolution, and wherein the ARC process includes resampling a portion of the current video block at a second resolution different from the first resolution, and based on the configuration, performing a conversion between the current video block and the bitstream representation of the current video block.

[0018] In yet another representative aspect, the above method is embodied in a form of processor-executable code and stored in a computer-readable program medium.

[0019] In yet another representative aspect, an apparatus configured or operable to perform the above method is disclosed. The apparatus can include a processor programmed to implement the method.

[0020] In yet another representative aspect, a video decoder apparatus can implement a method as described herein.

[0021] The above and other aspects and features of the disclosed technology are described in more detail in the accompanying drawings, descriptions, and claims. BRIEF DESCRIPTION OF DRAWINGS

[0022] FIG. 1 An example of an adaptive stream of two representations of the same content coded at different resolutions is shown.

[0023] FIG. 2 Another example of an adaptive stream of two representations of the same content coded at different resolutions is shown, where segments use a closed group of pictures (GOP) or open GOP prediction structure.

[0024] FIG. 3 An example of an open GOP prediction structure of two representations is shown.

[0025] FIG. 4 An example of a representation switched at an open GOP location is shown.

[0026] FIG. 5 An example of a Random Access Skipped Leading (RASL) picture decoded using a resampled reference picture from another bitstream as a reference is shown.

[0027] FIGS. 6A-6C An example of a partition-wise mixed resolution (RWMR) viewpoint-dependent 360 streaming based on motion-constrained tile sets (MCTS) is shown.

[0028] FIG. 7 An example of collocated subpicture representation with different intra random access point (IRAP) intervals and different sizes is shown.

[0029] FIG. 8 An example of a segment received when a viewing orientation change causes a resolution change at the beginning of the segment is shown.

[0030] FIG. 9 An example of a viewing orientation change is shown.

[0031] FIG. 10 An example of subpicture representation for two subpicture positions is shown.

[0032] FIG. 11 An example of encoder modification for adaptive resolution conversion (ARC) is shown.

[0033] FIG. 12 An example of decoder modification for ARC is shown.

[0034] FIG. 13 An example of tile group based resampling for ARC is shown.

[0035] FIG. 14 An example of ARC processing is shown.

[0036] FIG. 15 An example of optional temporal motion vector prediction (ATMVP) for a coding unit is shown.

[0037] FIG. 16A and FIG. 16B An example of a simplified affine motion model is shown.

[0038] FIG. 17 An example of an affine motion vector field (MVF) per subblock is shown.

[0039] FIG. 18A and FIG. 18B Examples of 4-parameter affine model and 6-parameter affine model are shown, respectively.

[0040] FIG. 19 An example of motion vector prediction (MVP) for AFJNTER for inherited affine candidate is shown.

[0041] FIG. 20 An example of MVP for AFJNTER for constructed affine candidate is shown.

[0042] FIG. 21A and 21B An example of candidate for AFJMERGE is shown.

[0043] FIG. 22 An example of candidate position for affine merge mode is shown.

[0044] FIG. 23 An example of deriving TMVP / ATMVP with ARC is shown.

[0045] FIGS. 24A-24J A flowchart of an example method for video processing is shown.

[0046] FIG. 25 is a block diagram of an example of a hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.

[0047] FIG. 26 is a block diagram of an example video processing system in which the disclosed techniques can be implemented. DETAILED DESCRIPTION

[0048] Embodiments of the disclosed techniques can be applied to existing video coding standards (e.g., HEVC, H.265) and future standards to improve compression performance. Section headings are used in this document to improve readability of the description, and do not limit the discussion or embodiments (and / or implementations) to the respective section in any way.

[0049] 1. Video coding introduction

[0050] Video coding methods and techniques are ubiquitous in modern technology due to the increasing demand for higher resolution video. Video codecs, which are typically electronic circuits or software that compress or decompress digital video, are constantly improving to provide higher coding efficiency. Video codecs convert uncompressed video into a compressed format and vice versa. There is a complex relationship between video quality, the amount of data used to represent the video (determined by the bit rate), the complexity of the encoding and decoding algorithms, sensitivity to data loss and errors, ease of editing, random access, and end-to-end delay (latency). The compressed format typically conforms to a standard video compression specification, such as the High Efficiency Video Coding (HEVC) standard (also known as H.265 or MPEG-H Part 2), the pending Versatile Video Coding standard, or other current and / or future video coding standards.

[0051] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Starting from H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction and transform coding are utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, the JVET adopted and introduced many new methods into the reference software named “Joint Exploration Model” (JEM). In April 2018, the Joint

[0052] AVC and HEVC do not have the ability to change resolution without having to introduce IDR or Intra Random Access Point (IRAP) pictures; this ability can be referred to as adaptive resolution change (ARC). ARC functionality can be beneficial in certain use cases or application scenarios, which include:

[0053] - Rate adaptation in video telephony and conferencing: In order to adapt the coded video to the changing network conditions, when the network conditions become worse and the available bandwidth becomes lower, the encoder can adapt it by coding smaller resolution pictures. Currently, the picture resolution can only be changed after an IRAP picture; this has several problems. The IRAP picture with reasonable quality will be much larger than the inter-coded pictures and will also be much more complex to decode: this will waste time and resources. If the decoder requests a change of resolution for loading reasons, a problem arises. It can also break the low delay buffer conditions, forcing audio resynchronization and the end-to-end delay of the stream will increase, at least temporarily. This will give a bad experience to the user.

[0054] - Change of active speaker in multi-party video conferencing: For multi-party video conferencing, the active speaker is usually displayed with a larger video size than the videos of the other conference participants. When the active speaker changes, it can also be necessary to adjust the picture resolution of each participant. The need for an ARC feature becomes particularly important when such changes occur frequently in a talker- dominant conference.

[0055] - Fast start in streaming: For streaming applications, it is common for the application to buffer up to a certain length of decoded pictures before starting the display. Starting a bitstream at a smaller resolution will allow the application to have enough pictures in the buffer to start the display faster.

[0056] Adaptive stream switching in streaming: The Dynamic Adaptive Streaming over HTTP (DASH) specification includes a feature called @mediaStreamStructureId. This enables switching between different representations at open GOP random access points with non-decodable leading pictures (e.g. CRA pictures with associated RASL pictures in HEVC). When two different representations of the same video have different bitrates but have the same spatial resolution, switching between the two representations can be performed on a CRA picture with associated RASL pictures and with the same @mediaStreamStructureId value, and with the RASL pictures associated with the CRA picture at the switching point can be decoded with acceptable quality, enabling a seamless switch. Using ARC, the @mediaStreamStructureId feature can also be used to switch between DASH representations with different spatial resolutions.

[0057] ARC is also known as dynamic resolution conversion.

[0058] ARC can also be considered as a special case of reference picture resampling (RPR), such as H.263 Annex P.

[0059] 1.1 Reference picture resampling in H.263 Annex P

[0060] This mode describes an algorithm to warp the reference picture before it is used for prediction. It can be useful to have a reference picture with a different source format than the picture being predicted. By warping the shape, size and position of the reference picture, it can also be used for global motion estimation or rotational motion estimation. The syntax includes the warping parameters to be used and the resampling algorithm. The simplest level of operation for the reference picture resampling mode is 4x resampling with implicit factors, since only FIR filters are needed for the up- and down-sampling process. In this case, no additional signaling overhead is needed when the size of the new picture (indicated in the picture header) is different from the size of the previous picture, since its usage is understood.

[0061] 1.2 ARC contribution to VVC

[0062] 1.2.1 JVET-M0135

[0063] Just to initiate the discussion, a preliminary design of ARC as described below (some parts taken from JCTVC-F158) is proposed as a place holder.

[0064] 1.2.1.1 Description of the basic tools

[0065] The basic tools to support ARC are constrained as follows:

[0066] The spatial resolution in both dimensions can differ from the nominal resolution by a factor of 0.5. The spatial resolution can be increased or decreased, resulting in scaling ratios of 0.5 and 2.0.

[0067] The aspect ratio of the video format and the chroma format are not changed.

[0068] The cropped area is scaled by the ratio of the spatial resolutions.

[0069] The reference pictures are simply rescaled as needed and inter prediction is applied as usual.

[0070] 1.2.1.2 Scaling operations

[0071] It is proposed to use simple, zero-phase separable filters for down- and up-scaling. Note that these filters are only used for prediction; the decoder can use more complex scaling for output purposes.

[0072] The following 1:2 downsampling filter with zero phase and 5 taps is used:

[0073] (-1, 9, 16, 9, -1) / 32

[0074] Downsampled samples are located at even sample positions and at the same positions. Luma and chroma use the same filter.

[0075] For 2:1 upsampling, the other samples are generated at odd grid positions using the half-pel motion-compensated interpolation filter coefficients in the latest VVC WD.

[0076] The combined upsampling and downsampling will not change the positions of the phase or chroma sample points.

[0077] 1.2.1.3 Resolution description in parameter sets

[0078] The signaling of the picture resolution in the SPS is changed as follows, which is marked with double brackets for deletion in the following description and the rest of the description in this document (e.g., [[a]] means the deletion of the character "a").

[0079] Sequence parameter set RBSP syntax and semantics

[0080]

[0081] [[pic_width_in_luma_samples specifies the width of each decoded picture in units of luma samples. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0082] pic_height_in_luma_samples specifies the height of each decoded picture in units of luma samples. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.]]

[0083] num_pic_size_in_luma_samples_minus1 plus 1 specifies the number of picture sizes (width and height) in units of luma samples that can be present in the coded video sequence.

[0084] pic_width_in_luma_samples[ i ] specifies the i-th width of a decoded picture in units of luma samples that can be present in the coded video sequence.

[0085] pic_width_in_luma_samples[ i ] shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0086] pic_height_in_luma_samples[ i ] specifies the i-th height of the decoded pictures, in units of luma samples, that can be present in the coded video sequence.

[0087] pic_height_in_luma_samples[ i ] shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0088] Picture parameter set RBSP syntax and semantics

[0089]

[0090] pic_size_idx specifies the index of the i-th picture size in the sequence parameter set. The width of the picture referring to the picture parameter set is pic_width_in_luma_samples[ pic_size_idx ] in luma samples. Likewise, the height of the picture referring to the picture parameter set is pic_height_in_luma_samples[ pic_size_idx ] in luma samples.

[0091] 1.2.2 JVET-M0259

[0092] 1.2.2.1 Background: Sub-picture

[0093] In the Omnidirectional Media Format (OMAF), the term sub-picture track is defined as follows: a track that has a spatiotemporal relationship with other tracks and represents a spatiotemporal subset of the original video content (before video encoding at the content producing side, the original video content is divided into spatiotemporal subsets). A sub-picture track of HEVC can be constructed by rewriting the parameter sets and slice segment headers of a motion-constrained tile set so that it becomes an independent HEVC bitstream. A sub-picture representation can be defined as a DASH representation that carries a sub-picture track.

[0094] JVET-M0261 uses the term sub-picture as the spatial partitioning unit for VVC, summarized as follows:

[0095] 1. Pictures are divided into sub-pictures, tile groups, and tiles.

[0096] 2. A sub-picture is a rectangular set of tile groups starting with a tile group whose tile_group_address is equal to 0.

[0097] 3. Each sub-picture can refer to its own PPS and can thus have its own tile partitioning.

[0098] 4. In the decoding process, a subpicture is treated as a picture.

[0099] 5. The reference picture for decoding a subpicture is generated by extracting the region collocated with the current subpicture from the reference pictures in the decoded picture buffer. The extracted region shall be the decoded subpicture, i.e. inter prediction occurs between subpictures of the same size and at the same position in the picture.

[0100] 6. A slice group is a sequence of slices in the tile raster scan of a subpicture.

[0101] In this contribution, we refer to the term subpicture as defined in JVET-M0261. However, the track encapsulating a sequence of subpictures as defined in JVET-M0261 has very similar properties as the subpicture track defined in OMAF, the examples given below apply in both cases.

[0102] 1.2.2.2 Use cases

[0103] 1.2.2.2.1 Adaptive resolution change in streaming

[0104] Requirements for support of adaptive streaming

[0105] Section 5.13 (“Support of adaptive streaming”) of MPEG N17074 includes the following requirements for VVC:

[0106] In case an adaptive streaming service provides multiple representations of the same content, each representation having different properties (e.g. spatial resolution or sample bit depth), the standard shall support fast representation switching. The standard shall allow the use of efficient prediction structures (e.g. so-called open GOPs) without compromising the ability for fast, and seamless representation switching between representations of different properties such as different spatial resolutions.

[0107] Example of open GOP prediction structure with representation switching

[0108] Content generation for adaptive bitrate streaming includes the generation of different representations, which can have different spatial resolutions. A client requests segments from the representations, and can thus decide in which resolution and bitrate to receive the content. At the client, segments of different representations are concatenated, decoded and played. The client should be able to achieve seamless playback with one decoder instance. As FIG. 1 shown, a closed GOP structure (starting with an IDR picture) is typically used.

[0109] Open GOP prediction structure (starting with a CRA picture) has better compression performance than the respective closed GOP prediction structure. For example, in the case of an IRAP picture interval of 24 pictures, the average bit rate is reduced by 5.6% in terms of Bjontegaard delta bit rate.

[0110] It is reported that open GOP prediction structure also reduces subjective visible quality pumping.

[0111] A challenge in using open GOP in streaming is that RASL pictures cannot be decoded with the correct reference pictures after switching representation. We will describe this challenge in the following with respect to the representation in FIG. 2 .

[0112] A segment starting with a CRA picture includes RASL pictures for which at least one reference picture is located in a previous segment. This is illustrated in FIG. 3 , where picture 0 in both bitstreams is located in a previous segment and used as a reference for predicting a RASL picture.

[0113] The representation switch marked with a dashed rectangle in FIG. 2 is illustrated in FIG. 4 . It can be seen that the reference picture for the RASL picture ("picture 0") is not decoded. Therefore, the RASL picture cannot be decoded and there will be a gap in the video playback.

[0114] However, as described by embodiments of the present invention, it has been found that it is subjectively acceptable to decode a RASL picture with a resampled reference picture. The resampling of "picture 0" and its use as a reference picture for decoding a RASL picture is illustrated in FIG. 5 .

[0115] 1.2.2.2.2 Region-wise mixed resolution (RWMR) 360° video streaming with viewport change

[0116] Background: HEVC-based RWMR streaming

[0117] RWMR 360° streaming provides an increased effective spatial resolution over the viewport. The tile covering the viewport originates from a 6K (6144x3072) ERP picture or equivalent CMP resolution scheme as shown in Figure 6, where the "4K" decoding capability (HEVC level 5.1) is included in OMAF sections D.6.3 and D.6.4, and also adopted in the VR Industry Forum guidelines with the "4K" decoding capability (HEVC level 5.1). Such resolution is claimed to be suitable for head-mounted displays using quad-HD (2560x1440) display panels.

[0118] Encoding : The content is encoded at two spatial resolutions, respectively in cubic dimensions 1536x1536 and 768x768. In both bitstreams, a 6x4 tile grid is used and a motion-constrained tile set (MCTS) is encoded for each tile position.

[0119] Packaging : Each MCTS sequence is encapsulated as a subpicture track and can be used as a subpicture representation in DASH.

[0120] Selection of MCTSs for a stream : From the high-resolution bitstream, 12 MCTS are selected and the complementary 12 MCTS are extracted from the low-resolution bitstream. Thus, a hemisphere (180°x180°) of the stream content originates from the high-resolution bitstream.

[0121] Merging MCTSs into a bitstream to be decoded : The received MCTS of a single time instance are merged into a coded picture of 1920x4608, which complies with HEVC level 5.1. Another option for the merged picture is to have four tile columns of width 768, two tile columns of width 384 and three tile rows of height 768 luma samples, resulting in a picture of 3840x2304 luma samples.

[0122] Background: Several representations of different IRAP intervals for viewport-dependent 360° streaming

[0123] When the viewing orientation in a HEVC-based viewport-dependent 360° streaming changes, a new selection of subpicture representations can take effect at the next IRAP-aligned segment boundary. The subpicture representations are merged into a coded picture for decoding, so the VCL NAL unit types are aligned in all selected subpicture representations.

[0124] To provide a trade-off between the response time to changes in viewing orientation and rate-distortion performance when the viewing orientation is stable, multiple versions of the content can be encoded at different IRAP intervals. In FIG. 7 A set of collocated subpicture representations for encoding is shown in Table 1 for the scheme presented in Figure 6.

[0125] FIG. 8 An example is shown where a viewing orientation change results in a higher resolution (768x768) sub-picture location being selected. In this example, a viewing orientation change occurs such that segment 4 is received from the short IRAP interval sub-picture representation. Thereafter, the viewing orientation is stable and thus the long IRAP interval version can be used starting from segment 5.

[0126] Problem statement

[0127] Since the viewing orientation is gradually moving in a typical viewing situation, the resolution changes in RWMR viewport dependent streaming only in a subset of the sub-picture locations. FIG. 9 A viewing orientation change from the perspective shown slightly above and to the right of Figure 6 is shown. The perspective partitions with different resolutions than previously are denoted with a "C". It can be observed that out of the 24 perspective partitions, the resolution of 6 of them changed. However, as mentioned above, in response to the viewing orientation change, segments starting with an IRAP picture need to be received for all 24 perspective partitions. In terms of rate-distortion performance of the streaming, it is inefficient to update all sub-picture locations with segments starting with an IRAP picture.

[0128] In addition, the ability to use open GOP prediction structure with sub-picture representation for RWMR 360° streaming is desirable to improve rate-distortion performance and avoid visible picture quality pumping caused by closed GOP prediction structure.

[0129] Proposed design goals

[0130] The following design goals are proposed:

[0131] 1. The VVC design should allow merging of a sub-picture originating from a random access picture and another sub-picture originating from a non-random access picture into the same coded picture conforming to VVC.

[0132] 2. The VVC design should allow using open GOP prediction structure in sub-picture representation without compromising fast and seamless representation switching capability between sub-picture representations of different properties such as different spatial resolutions, while allowing merging of sub-picture representations into a single VVC bitstream.

[0133] The following design goals can be achieved with FIG. 10To illustrate the design goal, a subpicture representation is given with two subpicture positions. For both subpicture positions, a separate version of the content is encoded for each combination between two resolutions and two random access intervals. Some segments start with an open GOP prediction structure. The change of viewing orientation causes the resolution of subpicture position 1 to switch at the beginning of segment 4. Since segment 4 starts with a CRA picture associated with a RASL picture, the reference pictures of those RASL pictures in segment 3 need to be resampled. It should be noted that this resampling applies to subpicture position 1, while the decoded subpictures of the other subpicture positions are not resampled. In this example, the change of viewing orientation does not cause a change of resolution for subpicture position 2, so the decoded subpictures of subpicture position 2 are not resampled. In the first pictures of segment 4, the segment for subpicture position 1 includes subpictures originating from a CRA picture, while the segment for subpicture position 2 includes subpictures originating from a non-random access picture. It is proposed to allow in VVC to merge these subpictures into a coded picture.

[0134] 1.2.2.2.3 Adaptive resolution change in video conferencing

[0135] JCTVC-F158 proposes an adaptive resolution change mainly for video conferencing. The following subsection is copied from JCTVC-F158 and introduces the use case for which the adaptive resolution is claimed to be efficient.

[0136] Seamless network adaptation and error resilience

[0137] Applications such as video conferencing and streaming over packet networks often require the encoded stream to adapt to changing network conditions, especially when the bit rate is too high and data is lost. Such applications usually have a return channel that allows the encoder to detect errors and perform adjustments. The encoder has two main tools at its disposal: reducing the bit rate and changing the temporal or spatial resolution. Changing the temporal resolution can be efficiently achieved by encoding with a hierarchical prediction structure. However, to achieve the best quality, changing the spatial resolution is required, and is part of a well-designed encoder for video communication.

[0138] Changing the spatial resolution within AVC requires sending an IDR frame and resetting the stream. This causes serious problems. An IDR frame with reasonable quality will be much larger than an inter picture and will also be correspondingly more complex to decode: this wastes time and resources. If the decoder requests a change of resolution for loading reasons, a problem arises. It can also break the low-delay buffer conditions, forcing audio resynchronization, and will increase (at least temporarily) the end-to-end delay of the stream. This gives a poor experience to the user.

[0139] To minimize these problems, IDRs are typically sent at a similar number of bits as P-frames, at low quality, and it takes a significant amount of time to recover to full quality for a given resolution. To get sufficiently low latency, the quality can indeed be very low, and there is often visible blurring before the image "refocuses". In fact, in terms of compression, intra-frames do very little work: it is just a way to restart the stream.

[0140] Thus, there is a need for a method in HEVC that allows changing resolution, especially under challenging network conditions, and with minimal impact on the subjective experience.

[0141] Fast start

[0142] It would be useful to have a "fast start" mode, where the first frame is sent at a reduced resolution, and the resolution is increased over the next few frames, in order to reduce latency and reach normal quality more quickly without unacceptable image blurring at the start.

[0143] Conference "composition"

[0144] Video conferences often also have the feature of showing the speaker in full screen, and the other participants in smaller resolution windows. To efficiently support this feature, the smaller pictures are typically sent at a lower resolution. Then, when a participant becomes the speaker and is shown full screen, this resolution is increased. Sending an intra-frame at this point causes an unpleasant hiccup in the video stream. This effect can be very noticeable and annoying if the speakers alternate quickly.

[0145] 1.2.2.3 Proposed design goals

[0146] The following are the high-level design choices proposed for VVC version 1:

[0147] 1. Propose to include the reference picture resampling process in VVC version 1 for the following use cases:

[0148] - Use of efficient prediction structures in adaptive streaming (e.g. so-called open GOPs) without compromising the ability for fast and seamless representation switching between representations of different properties such as different spatial resolutions.

[0149] - Adaptation of low-latency conversational video content to network conditions and application-induced resolution changes without noticeable latency or latency changes.

[0150] 2. It is proposed that the VVC design allows merging a sub-picture originating from a random access picture and another sub-picture originating from a non-random access picture into the same coded picture conforming to VVC. It is claimed that it enables efficient handling of viewing orientation changes in mixed quality and mixed resolution viewport adaptive 360° streaming.

[0151] 3. It is proposed to include resampling handling in terms of sub-pictures in VVC version 1. It is claimed that it enables efficient prediction structures to be enabled for more efficient handling of view orientation changes in mixed resolution viewport adaptive 360° streaming.

[0152] 1.2.3 JVET-N0048

[0153] The use cases and design goals for adaptive resolution change (ARC) are discussed in detail in JVET-M0259. The summary is as follows:

[0154] 1. Real-time communication

[0155] JCTVC-F158 originally included the following use cases for adaptive resolution change:

[0156] a. Seamless network adaptation and error recovery (through dynamic adaptive resolution change);

[0157] b. Fast startup (gradually increasing resolution at session startup or reset);

[0158] c. Conference “formation” (providing higher resolution for the speaker);

[0159] 2. Adaptive streaming

[0160] Section 5.13 (“Support of adaptive streaming”) of MPEG N17074 includes the following requirements for VVC:

[0161] In case of an adaptive streaming service providing multiple representations with the same content, each with different properties (e.g., spatial resolution or sample bit depth), the standard shall support fast representation switching. The standard shall allow the use of efficient prediction structures (e.g., so-called open GOPs) without compromising the fast and seamless representation switching capability between different properties such as different spatial resolutions.

[0162] JVET-M0259 discusses how this requirement can be met by resampling the reference pictures of the leading pictures.

[0163] 3. 360° viewport related streaming

[0164] JVET-M0259 discusses how this use case can be addressed by resampling certain independently coded picture regions of the reference pictures of the leading pictures.

[0165] This contribution proposes an adaptive resolution coding method, which claims to satisfy all the use cases and design goals mentioned above. This proposal, together with JVET-N0045 (which proposes independent subpicture layers), handles the streaming and conferencing "composition" use cases related to 360° viewport.

[0166] Proposed specification text

[0167] Signaling

[0168] sps_max_rpr

[0169]

[0170] sps_max_rpr specifies the maximum number of active reference pictures in reference picture list 0 or 1 for any slice group in the CVS where pic_width_in_luma_samples and pic_height_in_luma_samples are not equal to pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively.

[0171] Picture width and height

[0172]

[0173]

[0174]

[0175] max_width_in_luma_samples specifies that pic_width_in_luma_samples in any active PPS must be less than or equal to max_width_in_luma_samples for any picture of the CVS for which this SPS is active, which is a requirement for bitstream conformance.

[0176] max_height_in_luma_samples specifies that pic_height_in_luma_samples in any active PPS must be less than or equal to max_height_in_luma_samples for any picture of the CVS for which this SPS is active, which is a requirement for bitstream conformance.

[0177] Advanced decoding process

[0178] The decoding process operation for the current picture CurrPic is as follows:

[0179] 1. Decoding of NAL units is specified in clause 8.2.

[0180] 2. The processing in clause 8.3 specifies the following decoding processes using the syntax elements in and above the slice header:

[0181] - Derive variables and functions related to picture order count as specified in clause 8.3.1. This method needs to be invoked only for the first slice group of a picture.

[0182] - At the beginning of the decoding process for each slice group of a non-IDR picture, invoke the decoding process for reference picture list construction specified in clause 8.3.2 to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).

[0183] - Invoke the decoding process for reference picture marking in clause 8.3.3, where a reference picture can be marked as "unused for reference" or "used for long-term reference". This method needs to be invoked only for the first slice group of a picture.

[0184] - For each active reference picture (with pic_width_in_luma_samples or pic_height_in_luma_samples not equal to pic_width_in_luma_samples or pic_height_in_luma_samples of CurrPic, respectively) in RefPicList[0] and RefPicList[1], the following applies:

[0185] - Invoke the resampling process in clause X.Y.Z [Ed.(MH): details of the invocation parameters to be added] whose output has the same reference picture marking and picture order count as the input.

[0186] - The reference picture used as input to the resampling process is marked as "unused for reference".

[0187] 3. [Ed.(YK): Add here the invocation of the decoding processes for coding tree units, scaling, transform, loop filtering, etc.]

[0188] 4. After all slice groups of the current picture have been decoded, mark the current decoded picture as "used for short-term reference".

[0189] Resampling process

[0190] The SHVC resampling process (HEVC clause H.8.1.4.2) is proposed with the following additions:

[0191]

[0192] If sps_ref_wraparound_enabled_flag is equal to 0, the sample values tempArray[ n ] for n = 0..7 are derived as follows:

[0193]

[0194] Otherwise, the sample values tempArray[ n ] for n = 0..7 are derived as follows:

[0195]

[0196]

[0197] If sps_ref_wraparound_enabled_flag is equal to 0, the sample values tempArray[ n ] for n = 0..3 are derived as follows:

[0198]

[0199] Adaptive resolution change has appeared as a concept in video compression standards at least since 1996; in particular, H.263+ related proposals on reference picture resampling (RPR, Annex P) and reduced resolution update (Annex Q). It has recently gained some attention, first in the context of proposals by Cisco during JCT-VC, then in the context of VP9 (now moderately widely deployed), and most recently in the context of VVC. ARC allows to reduce the number of samples that need to be encoded for a given picture, and to upsample the final reference picture to a higher resolution when needed.

[0200] We believe that ARC is of particular interest in two cases:

[0201] 1) Intra coded pictures (such as IDR pictures) are typically much larger than inter coded pictures. It is clear that, for whatever reason, downsampling a picture to be intra coded provides a better input for future prediction. It is also clearly advantageous from the point of view of rate control, at least in low delay applications.

[0202] 2) When operating a codec at places close to a discontinuity, as is typically done by at least some cable and satellite operators, it can be convenient to have ARC even for non-intra coded pictures, such as in scene transitions without hard transition points.

[0203] 3) Perhaps a bit too far ahead: can the concept of fixed resolution be generally justified? With the advent of CRTs and the prevalence of scaling engines in rendering devices, the hard binding between rendering and encoding resolution is in the past. In addition, we note that there is some available research that suggests that when a lot of activity occurs in a video sequence, most people cannot focus on fine details (possibly associated with high resolution) even if that activity is at other locations in the spatial domain. If this is true and generally accepted, then fine-grained resolution changes can be a better rate control mechanism than adaptive QP. Due to lack of data, we are currently putting this forward for discussion (feedback from the knowledgeable is welcome). Of course, eliminating the concept of fixed resolution bitstreams has myriad system-level and implementation implications that we are well aware of (at least at the level at which they exist if not their detailed nature).

[0204] Technically, ARC can be implemented as reference picture resampling. There are two main aspects to implementing reference picture resampling: the resampling filter and the signaling of resampling information in the bitstream. This document focuses on the latter and only touches on the former to the extent that we have implementation experience. More research on suitable filter design is encouraged.

[0205] Overview of existing ARC implementations

[0206] FIG. 11 and 12 show the existing ARC encoder / decoder implementations, respectively. In our implementation, the width and height of a picture can be changed according to the granularity of each picture, while ignoring the picture type. At the encoder, the input image data is down-sampled to the selected picture size for the current picture encoding. After the first input picture is encoded as an intra picture, the decoded picture is stored in the decoded picture buffer (DPB). When the subsequent pictures are down-sampled at different sampling rates and encoded as inter pictures, the reference pictures in the DPB are up-scaled / down-scaled according to the spatial ratio between the reference picture size and the current picture size. At the decoder, the decoded pictures are stored in the DPB without resampling. However, when used for motion compensation, the reference pictures in the DPB are up-scaled / down-scaled according to the spatial ratio between the reference and the current decoded picture. The decoded picture is up-sampled to the original picture size or the required output picture size when the decoded picture is highlighted for display. In the motion estimation / compensation process, the motion vectors are scaled according to the picture size ratio as well as the picture order count difference.

[0207] Signaling of ARC parameters

[0208] The term ARC parameters is used herein as the combination of any parameters needed to make ARC work. In the simplest case, it can be a zoom factor, or an index into a table with defined zoom factors. It can be a target resolution (e.g. in samples or maximum CU size granularity), or an index into a table providing target resolutions, as proposed in JVET-M0135. It also includes the filter selector of the upsampling / downsampling filter in use, and even filter parameters (up to filter coefficients).

[0209] From the start, we propose in this paper to at least conceptually allow different ARC parameters for different parts of a picture. We propose that the appropriate syntax structure according to the current VVC draft should be a rectangular tile group (TG). Those using scan order TGs will be limited to using ARC for the complete picture only, or to including scan order TGs in rectangular TGs to some extent (we do not recall that TG nesting has been discussed so far, and perhaps is a bad idea). The easy specification by bitstream constraint is possible.

[0210] As different TGs can have different ARC parameters, the appropriate location for the ARC parameters will be in the TG header or in a parameter set range-wise, and referenced by the TG header-adaptation parameter set in the current VVC draft, or more in detail, a table in a higher parameter set. Of these three choices, we propose to use the TG header at this time to encode the reference to the table entry including the ARC parameters, and that table to be in the SPS, with the maximum table value encoded in the (forthcoming) DPS. We can directly encode the zoom factor into the TG header without the need to use any parameter set value. If, as we do, each tile group signaling of the ARC parameters is a design criterion, the use of the PPS as reference is prohibited, as proposed in JVET-M0135.

[0211] For the table entry itself, we see many options:

[0212] • Encode the down-sampling factor, is it to use both dimensions simultaneously, or the X and Y dimensions separately? This is mainly a hardware (HW) implementation discussion, and some people can prefer a result where the zoom factor in the X dimension is quite flexible, but the scaling factor in the Y dimension is fixed to 1, or a choice is very limited. We propose that the syntax is the wrong place to express such constraints, and if needed, we prefer to express the constraints as conformance requirements. In other words, keep the syntax flexible.

[0213] • Encode the target resolution. This is what we propose in the following. There can be more or less complex constraints on these resolutions relative to the current resolution, perhaps expressed in bitstream conformance requirements.

[0214] • It is preferred to down-sample each slice group for picture composition / extraction. However, it is not critical from the signaling point of view. If the group makes an unwise decision to allow ARC only at picture granularity, we can always require all TGs to use the same ARC parameters for bitstream conformance requirement.

[0215] • Control information related to ARC. In our design below, it includes the size of the reference picture.

[0216] • Do we need flexibility in filter design? What is more important than a bunch of code points? If yes, put them into APS? (If no, please do not bring up the APS update discussion again. If the down-sampling filter is changed and ALF is kept, we propose that the bitstream must incur extra overhead.)

[0217] For now, to keep the proposed techniques consistent and simple (as much as possible), we propose

[0218] • Fixed filter design

[0219] • Target resolution in SPS table with bitstream constraint, TBD.

[0220] • Min / Max target resolution in DPS to facilitate cap exchange / coordination.

[0221] The resulting syntax is shown below:

[0222] Decoder parameter set RBSP syntax

[0223]

[0224] max_pic_width_in_luma_samples specifies the maximum width of decoded pictures in the bitstream in units of luma samples. max_pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. The value of dec_pic_width_in_luma_samples[ i ] shall not be greater than the value of max_pic_width_in_luma_samples.

[0225] max_pic_height_in_luma_samples specifies the maximum height of decoded pictures in units of luma samples. max_pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. The value of dec_pic_height_in_luma_samples[ i ] shall not be greater than the value of max_pic_height_in_luma_samples.

[0226] sequence parameter set RBSP syntax

[0227]

[0228]

[0229] adaptive_pic_resolution_change_flag equal to 1 specifies that the output picture size (output_pic_width_in_luma_samples, output_pic_height_in_luma_samples), the indication of the number of decoded picture sizes (num_dec_pic_size_in_luma_samples_minus1) and at least one decoded picture size (dec_pic_width_in_luma_samples[ i ], dec_pic_height_in_luma_samples[ i ]) are present in the SPS. The reference picture size (reference_pic_width_in_luma_samples, reference_pic_height_in_luma_samples) is conditionally present depending on the value of reference_pic_size_present_flag.

[0230] output_pic_width_in_luma_samples specifies the width of the output picture in units of luma samples. output_pic_width_in_luma_samples shall not be equal to 0.

[0231] output_pic_height_in_luma_samples specifies the height of the output picture in units of luma samples. output_pic_height_in_luma_samples shall not be equal to 0.

[0232] reference_pic_size_present_flag equal to 1 specifies that reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples are present.

[0233] reference_pic_width_in_luma_samples specifies the width of the reference picture in units of luma samples. output_pic_width_in_luma_samples shall not be equal to 0. If not present, the value of reference_pic_width_in_luma_samples is inferred to be equal to dec_pic_width_in_luma_samples[ i ].

[0234] reference_pic_height_in_luma_samples specifies the height of the reference picture in units of luma samples. output_pic_height_in_luma_samples shall not be equal to 0. If not present, the value of reference_pic_height_in_luma_samples is inferred to be equal to dec_pic_height_in_luma_samples[ i ].

[0235] NOTE 1 - The size of the output picture shall be equal to the values of output_pic_width_in_luma_samples and output_pic_height_in_luma_samples. When a reference picture is used for motion compensation, the size of the reference picture shall be equal to the values of reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples.

[0236] num_dec_pic_size_in_luma_samples_minus1 plus 1 specifies the number of decoded picture sizes (dec_pic_width_in_luma_samples[ i ], dec_pic_height_in_luma_samples[ i ]) in units of luma samples in the coded video sequence.

[0237] dec_pic_width_in_luma_samples[ i ] specifies the i-th width of the decoded picture size in luma samples in the coded video sequence. dec_pic_width_in_luma_samples[ i ] shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0238] dec_pic_height_in_luma_samples[ i ] specifies the i-th height of the decoded picture size in luma samples in the coded video sequence. dec_pic_height_in_luma_samples[ i ] shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0239] NOTE 2 - The size of the i-th decoded picture (dec_pic_width_in_luma_samples[ i ], dec_pic_height_in_luma_samples[ i ]) can be equal to the decoded picture size of the decoded picture in the coded video sequence.

[0240] Slice header syntax

[0241]

[0242] dec_pic_size_idx specifies that the width of the decoded picture shall be equal to pic_width_in_luma_samples[ dec_pic_size_idx ] and the height of the decoded picture shall be equal to pic_height_in_luma_samples[ dec_pic_size_idx ].

[0243] Filters

[0244] The proposed design conceptually includes four different filter sets: a down-sampling filter from the original picture to the input picture, an up / down-sampling filter for re-scaling the reference picture for motion estimation / compensation, and an up-sampling filter from the decoded picture to the output picture. The first and last one can be left as non-normative. Within the scope of the norm, either the up / down-sampling filter needs to be explicitly signaled in the appropriate parameter set or the up / down-sampling filter is pre-defined.

[0245] Our implementation uses the down-sampling filter of SHVC (SHM version 12.4), which is a 12-tap and 2D separable filter used for down-sampling to adjust the size of the reference picture to be used for motion compensation. In the current implementation, only dyadic sampling is supported. Therefore, by default, the phase of the down-sampling filter is set to be equal to zero. For up-sampling, an 8-tap interpolation filter with 16 phases is used to shift the phase and align the luma and chroma pixel positions with the original positions.

[0246] Table 1 and Table 2 provide the 8-tap filter coefficients fL[p, x] for luma up-sampling process for p = 0..15 and x = 0..7, and the 4-tap filter coefficients fC[p, x] for chroma up-sampling process for p = 0..15 and x = 0..3.

[0247] Table 3 provides the 12-tap filter coefficients for down-sampling process. The same filter coefficients are used for both luma and chroma down-sampling.

[0248] Table 1. Luma up-sampling filter with 16 phases

[0249]

[0250] Table 2. Chroma up-sampling filter with 16 phases

[0251]

[0252]

[0253] Table 3. Down-sampling filter coefficients for luma and chroma

[0254]

[0255] We have not experimented with other filter designs yet. We expect that subjective and objective gains can be expected (perhaps obvious) when using filters that are suitable for the content and / or scaling factors.

[0256] Tile group boundary discussion

[0257] Since there are indeed many works related to tile groups, our implementation for tile group (TG) based ARC is not fully completed yet. We tend to revisit the implementation after the discussion of spatially compositing and extracting multiple sub-pictures to a composed picture in the compressed domain produces at least a draft of the work. However, this does not prevent us from inferring the results to some extent and adjusting our signaling design accordingly.

[0258] So far, we believe that the slice group header is the right place for the proposed dec_pic_size_idx and the like for the reasons already stated. We use the single ue(v) codepoint dec_pic_size_idx that conditionally appears in the slice group header to indicate the employed ARC parameters. To match our implementation (i.e., ARC only per picture), we now need to do one thing in the specification space, namely, either encode only a single slice group or have all TG headers for a given coded picture have the same dec_pic_size_idx (if present) as a condition for bitstream conformance.

[0259] The parameter dec_pic_size_idx can be moved to any header that starts a subpicture. Our current feeling is that it will most likely continue to be the slice group header.

[0260] In addition to these syntactical considerations, some additional work is needed to enable slice group or subpicture based ARC. Perhaps the most difficult part is how to deal with the unwanted samples in a picture where a subpicture has been resampled to a smaller size.

[0261] Consider FIG. 13 the right part of which consists of four subpictures (possibly represented as four rectangular slice groups in the bitstream syntax). On the left, the lower right TG is subsampled to half size. How do we deal with the samples outside the relevant area marked “Half”?

[0262] A commonality of some existing video coding standards is that they do not support spatial extraction of parts of a picture in the compressed domain. This means that each sample of a picture is represented by one or more syntax elements and each syntax element affects at least one sample. If we want to preserve, we need to somehow fill the area around the samples covered by the TG marked “Half” after downsampling. H.263+ Annex P solves this problem by padding; in fact, the sample values of the padded samples can be signaled in the bitstream (within certain strict limitations).

[0263] An alternative that can constitute a major departure from the previous assumption can relax the current understanding, namely, that each sample of a reconstructed picture must be represented by something in the coded picture (even if that something is just a skipped block), but such an alternative can be needed in any case if we want to support sub-bitstream extraction (and composition) based on rectangular parts of a picture. Reconstruct

[0264] Implementation notes, system implications, and profiles / levels

[0265] We propose to include the basic ARC in the "base / main" profile. If not needed for some application scenarios, it can be removed using sub-profiles. Some restrictions are acceptable. In this regard, we note that some H.263+ profiles and "Recommended Modes" (previous profiles) include a restriction to Annex P that can only be used as "implicit factor 4", i.e. binary down-sampling in both dimensions. This is sufficient to support fast start (fast I-frame acquisition) in video conferencing.

[0266] The design leads us to believe that all filtering can be done "on the fly" and that there is no increase in memory bandwidth, or only a negligible increase. As far as we are concerned, there is no need to push ARC into exotic profiles.

[0267] We do not believe that complex tables, etc. can be effectively used for performance exchange, as proposed in Marrakech with JVET-M0135. Assuming a request-response and similar limited depth signal exchange, the number of choices is too large to allow meaningful cross-supplier interoperability. As far as we are concerned, in reality, to support ARC in a meaningful way in a performance exchange scenario, we have to fall back to a small number of interoperability points. For example: no ARC, ARC with implicit factor 4, full ARC. As an alternative, we can specify the required support for all ARCs and leave the restriction of bitstream complexity to higher level SDOs. In any case, this is a strategic discussion we should have (in addition to the one already existing in the context of sub-profiles and flags).

[0268] Regarding levels: we believe that the basic design principle must be that, as a condition for bitstream conformance, the sample count of up-sampled pictures must fit the level of the bitstream, regardless of how many up-samplings are signaled in the bitstream, and all samples must fit the coded picture of the up-sampling. We note that this is not the case in H263+. There can be some samples missing there.

[0269] 1.2.5 JVET-N0118

[0270] The following aspects are proposed:

[0271] 1) Signaling a list of picture resolutions in the SPS and signaling the index of this list in the PPS to specify the size of individual pictures.

[0272] 2) For any picture that is going to be output, the decoded picture before resampling is cropped (as needed) and output, i.e. the resampled picture is not used for output, only for inter prediction reference.

[0273] 3) Support of 1.5 and 2 times resampling rates. No support of arbitrary resampling rates. Further study whether one or more than two other resampling rates are needed.

[0274] 4) Between picture-level resampling and block-level resampling, the preference is for block-level resampling.

[0275] a. However, if picture-level resampling is chosen, the following aspects are proposed:

[0276] i. When a reference picture is resampled, both the resampled version of the reference picture and the original, resampled version are stored in the DPB, so both affect the fullness of the DPB.

[0277] ii. When the corresponding non-resampled reference picture is marked as "unused for reference", the resampled reference picture is marked as "unused for reference".

[0278] iii. The RPL signaling syntax remains unchanged, while the construction process of RPL is modified as follows: when a reference picture needs to be included in the RPL entry, and the version of that reference picture with the same resolution as the current picture is not in the DPB, the picture resampling process is invoked and the resampled version of that reference picture is included in the RPL entry.

[0279] iv. The number of resampled reference pictures that can exist in the DPB should be limited to, for example, less than or equal to 2.

[0280] b. Otherwise (block-level resampling is chosen), the following is recommended:

[0281] i. To limit the worst-case decoder complexity, it is proposed that bi-prediction of a block from a reference picture with a different resolution than the current picture is not allowed.

[0282] ii. Another option is that when resampling and quarter-pel interpolation are needed, the two filters are combined together and the operation is applied immediately.

[0283] 5) Regardless of which picture-based resampling method and block-based resampling method is chosen, it is proposed that temporal motion vector scaling is applied when needed.

[0284] 1.2.5.1 Implementation

[0285] The ARC software is implemented on top of VTM-4.0.1, but with the following changes:

[0286] - The list of supported resolutions is signaled in the SPS.

[0287] - The spatial resolution signaling is moved from SPS to PPS.

[0288] - A picture-based resampling scheme is implemented to resample reference pictures. After a picture is decoded, the reconstructed picture can be resampled to a different spatial resolution. Both the original reconstructed picture and the resampled reconstructed picture are stored in the DPB for future pictures to be referenced in decoding order.

[0289] - The implemented resampling filters are based on the filters tested in JCTVC-H0234 as follows:

[0290] o Up-sampling filter: 4-tap + / - quarter phase DCTIF with taps (-4, 54, 16, -2) / 64.

[0291] o Down-sampling filter: hll filter with taps (1, 0, -3, 0, 10, 16, 10, 0, -3, 0, 1) / 32.

[0292] - When constructing the reference picture lists (i.e., L0 and L1) of a current picture, only reference pictures with the same resolution as the current picture are used. Note that a reference picture can be available either at its original size or at its resampled size.

[0293] - TMVP and ATMVP can be enabled; however, when the original coded resolution of a current picture and a reference picture are different, TMVP and ATMVP will be disabled for that reference picture.

[0294] - To facilitate and simplify the encoder implementation, the encoder outputs the highest available resolution when outputting a picture.

[0295] 1.2.5.2 Signaling on picture size and picture output

[0296] 1. On the list of spatial resolutions of coded pictures in the bitstream

[0297] Currently, all coded pictures in a CVS have the same resolution. Therefore, only one resolution (i.e., the width and height of the picture) is signaled directly in the SPS. With the support of ARC, the list of picture resolutions needs to be signaled instead of one resolution, and we propose to signal the list in the SPS and signal the index of the list in the PPS to specify the size of an individual picture.

[0298] 2. On picture output

[0299] We propose that for any picture that is going to be output, the decoded picture before resampling (as needed) is cropped and output, i.e., the resampled picture is not used for output, only for inter prediction reference. The ARC resampling filter should be designed to optimize the use of the resampled picture for inter prediction, and such filter can not be optimal for picture output / display purposes, for which the video terminal device usually has implemented optimized output scaling functions.

[0300] 1.2.5.3 On resampling

[0301] The resampling of decoded pictures can be either picture-based or block-based. For the final ARC design in VVC, we are more inclined to block-based resampling compared to picture-based resampling. We suggest that discussion on both approaches be conducted, and the JVET decides which one of the two approaches should be specified for the ARC support in VVC.

[0302] Picture-based resampling

[0303] In picture-based resampling for ARC, a picture is resampled only once for a certain resolution, and then it is stored in the DPB, while the non-resampled version of the same picture is also kept in the DPB.

[0304] There are two issues with adopting picture-based resampling for ARC: 1) extra DPB buffer is needed to store the resampled reference picture, and 2) extra memory bandwidth is needed due to the increased operations of reading reference picture data from the DPB and writing reference picture data to the DPB.

[0305] It is not a good idea to keep only one version of a reference picture in the DPB for picture-based resampling. If we keep only the non-resampled version, the reference picture can need to be resampled multiple times, as multiple pictures can reference the same reference picture. On the other hand, if the reference picture is resampled and we keep only the resampled version, then we need to apply inverse resampling when the reference picture needs to be output, as it is better to output the non-resampled picture as mentioned above. This is an issue because the resampling process is not a lossless operation. Take a picture A, then downsample it, and then upsample it to get A' with the same resolution as A. A and A' will not be the same; A' includes less information than A because some high frequency information is lost during the downsample and upsample process.

[0306] To address the issue of extra DPB buffer and memory bandwidth, we propose that if picture-based resampling is used for the ARC design in VVC, the following applies:

[0307] 1. When a reference picture is resampled, both the resampled version of the reference picture and the original, resampled version are stored in the DPB, so both affect the fullness of the DPB.

[0308] 2. When the corresponding non-resampled reference picture is marked as "unused for reference", the resampled reference picture is marked as "unused for reference".

[0309] 3. The reference picture list (RPL) of each slice group includes reference pictures with the same resolution as the current picture. Although the RPL signaling syntax does not need to be changed, the RPL construction process can be modified to ensure the content in the previous sentence, as follows: when a reference picture needs to be included in an RPL entry, if a version of the reference picture with the same resolution as the current picture is not available, the picture resampling process is invoked and the resampled version of the reference picture is included.

[0310] 4. The number of resampled reference pictures that can exist in the DPB should be limited to, for example, less than or equal to 2.

[0311] In addition, to enable the use of temporal MVs (e.g., Merge mode and ATMVP) in the case where the temporal MV comes from a reference frame with a different resolution than the current frame, we propose to scale the temporal MV to the current resolution as needed.

[0312] Block-based ARC resampling

[0313] In block-based resampling for ARC, reference blocks are resampled as needed and the resampled pictures are not stored in the DPB.

[0314] The main issue here is the additional decoder complexity. This is because a block in a reference picture can be referenced by multiple blocks in another picture and blocks in multiple pictures.

[0315] When a block in a reference picture is referenced by a block in the current picture and the resolutions of the reference picture and the current picture are different, the reference block is resampled by invoking an interpolation filter so that the reference block has integer pixel resolution. When the motion vector is in quarter-pixel, the interpolation process is invoked again to obtain a resampled reference block in quarter-pixel resolution. Therefore, for each motion compensation operation of a current block that comes from a reference block involving different resolutions, up to two, instead of one, interpolation filter operations are needed. Without ARC support, at most only one interpolation filter operation is needed (i.e., to generate a reference block in quarter-pixel resolution).

[0316] To limit the worst-case complexity, we propose that if the ARC design in VVC uses block-based resampling, the following applies:

[0317] - Bi-prediction of a block from a reference picture with a different resolution than the current picture is not allowed.

[0318] - More precisely, the constraint is as follows: for a current block blkA in a current picture picA that references a reference block blkB in a reference picture picB, when picA and picB have different resolutions, block blkA shall be a uni-predicted block.

[0319] With this constraint, the worst case number of interpolation operations required to decode a block is limited to 2. If the block references a block from a picture with a different resolution, then as mentioned above, the number of interpolation operations required is 2. This is the same as the case where the block references a reference block from a picture with the same resolution and is coded as a bi-predicted block, since the number of interpolation operations is also 2 (i.e., one for obtaining the quarter-pel resolution for each reference block).

[0320] To simplify the implementation, we propose another variant, i.e., if the ARC design in VVC uses block-based resampling, then the following applies:

[0321] - If the resolution of the reference frame and the current frame are different, then the corresponding position of each pixel of the predictor is first computed, and then only one interpolation is applied. That is, two interpolation operations (i.e., one for resampling, one for quarter-pel interpolation) are combined into only one interpolation operation. The sub-pixel interpolation filter in the current VVC can be reused, but in this case, the granularity of the interpolation should be enlarged, but the number of interpolation operations is reduced from 2 to 1.

[0322] - To enable temporal MV usage (e.g., Merge mode and ATMVP) in the case where the temporal MV comes from a reference frame with a different resolution than the current frame, we propose to scale the temporal MV to the current resolution as needed.

[0323] Resampling rate

[0324] In JVET-M0135, to start the discussion on ARC, it was proposed that for the starting point of ARC, only a resampling rate of 2x is considered (meaning, 2x2 for upsampling, 1 / 2 x 1 / 2 for downsampling). Through further discussions on this topic after the Marrakech meeting, we learned that supporting only a resampling rate of 2x is very limited, because in some cases, a smaller difference between the resampled and the non-resampled resolution would be more beneficial.

[0325] Although it would be desirable to support arbitrary resampling rates, support seems difficult. This is because in order to support arbitrary resampling rates, the number of resampling filters that must be defined and implemented seems too large and puts a heavy burden on the implementation of decoders.

[0326] We propose that more than one but a small number of resampling rates should be supported, at least 1.5x and 2x resampling rates should be supported, and arbitrary resampling rates should not be supported.

[0327] 1.2.5.4 Maximum DPB buffer size and buffer fullness

[0328] With ARC, a DPB can include decoded pictures with different spatial resolutions within the same CVS. For DPB management and related aspects, it is no longer valid to calculate DPB size and fullness in units of decoded pictures.

[0329] Below is a discussion of some specific aspects that need to be addressed if ARC is supported, and possible solutions in the final VVC specification (we do not propose to adopt the possible solutions in this meeting):

[0330] 1. Instead of using the value of PicSizeInSamplesY (i.e., PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples) to derive MaxDpbSize (i.e., the maximum number of reference pictures that can exist in the DPB), MaxDpbSize is derived based on the value of MinPicSizeInSamplesY. MinPicSizeInSampleY is defined as follows:

[0331] MinPicSizeInSampleY = (width of the smallest picture resolution in the bitstream) * (height of the smallest resolution in the bitstream)

[0332] The derivation of MaxDpbSize is modified as follows (based on the HEVC equation):

[0333]

[0334]

[0335] 2. Each decoded picture is associated with a value called PictureSizeUnit. PictureSizeUnit is an integer value that specifies how large the decoded picture size is relative to MinPicSizeInSampleY. The definition of PictureSizeUnit depends on which resampling rates are supported in VVC for ARC.

[0336] For example, if ARC only supports resampling rates of 2, then PictureSizeUnit is defined as follows:

[0337] - The decoded picture with the lowest resolution in the bitstream is associated with a PictureSizeUnit of 1.

[0338] - The decoded picture with a resolution of 1.5 x 1.5 of the smallest resolution in the bitstream is associated with a PictureSizeUnit of 9 (i.e. 2.25*4).

[0339] As another example, if the ARC supports resampling rates of 1.5 and 2, the PictureSizeUnit is defined as follows:

[0340] - The decoded picture with the lowest resolution in the bitstream is associated with a PictureSizeUnit of 4.

[0341] - The decoded picture with a resolution of 1.5 x 1.5 of the smallest resolution in the bitstream is associated with a PictureSizeUnit of 9 (i.e. 2.25*4).

[0342] - The decoded picture with a resolution of 2x2 of the smallest resolution in the bitstream is associated with a PictureSizeUnit of 16 (i.e. 4*4).

[0343] For other resampling rates supported by the ARC, the value of PictureSizeUnit for each picture size should be determined using the same principle as in the above examples.

[0344] 3. Let the variable MinPictureSizeUnit be the smallest possible value of PictureSizeUnit. That is, if the ARC supports only a resampling rate of 2, MinPictureSizeUnit is 1; if the ARC supports resampling rates of 1.5 and 2, MinPictureSizeUnit is 4; and so on, using the same principle to determine the value of MinPictureSizeUnit.

[0345] 4. The value range of sps_max_dec_pic_buffering_minusl[i] is specified to be in the range of 0 to (MinPictureSizeUnit * (MaxDpbSize - 1)). The variable MinPictureSizeUnit is the smallest possible value of PictureSizeUnit.

[0346] 5. The DPB fullness operation is specified based on PictureSizeUnit as shown below:

[0347] - Initialize the HRD at decoding unit 0 with both the CPB and the DPB set to empty (set the DPB fullness to equal 0).

[0348] - When the DPB is flushed (i.e., all pictures are removed from the DPB), the DPB fullness is set equal to 0.

[0349] - When a picture is removed from the DPB, the DPB fullness is decremented by the value of the PictureSizeUnit associated with the removed picture.

[0350] - When a picture is inserted into the DPB, the DPB fullness is incremented by the value of the PictureSizeUnit associated with the inserted picture.

[0351] 1.2.5.5 Resampling filter

[0352] In software implementation, the implemented resampling filter is simply taken from the previously available filter described in JCTVC-H0234. If other resampling filters have better performance and / or lower complexity, they should be tested and used. We propose to test various resampling filters to trade-off between complexity and performance. Such testing can be done in CE.

[0353] 1.2.5.6 Other necessary modifications to existing tools

[0354] To support ARC, some modifications and / or other operations to certain existing coding tools can be needed. For example, in picture-based resampling of ARC software implementation, we disabled TMVP and ATMVP for simplicity when the original coding resolutions of the current picture and the reference picture are different.

[0355] 1.2.6 JVET-N0279

[0356] According to the “Requirements on future video coding standard”, “In the case of adaptive streaming services that provide multiple representations of the same content, each with different properties (e.g., spatial resolution or sampling bit depth), the standard shall support fast representation switching”. In real-time video communication, allowing resolution change within a coded video sequence without the need to insert I-pictures not only enables the video data to adapt seamlessly to dynamic channel conditions or user preferences, but also eliminates the juddering effect caused by I-pictures. An example of the assumption of adaptive resolution change is shown in FIG. 14 , where a current picture is predicted from a reference picture of different size.

[0357] This contribution proposes high-level syntax to signal adaptive resolution changes and modifications to the prediction process of the current motion compensation in VTM. These modifications are limited to motion vector scaling and sub-pixel position derivation without changing the existing motion compensation interpolator. This will allow reusing the existing motion compensation interpolator without the need for a new processing block to support adaptive resolution changes, which would incur additional costs.

[0358] 1.2.6.1 Adaptive resolution change signaling

[0359] 1.2.6.1.1 SPS

[0360]

[0361]

[0362] [[pic_width_in_luma_samples specifies the width of each decoded picture in units of luma samples. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.]]

[0363] [[pic_height_in_luma_samples specifies the height of each decoded picture in units of luma samples. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.]]

[0364] max_pic_width_in_luma_samples specifies the maximum width of decoded pictures referring to the SPS in units of luma samples. max_pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0365] max_pic_height_in_luma_samples specifies the maximum height of decoded pictures referring to the SPS in units of luma samples. max_pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY.

[0366] 1.2.6.1.2 PPS

[0367]

[0368] pic_size_different_from_max_flag equal to 1 specifies that the PPS signals a different picture width and picture height than max_pic_width_in_luma_samples and max_pic_height_in_luma_sample in the referenced SPS. pic_size_different_from_max_flag equal to 0 specifies that pic_width_in_luma_samples and pic_height_in_luma_sample are the same as max_pic_width_in_luma_samples and max_pic_height_in_luma_sample in the referenced SPS.

[0369] pic_width_in_luma_samples specifies the width of each decoded picture in luma samples. pic_width_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. If pic_width_in_luma_samples is not present, it is inferred to be equal to max_pic_width_in_luma_samples.

[0370] pic_height_in_luma_samples specifies the height of each decoded picture in luma samples. pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of MinCbSizeY. When pic_height_in_luma_samples is not present, it is inferred to be equal to max_pic_height_in_luma_samples.

[0371] A requirement for bitstream conformance is that the horizontal and vertical scaling ratios shall be in the range of 1 / 8 to 2, inclusive, for each active reference picture. The scaling ratios are defined as follows:

[0372] - horizontal_scaling_ratio = ((reference_pic_width_in_luma_samples « 14) + (pic_width_in_luma_samples / 2)) / pic_width_in_luma_samples

[0373] - vertical scaling ratio = ((reference_pic_height_in_luma_samples « 14) + (pic_height_in_luma_samples / 2)) / pic_height_in_luma_samples

[0374]

[0375] Reference picture scaling process

[0376] When a resolution change occurs within a CVS, a picture can have a different size than one or more of its reference pictures. This proposal normalizes all motion vectors to the current picture grid, rather than its corresponding reference picture grid. This is purported to be beneficial for keeping the design consistent and making resolution changes transparent to the motion vector prediction process. Otherwise, due to the different scales, neighboring motion vectors pointing to a reference picture of a different size cannot be directly used for spatial motion vector prediction.

[0377] When a resolution change occurs, both the motion vector and the reference block must be scaled when performing motion-compensated prediction. The scaling range is limited to [1 / 8, 2], i.e., the upscaling by a ratio is limited to 1:8, while the downscaling by a ratio is limited to 2:1. Note that upscaling by a ratio refers to the case where the reference picture is smaller than the current picture, while downscaling by a ratio refers to the case where the reference picture is larger than the current picture. In the following sections, the scaling process will be detailed.

[0378] Luma block

[0379] The scaling factor and its fixed-point representation are defined as

[0380]

[0381]

[0382] The scaling process consists of two parts:

[0383] 1. Mapping the top-left pixel of the current block to the reference picture;

[0384] 2. Using the horizontal and vertical steps to determine the reference locations for the other pixels of the current block.

[0385] If the coordinates of the top-left pixel of the current block are (x, y), the sub-pixel position (x', y') in the reference picture pointed to by the motion vector (mvX, mvY) is specified in 1 / 16-pixel units as follows:

[0386] • The horizontal position in the reference picture is

[0387] x′=((x<<4)+mvX)·hori_scale_fp, (3)

[0388] And x′ is further reduced to retain only 10 decimal places

[0389] x′=Sign(x′)·((Abs(x′)+ (1<<7))》8). (4)

[0390] ●Similarly, the vertical position of the reference image is:

[0391] y′=((y《4)+mvY)·vert_scale_fp, (5)

[0392] And y′ is further reduced to

[0393] y′=Sign(y′)·((Abs(y′)+(1<<7))>>8) (6)

[0394] At this point, the reference position of the top left pixel of the current block is (x', y'). Other reference sub-pixel / pixel positions are calculated relative to (x', y') with horizontal and vertical steps. These steps are derived with 1 / 1024 pixel accuracy based on the horizontal and vertical scaling factors mentioned above, as shown below:

[0395] x_step=(hori_scale_fp+8)>> 4, (7)

[0396] y_step=(vert_scale_fp+8)>>4. (8)

[0397] As an example, if a pixel in the current block is i columns and j rows away from the top-left pixel, the horizontal and vertical coordinates of its corresponding reference pixel can be obtained as follows:

[0398] x′ i =x′+i*x_step, (9)

[0399] y′ j =y′+j*y_step. (10)

[0400] In sub-pixel interpolation, x′ must be i and y′ j Decomposed into full pixel part and fractional pixel part:

[0401] ● The complete pixel portion used to address the reference block is equal to

[0402] (x′ i +32)>>10, (11)

[0403] (y′j +32)>>10. (12)

[0404] The fractional pixel portion used to select the interpolation filter is equal to

[0405] Δx=((x′ i +32)>>6)&15, (13)

[0406] Δy=((y′ j +32)>>6)&15. (14)

[0407] Once the full pixel and fractional pixel positions in the reference picture are determined, the existing motion compensation interpolator can be used without any other changes. The full pixel positions will be used to obtain the reference block patch from the reference picture, while the fractional pixel positions will be used to select the appropriate interpolation filter.

[0408] Chroma Block

[0409] When the chroma format is 4:2:0, the accuracy of the chroma motion vector is 1 / 32 pixel. In addition to the adjustments related to the chroma format, the scaling process of the chroma motion vector and the chroma reference block is almost the same as that of the luma block.

[0410] When the coordinates of the upper left corner pixel of the current chroma block are (x c ,y c ), the initial horizontal and vertical positions in the reference chrominance picture are

[0411] x c ′=((x c <<5)+mvX)·hori_scale_fp, (1)

[0412] y c ′=((y c <<5)+mvY)·vert_scale_fp, (2)

[0413] Where mvX and mvY are the original luma motion vectors, but now should be checked with 1 / 32 pixel precision.

[0414] x c ′ and y c 'Further reduction to maintain 1 / 1024 pixel accuracy

[0415] x c ′=Sign(x c ′)·((Abs(x c ′)+(1<<8))>>9), (3)

[0416] y c ′=Sign(yc ′) · ((Abs(y c ′) + (1 « 8)) » 9). (4)

[0417] The above right shift adds one bit compared to the associated luma equation.

[0418] The step size used is the same for luma. For a chroma sample located at (i, j) relative to the top-left pixel, the horizontal and vertical coordinates of the reference samples are obtained by

[0419]

[0420]

[0421] In sub-pixel interpolation, and are also split into a full-pixel part and a fractional-pixel part:

[0422] • The full-pixel part used to address the reference block is equal to

[0423]

[0424]

[0425] • The fractional-pixel part used to select the interpolation filter is equal to

[0426]

[0427]

[0428] Interaction with other coding tools

[0429] Since the interaction of certain coding tools with the reference picture scaling incurs additional complexity and memory bandwidth, it is proposed to add the following restrictions to the VVC specification:

[0430] - When tile_group_temporal_mvp_enabled_flag is equal to 1, the size of the current picture and its collocated picture shall be the same.

[0431] - When changing resolution is allowed in a sequence, the decoder motion vector refinement shall be turned off.

[0432] - When changing resolution is allowed in a sequence, sps_bdof_enabled_flag shall be equal to 0.

[0433] 1.3 Coding Tree Block (CTB) based Adaptive Loop Filter (ALF) in JVET-N0415

[0434] Strip-level temporal filter

[0435] Adaptive parameter set (APS) is adopted in VTM4. Each APS contains a set of signaled ALF filters, up to 32 APSs are supported. In this proposal, strip-level temporal filter is tested. Tile groups can reuse ALF information from APS to reduce overhead. APS is updated as a first-in-first-out (FIFO) buffer.

[0436] CTB-based ALF

[0437] For luma component, when ALF is applied to luma CTB, it is indicated to select among 16 fixed, 5 temporal or 1 signaled filter set. Only filter setting index is signaled. Only one new set (25 filters) can be signaled for one slice. If a new set is signaled for a slice, all luma CTBs in the same slice share the set. Fixed filter set can be used to predict new slice-level filter set and also as candidate filter set for luma CTB. Total number of filters is 64.

[0438] For chroma component, when ALF is applied to chroma CTB, if a new filter is signaled to the slice, the CTB will use the new filter, otherwise, the latest temporal chroma filter satisfying temporal scalability constraint will be applied.

[0439] As a strip-level temporal filter, APS is updated as a first-in-first-out (FIFO) buffer.

[0440] 1.4 Alternative temporal motion vector prediction (ATMVP) prediction (also referred to as subblock-based temporal Merge candidate in VVC)

[0441] In the ATMVP (alternative temporal motion vector prediction) method, the temporal motion vector prediction (TMVP) of the motion vector is modified by extracting multiple sets of motion information (including motion vector and reference index) from blocks smaller than the current CU. As shown in FIG. 15 The sub-CU is a square N x N block (N is set to 8 by default).

[0442] The ATMVP predicts the motion vector of the sub-CU within the CU in two steps. The first step is to identify the corresponding block in the reference picture using the so-called temporal vector. The reference picture is called the motion source picture. The second step is to divide the current CU into sub-CUs and obtain the motion vector from the block corresponding to each sub-CU and the reference index of each sub-CU, as shown in FIG. 15

[0443] ​In the first step, the reference picture and the corresponding block are determined from the motion information of the spatial neighboring blocks of the current CU. To avoid the repeated scanning process of the neighboring blocks, the Merge candidate from block 0 (left block) in the Merge candidate list of the current CU is used. The first available motion vector from block A0, which refers to the collocated reference picture, is set as the temporal vector. In this way, in ATMVP, the corresponding block, which is sometimes referred to as the collocated block, can be identified more accurately compared to TMVP, where the corresponding block is always located at the right bottom or center position relative to the current CU.

[0444] In the second step, the corresponding block of a sub-CU is identified by the temporal vector in the motion source picture by adding the temporal vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (covering the smallest motion grid of the center sample) is used to derive the motion information of the sub-CU. After the motion information of the corresponding NxN block is identified, it is converted to the reference index and motion vector of the current sub-CU in the same way as TMVP of HEVC, where the motion scaling and other processing also apply.

[0445] 1.5 Affine Motion Prediction

[0446] In HEVC, the motion-compensated prediction (MCP) only applies the translational motion model. However, there can be various motions in the real world, such as zoom-in / zoom-out, rotation, perspective motion, and other irregular motions. In VVC, the simplified affine transform motion-compensated prediction applies to both the 4-parameter affine model and the 6-parameter affine model. As shown in FIG. 16, the affine motion field of a block is described by two control point motion vectors (CPMV) for the 4-parameter affine model and 3 CPMVs for the 6-parameter affine model.

[0447] The motion vector field (MVF) of a block is described by the following equations with the 4-parameter affine model in Equation (1) (where the 4 parameters are defined as variables a, b, c, d, e, and f) and the 6-parameter affine model in Equation (2) (where the 4 parameters are defined as variables a, b, e, and f):

[0448]

[0449]

[0450] where (mv h 0, mv h 0) is the motion vector of the top-left control point, (mv h 1, mv h 1) is the motion vector of the top-right control point, and (mv h 2, mv h2) is the motion vector of the bottom-left corner motion vector control point, all these three motion vectors are called control point motion vectors (CPMV), (x, y) represents the coordinates of the representative point relative to the top-left sample within the current block, and (mv h (x,y),mv v (x,y)) is the motion vector derived for the sample located at (x, y). CP motion vectors can be signaled (e.g. in affine AMVP mode) or derived on the fly (e.g. in affine Merge mode). w and h are the width and height of the current block. In practice, the division is implemented by taking the integer and right shifting. In VTM, the representative point is defined as the center position of the sub-block, e.g. when the top-left corner of the sub-block has coordinates (xs, ys) relative to the top-left sample within the current block, the coordinates of the representative point are defined as (xs+2, ys+2). For each sub-block (i.e. 4x4 in VTM), the representative point is used to derive the motion vector for the whole sub-block.

[0451] To further simplify the motion compensation prediction, a sub-block based affine transform prediction is applied. To derive the motion vector for each MxN (in current VVC, both M and N are set to 4) sub-block, as shown in FIG. 17 Equations (1) and (2), the motion vector of the center sample of each sub-block is calculated and is rounded to the 1 / 16 fractional precision. Then, a motion compensation interpolation filter suitable for 1 / 16 pixel is applied to generate the prediction for each sub-block with the derived motion vector. The affine mode introduces the 1 / 16 pixel interpolation filter.

[0452] After MCP, the high-precision motion vector for each sub-block is rounded and saved to the same precision as the regular motion vector.

[0453] 1.5.1 Signaling of affine prediction

[0454] Similar to the translational motion model, due to affine prediction, there are also two modes for signaling the side information. They are AFFINE INTER and AFFINE MERGE modes.

[0455] 1.5.2 AF INTER mode

[0456] For a CU with both width and height larger than 8, AF INTER mode can be applied. An affine flag at CU level is signaled in the bitstream to indicate whether AF INTER mode is used.

[0457] In this mode, for each reference picture list (list 0 or list 1), the affine AMVP candidate list consists of three types of affine motion predictors in the following order, where each candidate includes the estimated CPMV of the current block. The difference between the best CPMV found on the encoder side (such as mv0, mv1, mv2 in Figure 18) and the estimated CPMV is signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is further signaled.

[0458] (1) Inherited affine motion predictor

[0459] The checking order is similar to the spatial MVP in HEVC AMVP list construction. First, the left inherited affine motion predictor is derived from the first block in {A1, A0}, which is affine coded and has the same reference picture as the current block. Second, the above inherited affine motion predictor is derived from the first block in {B1, B0, B2}, which is affine coded and has the same reference picture as the current block. FIG. 19 Five blocks A1, A0, B1, B0, B2 are shown in FIG.

[0460] Once a neighboring block is found to be coded in affine mode, the CPMV of the coding unit covering the neighboring block is used to derive the predictor of the CPMV of the current block. For example, if A1 is coded in non-affine mode and A0 is coded in 4-parameter affine mode, the left-inherited affine MV predictor will be derived from A0. In this case, the CPMV of the CU covering A0, such as FIG. 21B CPMV in the upper left corner and CPMV in the upper right corner As shown, the estimated CPMV for deriving the current block is composed of the upper left (coordinate (x0, y0)), upper right (coordinate (x1, y1)) and lower right (coordinate (x2, y2)) positions of the current block. express.

[0461] (2) Constructed affine motion predictor

[0462] The constructed affine motion predictor consists of control point motion vectors (CPMVs) derived from adjacent inter-coded blocks with the same reference picture, as FIG. 20 If the current affine motion model is 4-parameter affine, the number of CPMVs is 2. Otherwise, if the current affine motion model is 6-parameter affine, the number of CPMVs is 3. The CPMV in the upper left corner It is derived from the MV of the first block in the group {A, B, C}, which is inter-coded and has the same reference picture as the current block. is derived from the MV at the first block in the group {D, E} which is inter coded and has the same reference picture as the current block. The bottom-left CPMV is derived from the MV at the first block in the group {F, G} which is inter coded and has the same reference picture as the current block.

[0463] - If the current affine motion model is 4-parameter affine, the constructed affine motion predictor is inserted into the candidate list only if and are known, i.e. if and are used as the estimated CPMVs for the top-left (coordinates (x0, y0)), top-right (coordinates (x1, y1)) and bottom-right (coordinates (x2, y2)) positions of the current block.

[0464] - If the current affine motion model is 6-parameter affine, the constructed affine motion predictor is inserted into the candidate list only if and are known, i.e. if and are used as the estimated CPMVs for the top-left (coordinates (x0, y0)), top-right (coordinates (x1, y1)) and bottom-right (coordinates (x2, y2)) positions of the current block.

[0465] When the constructed affine motion predictor is inserted into the candidate list, no clipping process is applied.

[0466] 1) Regular AMVP motion predictor

[0467] The following applies until the number of affine motion predictors reaches the maximum value.

[0468] 1) Derive an affine motion predictor by setting all CPMVs equal to (if available).

[0469] 2) Derive an affine motion predictor by setting all CPMVs equal to (if available).

[0470] 3) Derive an affine motion predictor by setting all CPMVs equal to (if available).

[0471] 4) Derive an affine motion predictor by setting all CPMVs equal to the HEVC TMVP (if available).

[0472] 5) Derive an affine motion predictor by setting all CPMVs to zero MV.

[0473] Note that, has been derived in the constructed affine motion predictor.

[0474] In AF_INTER mode, when 4 / 6 parameter affine mode is used, 2 / 3 control points are needed, thus, 2 / 3 MVDs need to be coded for these control points, as shown in FIG. 18A and 18B In JVET-K0337, it is proposed to derive MVs as follows, i.e., predict mvd1 and mvd2 from mvd0.

[0475]

[0476]

[0477]

[0478] where mvd i and mv1 are the predicted motion vector, the motion vector difference and the motion vector of the top-left pixel (i = 0), the top-right pixel (i = 1) or the bottom-left pixel (i = 2), respectively, as shown in FIG. 18B Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two components, respectively, i.e., newMV = mvA + mvB, and the two components of newMV are set to (xA + xB) and (yA + yB), respectively.

[0479] 1.5.2.1 AF_MERGE mode

[0480] When a CU is applied in AF_MERGE mode, it obtains the first block encoded in affine mode from the valid neighbor reconstructed blocks. And the selection order of the candidate blocks is from left, top, top-right, bottom-left to top-left (denoted by A, B, C, D, E in turn), as shown in FIG. 21A For example, if the neighbor bottom-left block is encoded in affine mode, as shown in A0 in FIG. 21B , the control point (CP) motion vectors mv0 N , mv1 N and mv2 N of the top-left, top-right and bottom-left corners of the neighboring CU / PU including block A are obtained. Based on mv0 N , mv1 N and mv2 N , the motion vectors mv0 C , mv1 C and mv2 C of the top-left, top-right and bottom-left corners of the current CU / PU are calculated.(Only for 6-parameter affine models.) It should be noted that in VTM-2.0, if the current block is affine-coded, the sub-block at the top left (e.g., a 4×4 block in VTM) stores mv0, and the sub-block at the top right stores mv1. If the current block is coded using a 6-parameter affine model, the sub-block at the bottom left stores mv2; otherwise (the current block is coded using a 4-parameter affine model), LB stores mv2'. The other sub-blocks store the MV used for MC.

[0481] In deriving the CPMVmv0 of the current CU C 、mv1 C and mv2 C Afterwards, the MVF of the current CU is generated according to the simplified affine motion model equations (1) and (2). In order to identify whether the current CU is coded in AF_MERGE mode, an affine flag is signaled in the bitstream when at least one neighboring block is coded in affine mode.

[0482] In JVET-L0142 and JVET-L0632, the affine merge candidate list is constructed by the following steps:

[0483] 1) Insert inherited affine candidates

[0484] An inherited affine candidate is one that is derived from the affine motion model of its valid neighboring affine coded blocks. Up to two inherited affine candidates are derived from the affine motion models of neighboring blocks and inserted into the candidate list. For the left predictor, the scan order is {A0, A1}; for the above predictors, the scan order is {B0, B1, B2}.

[0485] 2) Insert the constructed affine candidate

[0486] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (e.g., 5), the constructed affine candidate is inserted into the candidate list. Constructing an affine candidate means constructing a candidate by combining the neighbor motion information of each control point.

[0487] a) First, FIG. 22 The motion information of the control point is derived from the specified spatial and temporal neighbors shown. CPk (k = 1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial locations used to predict CPk (k = 1, 2, 3); T is the temporal location used to predict CP4.

[0488] The coordinates of CP1, CP2, CP3 and CP4 are (0, 0), (W, 0), (H, 0) and (W, H), respectively, where W and H are the width and height of the current block.

[0489] The motion information of each control point is obtained according to the following priority order:

[0490] - For CP1, the priority is checked as B2 -> B3 -> A2. If B2 is available, B2 is used. Otherwise, if B2 is not available, B3 is used. If neither B2 nor B3 is available, A2 is used.

[0491] If none of the three candidates is available, the motion information of CP1 cannot be obtained.

[0492] - For CP2, the priority is checked as B1 -> B0;

[0493] - For CP3, the priority is checked as A1 -> A0;

[0494] - For CP4, T is used.

[0495] b) Secondly, affine Merge candidates are constructed using combinations of control points.

[0496] I. Three control points' motion information is needed to construct a 6-parameter affine candidate. The three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}). The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4} will be converted to a 6-parameter motion model represented by top-left, top-right and bottom-left control points.

[0497] II. Two control points' motion information is needed to construct a 4-parameter affine candidate. The two control points can be selected from one of the two combinations ({CP1, CP2}, {CP1, CP3}). The two combinations will be converted to a 4-parameter motion model represented by top-left and top-right control points.

[0498] III. The combinations of constructed affine candidates are inserted into the candidate list in the following order:

[0499] {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}

[0500] i. For each combination, check the reference indices of the list X of each CP, if they are all the same, this combination has a valid CPMV for the list X. If the combination has no valid CPMV for both list 0 and list 1, the combination is marked as invalid. Otherwise, it is valid and put the CPMV into the subblock Merge list.

[0501] 3) Fill with zero motion vector

[0502] If the number of candidates in the affine Merge candidate list is less than 5, insert a zero motion vector with zero reference index into the candidate list until the list is full.

[0503] More specifically, for subblock Merge candidate list, a 4-parameter Merge candidate with MV set to (0, 0) and prediction direction set to uni-prediction for P slice (for P slice) and bi-prediction for B slice (for B slice) of list 0.

[0504] 2. Drawbacks of existing implementations

[0505] When applied in VVC, the following issues can occur for ARC:

[0506] 1. It is not clear how to signal the information related to ARC.

[0507] 2. It is not clear how to apply affine / TMVP or ATMVP when the resolutions of the reference picture, collocated picture and current picture are different.

[0508] 3. The design of down-sampling or up-sampling filter for ARC can be better designed.

[0509] 3. Example method of adaptive resolution conversion

[0510] The following detailed inventions should be considered as examples to explain the general concepts. These inventions should not be interpreted in a narrow way. Furthermore, these inventions can be combined in any way.

[0511] In the following discussion, SatShift(x, n) is defined as

[0512]

[0513] Shift(x, n) is defined as Shift(x, n) = (x + offset0) » n.

[0514] In one example, offset0 and / or offset1 is set to (1 « n) » 1 or (1 « (n - 1)). In another example, offset0 and / or offset1 is set to 0.

[0515] In another example, offset0 = offset1 = ((1 « n) » 1) - 1 or ((1 « (n - 1)) - 1).

[0516] Clip3(min, max, x) is defined as

[0517]

[0518] Floor(x) is defined as the largest integer less than or equal to x.

[0519] Ceil(x) is the smallest integer greater than or equal to x.

[0520] Log2(x) is defined as the logarithm base 2 of x.

[0521] Signaling of ARCs

[0522] 1. Propose to signal the picture dimension information (width and / or height) related to ARC in a video unit other than DPS, VPS, SPS, PPS, APS, picture header, slice header, tile group header.

[0523] a. In one example, the picture dimension information related to ARC can be signaled in a Supplemental Enhancement Information (SEI) message.

[0524] b. In one example, the picture dimension information related to ARC can be signaled in a separate video unit for ARC. For example, the video unit can be named Resolution Parameter Set (RPS) or Conversion Parameter Set (CPS), or any other name.

[0525] i. In one example, more than one combination of width and height is signaled in a separate video unit for ARC, such as the RPS or CPS to be named.

[0526] 2. Propose not to signal the picture size (width or height) in 0th order Exponential Golomb code.

[0527] a. In one example, it can be encoded with a fixed length code or a unary code.

[0528] b. In one example, it can be encoded with a Kth (K>0) order Exponential Golomb code.

[0529] c. The dimension can be signaled in a video unit, such as DPS, VPS, SPS, PPS, APS, picture header, slice header, tile group header, etc., or in a separate video unit for ARC, such as the RPS or CPS to be named.

[0530] 3. Propose to signal the resolution ratio instead of multiple resolutions.

[0531] a. In one example, an indication of a base resolution can be signaled. In addition, an indication of allowed aspect ratio combinations (such as horizontal aspect ratio, vertical aspect ratio) can be further signaled.

[0532] b. In one example, an index of an indication of allowed aspect ratio combinations can be signaled in PPS to indicate the actual resolution of a picture.

[0533] 4. It is proposed that when more than one combination of picture width and height is signaled in a single video unit (such as DPS, VPS, SPS, PPS, APS, picture header, slice header, tile group header, etc.), or in a separate video unit (such as RPS or CPS to be named) for ARC, the width and height in the first combination are not allowed to be equal to the width and height in the second combination.

[0534] 5. It is proposed that the signaled dimensions (width W and height H) must be subject to restrictions.

[0535] a. For example, W should satisfy TW_min <= W <= TW_max.

[0536] b. For example, H should satisfy TH_min <= H <= TH_max.

[0537] c. In one example, W - TW_min - B can be signaled, where B is a fixed value such as 0.

[0538] d. In one example, H - TH_min - B can be signaled, where B is a fixed value such as 0.

[0539] e. In one example, TW_max - W - B can be signaled, where B is a fixed value such as 0.

[0540] f. In one example, TH_max - H - B can be signaled, where B is a fixed value such as 0.

[0541] g. In one example, TW_min and / or TH_min can be signaled.

[0542] h. In one example, TW_max and / or TH_max can be signaled.

[0543] 6. It is proposed that the signaled dimensions (width W and height H) must be in the form of W = w * X and H = h * Y, where X and Y are predefined integers, for example, X = Y = 4.

[0544] a. In one example, w and h are signaled. W and H are derived from w and h.

[0545] 7. It is proposed that picture dimension information (width and / or height) can be coded in a predictive manner.

[0546] a. In one example, the difference between the first width (W1) and the second width (W2), i.e. W2-W1, can be signaled.

[0547] i. Alternatively, W2-W1-B can be signaled, where B is a fixed value such as 1.

[0548] ii. In one example, W2 should be greater than W1.

[0549] iii. In one example, the difference can be coded using unary code, or truncated unary code, or fixed length code or fixed length coding.

[0550] b. In one example, the difference between the first height (H1) and the second height (H2), i.e. H2-H1, can be signaled.

[0551] i. Alternatively, H2-H1-B can be signaled, where B is a fixed value such as 1.

[0552] ii. In one example, H2 should be greater than H1.

[0553] iii. In one example, the difference can be coded using unary code, or truncated unary code, or fixed length code or fixed length coding.

[0554] c. In one example, the ratio between the first width (W1) and the second width (W2), i.e. W2 / W1, can be signaled. For example, if W = F*W1, then F is signaled. In another example, if W2 = Shift(F*W1, P), then F is signaled, where P is a number representing precision, e.g. P = 10.

[0555] i. Alternatively, F can be equal to (W2*P+W1 / 2) / W1, where P is a number representing precision, e.g. P = 10.

[0556] ii. Alternatively, F-B can be signaled, where B is a fixed value such as 1.

[0557] iii. In one example, W2 should be greater than W1.

[0558] iv. In one example, F can be coded using unary code, or truncated unary code, or fixed length code or fixed length coding.

[0559] d. In one example, the ratio between the first height (H1) and the second height (H2), i.e. H2 / H1, can be signaled. For example, if H2 = F*H1, then F is signaled. In another example, if H2 = Shift(F*H1, P), then F is signaled, where P is a number representing precision, e.g. P = 10.

[0560] i. Alternatively, F can be equal to (H2*P+H1 / 2) / H1, where P is a number representing precision, e.g. P = 10.

[0561] ii. Alternatively, F-B can be signaled, where B is a fixed value such as 1.

[0562] iii. In one example, H2 should be greater than H1.

[0563] iv. In one example, the difference can be coded with unary code, or truncated unary code, or fixed length code or fixed length coding.

[0564] e. In one example, W2 / W1 must be equal to H2 / H1, and only one of W2 / W1 or H2 / H1 is signaled.

[0565] 8. It is proposed that when different resolutions / resolution ratios are signaled, the following additional syntax elements can be further signaled.

[0566] a. The syntax element can be an indication of the CTU size.

[0567] b. The syntax element can be an indication of the minimum coding unit size.

[0568] c. The syntax element can be an indication of the maximum and / or minimum transform block size.

[0569] d. The syntax element can be an indication of the maximum depth of quad-tree and / or binary / trinary tree.

[0570] e. In one example, the additional syntax elements can be tied to a specific picture resolution.

[0571] Reference picture lists

[0572] 9. A conforming bitstream shall satisfy the following requirement: a reference picture with a different resolution than the current picture shall be assigned a larger reference index compared to a reference picture with the same resolution as the current picture.

[0573] a. Alternatively, before decoding one picture / slice / tile / slice group, the reference picture list can be reordered such that for the reference picture list, a reference picture with different resolution than the current picture should be assigned a larger reference index than a reference picture with the same resolution as the current picture.

[0574] Motion prediction from temporal blocks using ARCs (e.g., TMVP and ATMVP)

[0575] 10. Suppose there are two blocks, A and B. If the reference picture of block A is a reference picture with the same resolution as the current block, and the reference picture of block B is a reference picture with different resolution than the current block, it is proposed to prohibit using the motion information of block B to predict block A.

[0576] a. If the reference picture of block A is a reference picture with different resolution than the current block, and the reference picture of block B is a reference picture with the same resolution as the current block, it is proposed to prohibit using the motion information of block B to predict block A.

[0577] 11. It is proposed to prohibit prediction from a block in a reference picture with different resolution than the current picture.

[0578] 12. It is proposed that a reference picture cannot be a collocated reference picture if its width is different from the width of the current picture or its height is different from the height of the current picture.

[0579] b. Alternatively, if a reference picture's width is different from the width of the current picture and its height is different from the height of the current picture, the reference picture cannot be a collocated reference picture.

[0580] 13. How to find a collocated block in TMVP / ATMVP or a collocated sub-block in ATMVP can depend on whether the collocated reference picture and the current picture have the same picture width and height.

[0581] a. In one example, assuming the current picture dimension is W0*H0 and the collocated reference picture dimension is W1*H1, the position and / or dimension of the collocated block can depend on W0, H0, W1 and H1.

[0582] i. In one example, assuming the top-left coordinates of the current block or sub-block are (x, y), the collocated block or sub-block can be derived as a block in the collocated reference picture that covers the position (x', y'), where (x', y') can be calculated as x' = Rx*(x + offsetX) + offsetX' and y' = Ry*(y + offsetY) + offsetY'. The values of offsetX' and offsetY' are 0.

[0583] 1) In one example, (offsetX, offsetY) can be equal to (x+w, y+h) assuming the dimension of the current block or the current sub-block is w*h.

[0584] a) In an alternative example, (offsetX, offsetY) can be equal to (x+w / 2, y+h / 2).

[0585] b) In an alternative example, (offsetX, offsetY) can be equal to (x+w / 2-1, y+h / 2-1).

[0586] c) In an alternative example, (offsetX, offsetY) can be equal to (0, 0).

[0587] 2) In one example, Rx = W1 / W0.

[0588] 3) In one example, Ry = H1 / H0.

[0589] 4) In another example, x' = Shift(Rx*(x+offsetX), P), where P is a value representing precision, such as 10.

[0590] a) Rx can be derived as Rx = (W1*P+offset) / W0, where offset is an integer, such as 0 or W0 / 2.

[0591] 5) In another example, y' = Shift(Ry*(y+offsetY), P), where P is a value representing precision, such as 10.

[0592] a) Ry can be derived as Ry = (H1*P+offset) / H0, where offset is an integer, such as 0 or H0 / 2.

[0593] 14. In addition to upsampling or downsampling the reference samples in the reference picture at a resolution different from the current picture, it is proposed to further upsample or downsample the motion information / code information, and the upsampled and downsampled information can be used to encode subsequent blocks in other frames.

[0594] 15. It is proposed that if the width or height of the collocated reference picture is different from the width or height of the current picture, the buffer storing the MVs of the collocated reference picture can be upsampled or downsampled.

[0595] b. In one example, the upsampled or downsampled MV buffer and the MV buffer before upsampled or downsampled can be stored simultaneously.

[0596] i. Alternatively, the MV buffer before upsampled or downsampled can be removed.

[0597] c. In one example, the plurality of MVs in the upsampled MV buffer is copied from one MV in the MV buffer before upsampling.

[0598] i. For example, the plurality of MVs in the upsampled MV buffer can be in a region corresponding to the region of the MVs in the MV buffer before upsampling.

[0599] d. In one example, the one MV in the downsampled MV buffer can be selected from one of the plurality of MVs in the MV buffer before downsampling.

[0600] i. For example, the one MV in the downsampled MV buffer can be in a region corresponding to the region of the plurality of MVs in the MV buffer before downsampling.

[0601] 16. It is proposed that the derivation of the temporal MV in ATMVP can depend on the dimensions of the current picture W0*H0 and the collocated picture W1*H1.

[0602] c. For example, if the dimensions of the collocated picture and the dimensions of the current picture are different, the temporal MV denoted as tMV can be converted to tMV’.

[0603] i. For example, assuming tMV = (tMVx, tMVy), tMV’ = (tMVx’, tMVy’), tMVx’ can be calculated as tMVx’ = Rx*tMVx + offsetx, and tMVy’ can be calculated as tMVy’ = Ry*tMVy + offsety. offsetx and offsety are values such as 0.

[0604] 1) In one example, Rx = W1 / W0.

[0605] 2) In one example, Ry = H1 / H0.

[0606] 3) In another example, tMVx’ = Shift(Rx*(tMVx+offsetX), P) or SatShift(Rx*(tMVx+offsetX), P), where P is a value denoting precision, such as 10.

[0607] a) Rx can be derived as Rx = (W1*P+offset) / W0, where offset is an integer, such as 0 or W0 / 2.

[0608] 4) In another example, tMVy’ = Shift(Ry*(tMVy+offsetY), P) or SatShift(Ry*(tMVy+offsetY), P), where P is a value denoting precision, such as 10.

[0609] a) Ry can be derived as Ry = (H1 * P + offset) / H0, where offset is an integer, such as 0 or H0 / 2.

[0610] 17. The derivation of the MV prediction (MVP) of the current block or the current sub-block in TMVP / ATMVP can depend on the dimensions of the current picture, and / or the dimensions of the current picture referred by the MVP, and / or the dimensions of the collocated picture, and / or the dimensions of the reference picture of the collocated picture (referenced by the collocated MV). The collocated MV means the MV found in the collocated block. FIG. 23 An example is shown, where the current picture (CurPic), the reference picture referred by the MVP (RefPic), the collocated picture (ColPic), and the reference picture of the collocated picture referred by the collocated MV (RefColPic) are located at time (or with POC) TO, T1, T2, T3, respectively. And their dimensions are W0*H0, W1*H1, W2*H2, and W3*H3, respectively. The MVP of the current block / sub-block is denoted as MvCur = (MvCurX, MvCurY), and the collocated MV is denoted as MvCol (MvColX, MvColY).

[0611] d. MvCurX can be calculated as MvCurX = Rx * MvColX + offsetx, and MvCurY can be calculated as MvCurY = Ry * MvColY + offsety. offsetx and offsety are values such as 0.

[0612] e. In one example, Rx = W0 / W2.

[0613] f. In one example, Ry = H0 / H2.

[0614] g. In one alternative example, MvCurX = Shift(Rx * (MvColX + offsetX), P) or MvCurX = SatShift(Rx * (MvColX + offsetX), P), where P is a value representing precision, such as 10.

[0615] i. Rx can be derived as Rx = (W0 * P + offset) / W2, where offset is an integer, such as 0 or W2 / 2.

[0616] h. In one alternative example, MvCurY = Shift(Ry * (MvColY + offsetY), P) or MvCurY = SatShift(Ry * (MvColY + offsetY), P), where P is a value representing precision, such as 10.

[0617] i. Ry can be derived as Ry = (H0*P + offset) / H2, where offset is an integer such as 0 or H2 / 2.

[0618] i. In one example, (W3, H3) must be equal to (W2, H2). Otherwise, MvCol can be considered as unavailable.

[0619] j. In one example, (W0, H0) must be equal to (W1, H1), otherwise, MVCur can be considered as unavailable.

[0620] Interpolation and scaling in ARCs

[0621] 18. It is proposed that one or more down-sampling or up-sampling filter methods for ARC can be signaled in a video unit such as DPS, VPS, SPS, PPS, APS, picture header, slice header, tile group header, etc., or in a separate video unit for ARC such as RPS or CPS to be named.

[0622] 19. In the approach of JVET-N0279, hori_scale_fp and / or vert_scale_fp can be signaled in a video unit such as DPS, VPS, SPS, PPS, APS, picture header, slice header, tile group header, etc., or in a separate video unit for ARC such as RPS or CPS to be named.

[0623] 20. It is proposed that any division operation in the scaling methods in RAC such as the derivation of hori_scale_fp and / or vert_scale_fp, or the derivation of Rx and / or Ry in bullet 8 - bullet 11 can be replaced or approximated by one or more operations where one or more tables can be used. For example, the approach disclosed in P1905289301 can be used to replace or approximate the division operation.

[0624] The above examples can be incorporated in the context of the methods described below, for example, the methods 2400, 2410, 2420, 2430, 2440, 2450, 2460, 2470, 2480, and 2490 that can be implemented at a video decoder or a video encoder.

[0625] FIG. 24AA flowchart illustrating an example method for video processing is shown. The method 2400 includes, at step 2402, performing a conversion between a video comprising one or more video segments comprising one or more video units and a bitstream representation of the video, in some embodiments, the bitstream representation conforms to a format rule and includes information related to an adaptive resolution conversion (ARC) process, the format rule specifies applicability of the ARC process to a video segment, and includes, in a syntax structure other than a header syntax structure, a decoder parameter set (DPS), a video parameter set (VPS), a picture parameter set (PPS), a sequence parameter set (SPS), and an adaptation parameter set (APS), an indication that the one or more video units of the video segment are encoded at different resolutions.

[0626] In some embodiments, the bitstream representation conforms to a format rule and includes information related to an adaptive resolution conversion (ARC) process, a dimension of the one or more video units encoded with a K-th order exponential Golomb code is signaled in the bitstream representation, K is a positive integer, the format rule specifies applicability of the ARC process to a video segment, and an indication that the one or more video units of the video segment are encoded at different resolutions is included in the bitstream representation in a syntax structure.

[0627] In some embodiments, the bitstream representation conforms to a format rule and includes information related to an adaptive resolution conversion (ARC) process, a height (H) and a width (W) are signaled in the bitstream representation, H and W are positive integers and are constrained, the format rule specifies applicability of the ARC process to a video segment, and an indication that the one or more video units of the video segment are encoded at different resolutions is included in the bitstream representation in a syntax structure.

[0628] FIG. 24B A flowchart illustrating an example method for video processing is shown. The method 2410 includes, at step 2412, determining (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is the same as a resolution of a current picture, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is different from the resolution of the current picture.

[0629] The method 2410 includes, at step 2414, as a result of the determination, performing a conversion between the current video block and a bitstream representation of the video by disabling motion information of the second temporal neighboring block in a prediction of the first temporal neighboring block.

[0630] FIG. 24CA flowchart illustrating an exemplary method for video processing is shown. The method 2420 includes, in step 2422, determining that (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is different from a resolution of a current picture, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is the same as the resolution of the current picture.

[0631] The method 2420 includes, in step 2424, in response to the determination, performing a conversion between the current video block and a bitstream representation of the video by disabling, in a prediction of the first temporal neighboring block, motion information of the second temporal neighboring block.

[0632] FIG. 24D A flowchart illustrating an exemplary method for video processing is shown. The method 2430 includes, in step 2432, for a current video block of a video, determining that a resolution of a reference picture comprising a video block associated with the current video block is different from a resolution of a current picture comprising the current video block.

[0633] The method 2430 includes, in step 2434, in response to the determination, performing a conversion between the current video block and a bitstream representation of the video by disabling a prediction process based on the video block in the reference picture.

[0634] FIG. 24E A flowchart illustrating an exemplary method for video processing is shown. The method 2440 includes, in step 2442, making a decision as to whether a picture is allowed to be used as a collocated reference picture for a current video block of a current picture based on at least one dimension of the picture.

[0635] The method 2440 includes, in step 2444, in response to the decision, performing a conversion between the current video block of the video and a bitstream representation of the video.

[0636] FIG. 24F A flowchart illustrating an exemplary method for video processing is shown. The method 2450 includes, in step 2452, identifying a collocated block for a prediction of a current video block of a video based on a determination that a dimension of a collocated reference picture comprising the collocated block is the same as a dimension of a current picture comprising the current video block.

[0637] The method 2450 includes, in step 2454, performing a conversion between the current video block and a bitstream representation of the video using the collocated block.

[0638] FIG. 24G A flowchart illustrating an exemplary method for video processing is shown. The method 2460 includes, in step 2462, for a current video block of a video, determining that a reference picture associated with the current video block has a resolution that is different from a resolution of a current picture comprising the current video block.

[0639] The method 2460 includes, in step 2464, performing an upsampling operation or a downsampling operation on one or more reference samples of the reference picture and motion information of the current video block or coding information of the current video block as part of the conversion between the current video block and the bitstream representation of the video.

[0640] FIG. 24H A flowchart of an example method for video processing is shown. The method 2470 includes, in step 2472, determining, for a conversion between a current video block of a video and a bitstream representation of the video, that a height or a width of a current picture including the current video block is different from a height or a width of a collocated reference picture associated with the current video block.

[0641] The method 2470 includes, in step 2474, based on the determination, performing an upsampling operation or a downsampling operation on a buffer storing one or more motion vectors of the collocated reference picture.

[0642] FIG. 24I A flowchart of an example method for video processing is shown. The method 2480 includes, in step 2482, based on dimensions of a current picture including a current video block of a video and dimensions of a collocated picture associated with the current video block, deriving information for an optional temporal motion vector prediction (ATMVP) process applied to the current video block.

[0643] The method 2480 includes, in step 2484, using the temporal motion vector to perform the conversion between the current video block and the bitstream representation of the video.

[0644] FIG. 24J A flowchart of an example method for video processing is shown. The method 2490 includes, in step 2492, configuring a bitstream representation of a video for an application of an adaptive resolution conversion (ARC) process to a current video block of the video. In some embodiments, information related to the ARC process is signaled in the bitstream representation, including that a current picture of the current video block has a first resolution and the ARC process includes resampling a portion of the current video block at a second resolution different from the first resolution.

[0645] The method 2490 includes, in step 2494, based on the configuration, performing the conversion between the current video block and the bitstream representation of the current video block.

[0646] 4. Example implementations of the disclosed technology

[0647] FIG. 25is a block diagram of a video processing device 2500. The device 2500 can be used to implement one or more methods described herein. The device 2500 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 2500 can include one or more processors 2502, one or more memories 2504, and video processing hardware 2506. The processor(s) 2502 can be configured to implement one or more methods described in the present document, including but not limited to methods 1500, 1600, and 1700. The memory(ies) 2504 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 2506 can be used to implement, in hardware circuitry, some of the techniques described in the present document.

[0648] In some embodiments, a video encoding method can use information about FIG. 25 The described apparatus implemented on a hardware platform can be implemented.

[0649] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, an encoder will use or implement the tool or mode in processing of a video block, but does not necessarily modify a resulting bitstream based on the use of the tool or mode. That is, a conversion from a video block to a bitstream representation of a video will use the video processing tool or mode based on the decision or determination to enable the video processing tool or mode. In another example, when a video processing tool or mode is enabled, a decoder will process a bitstream based on the knowledge that the bitstream has been modified based on the video processing tool or mode being enabled. That is, a conversion from a bitstream representation of a video to a video block will be performed using the video processing tool or mode enabled based on the decision or determination.

[0650] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In an example, when a video processing tool or mode is disabled, an encoder will not use the tool or mode in a conversion of a video block to a bitstream representation of a video. In another example, when a video processing tool or mode is disabled, a decoder will process a bitstream based on the knowledge that the bitstream has not been modified using the video processing tool or mode enabled based on the decision or determination.

[0651] FIG. 26is a block diagram illustrating an example video processing system 2600 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 2600. The system 2600 can include an input 2602 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values) or can be received in a compressed or encoded format. The input 2602 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0652] The system 2600 can include an encoding component 2604 that can implement various encoding or coding methods described herein. The encoding component 2604 can reduce the average bitrate of video from the input 2602 to the output of the encoding component 2604 to generate an encoded representation of the video. Thus, the encoding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the encoding component 2604 can be stored or transmitted via a connected communication as represented by component 2606. The stored or transmitted bitstream (or encoded) representation of the video received at the input 2602 can be used by component 2608 to generate pixel values or displayable video that is sent to a display interface 2610. The process of generating user- visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “encoding” operations or tools, it should be understood that the encoding tools or operations are used at an encoder and that corresponding decoding tools or operations that reverse the results of the encoding will be performed by a decoder.

[0653] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or Displayport, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described herein can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.

[0654] In some embodiments, the following technical solutions can be implemented:

[0655] A1. A method of video processing, comprising performing a conversion between a video comprising one or more video segments and a bitstream representation of the video, the video segments comprising one or more video units, wherein the bitstream representation conforms to a format rule and includes information related to an adaptive resolution conversion (ARC) process, wherein a dimension of one or more video units encoded with a K-th order exponential Golomb code is signaled in the bitstream representation, wherein K is a positive integer, wherein the format rule specifies an applicability of the ARC process to the video segments, and wherein an indication of one or more video units of the video segments being encoded at different resolutions is included in the bitstream representation in a syntax structure.

[0656] A2. The method of solution A1, wherein the dimension comprises at least one of a width of a video unit in the one or more video units and a height of the video unit.

[0657] A3. The method of solution A1, wherein the one or more video units comprise a picture.

[0658] A4. The method of solution A1, wherein the syntax structure is a decoder parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, or a tile group header.

[0659] A5. The method of solution A1, wherein the syntax structure is a resolution parameter set (RPS) or a conversion parameter set (CPS).

[0660] A6. A method for video processing, comprising performing a conversion between a video comprising one or more video segments and a bitstream representation of the video, the video segments comprising one or more video units, wherein the bitstream representation conforms to a format rule and includes information related to an adaptive resolution conversion (ARC) process, wherein a height (H) and a width (W) of a video unit of the one or more video units are signaled in the bitstream representation, wherein H and W are positive integers and are constrained, wherein the format rule specifies an applicability of the adaptive resolution conversion (ARC) process to the video segments, and wherein an indication of one or more video units of the video segments being encoded at different resolutions is included in the bitstream representation in a syntax structure.

[0661] A7. The method of solution A6, wherein W < TW max , and wherein TW max is a positive integer.

[0662] A8. The method of solution A7, wherein TW max is signaled in the bitstream representation.

[0663] A9. The method of solution A6, wherein TW min ≤ W, and wherein TW min is a positive integer.

[0664] A10. The method of solution A9, wherein TW min is signaled in the bitstream representation.

[0665] A11. The method of solution A6, wherein H ≤ TH max , and wherein TH max is a positive integer.

[0666] A12. The method of solution A11, wherein TH max is signaled in the bitstream representation.

[0667] A13. The method of solution A6, wherein TH min ≤ H, and wherein TH min is a positive integer.

[0668] A14. The method of solution A13, wherein TH min is signaled in the bitstream representation.

[0669] A15. The method of solution A6, wherein a height H = h x Y and a width W = w x X, wherein w, h, X, and Y are positive integers, and wherein w and h are signaled in the bitstream representation.

[0670] A16. The method of solution A15, wherein X = Y = 4.

[0671] A17. The method of solution A15, wherein X and Y are pre-defined integers.

[0672] A18. The method of solution A6, wherein the one or more video units comprise a picture.

[0673] A19. A method for video processing, comprising: performing a conversion between a video comprising one or more video segments comprising one or more video units and a bitstream representation of the video, wherein the bitstream representation conforms to a format rule, and comprises information related to an adaptive resolution conversion (ARC) process, wherein the format rule specifies an applicability of the ARC process to the video segment, wherein an indication of encoding one or more video units of the video segment at different resolutions is included in the bitstream representation in a syntax structure that is different from a header syntax structure, a decoder parameter set (DPS), a video parameter set (VPS), a picture parameter set (PPS), a sequence parameter set (SPS), and an adaptation parameter set (APS).

[0674] A20. The method of solution A19, wherein the information related to ARC processing comprises a height (H) or a width (W) of a picture, the picture comprising one or more video units.

[0675] A21. The method of solution A19 or A20, wherein the information related to ARC processing is signaled in a supplemental enhancement information (SEI) message.

[0676] A22. The method of solution A19 or A20, wherein the header syntax structure comprises a picture header, a slice header, or a tile group header.

[0677] A23. The method of solution A19 or A20, wherein the information related to ARC processing is signaled in a resolution parameter set (RPS) or a conversion parameter set (CPS).

[0678] A24. The method of solution A19, wherein the information related to ARC processing comprises a ratio of a height to a width of a picture, the picture comprising one or more video units.

[0679] A25. The method of solution A19, wherein the information related to ARC processing comprises a plurality of ratios of different heights to different widths of a picture, the picture comprising one or more video units.

[0680] A26. The method of solution A25, wherein an index corresponding to an allowed ratio of the plurality of ratios is signaled in a picture parameter set (PPS).

[0681] A27. The method of solution A25, wherein any one ratio of the plurality of ratios is different from any other ratio of the plurality of ratios.

[0682] A28. The method of solution A19, wherein the information comprises at least one of (i) a difference between a first width and a second width, (ii) a difference between a first height and a second height, (iii) a ratio between the first width and the second width, or (iv) a ratio between the first height and the second height.

[0683] A29. The method of solution A28, wherein the information is encoded using a unary code, a truncated unary code, or a fixed-length code.

[0684] A30. The method of solution A19, wherein the bitstream representation further comprises at least one of: a syntax element indicative of a coding tree unit (CTU) size, a syntax element indicative of a minimum coding unit (CU) size, a syntax element indicative of a maximum or minimum transform block (TB) size, a syntax element indicative of a maximum depth of a partitioning process applicable to one or more video units, or a syntax element configured to be bound to a particular picture resolution.

[0685] A31. The method of solution A19, wherein a first reference picture associated with a current picture comprising one or more video units has a first resolution equal to a resolution of the current picture, wherein a second reference picture associated with the current picture has a second resolution greater than the resolution of the current picture, and wherein a reference index of the second reference picture is greater than a reference index of the first reference picture.

[0686] A32. The method of any of solutions A19 to A31, wherein the conversion generates the one or more video units from the bitstream representation.

[0687] A33. The method of any of solutions A19 to A31, wherein the conversion generates the bitstream representation from the one or more video units.

[0688] A34. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of any of solutions A19 to A33.

[0689] A35. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for executing the method of any of solutions A19 to A33.

[0690] In some embodiments, the following technical solutions can be implemented:

[0691] B1. A method for video processing, comprising: determining (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is the same as a resolution of a current picture comprising the current video block, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is different from the resolution of the current picture; and as a result of the determination, performing a conversion between the current video block and a bitstream representation of the video by disabling motion information of the second temporal neighboring block in a prediction of the first temporal neighboring block.

[0692] B2. A method for video processing, comprising determining that (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is different from a resolution of a current picture that includes the current video block, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is the same as the resolution of the current picture; and in response to the determination, performing a conversion between a picture block of the video and a bitstream representation of the video by disabling motion information of the second temporal neighboring block in prediction of the first temporal neighboring block.

[0693] B3. A method for video processing, comprising determining, for a current video block of a video, that a resolution of a reference picture that includes a video block associated with the current video block is different from a resolution of a current picture that includes the current video block; and in response to the determination, performing a conversion between the current video block and a bitstream representation of the video by disabling a prediction process based on the video block in the reference picture.

[0694] B4. A method for video processing, comprising making a decision, based on at least one dimension of a picture, as to whether the picture is allowed to be used as a collocated reference picture of a current video block of the picture; and performing a conversion between the current video block of the video and a bitstream representation of the video based on the decision.

[0695] B5. The method according to solution B4, wherein the at least one dimension of the reference picture is different from a corresponding dimension of the current picture that includes the current video block, and wherein the reference picture is not designated as a collocated reference picture.

[0696] B6. A method for video processing, comprising identifying, for a prediction of a current video block of a video, a collocated block based on a determination that a dimension of a collocated reference picture that includes the collocated block is the same as a dimension of a current picture that includes the current video block; and performing a conversion between the current video block and a bitstream representation of the video using the collocated block.

[0697] B7. The method according to solution B6, wherein the prediction comprises temporal motion vector prediction (TMVP) processing or alternative temporal motion vector prediction (ATMVP) processing.

[0698] B8. The method according to solution B7, wherein the dimension of the current picture is W0xH0, wherein the dimension of the collocated reference picture is W1xH1, wherein a position or a size of the collocated block is based on at least one of W0, H0, W1, or H1, and wherein W0, H0, W1, and H1 are positive integers.

[0699] B9. The method according to solution B8, wherein a derivation of a temporal motion vector in the ATMVP processing is based on at least one of W0, H0, W1, or H1.

[0700] B10. The method of solution B8, wherein the derivation of the motion vector prediction for the current video block is based on at least one of WO, HO, Wl, or HI.

[0701] B11. A method for video processing, comprising: for a current video block of a video, determining that a reference picture associated with the current video block has a resolution different from a resolution of a current picture that includes the current video block; and as part of a conversion between the current video block and a bitstream representation of the video, performing an upsampling operation or a downsampling operation on one or more reference samples of the reference picture and motion information of the current video block or coding information of the current video block.

[0702] B12. The method of solution B11, further comprising using information related to the upsampling operation or the downsampling operation in a frame different from a current frame that includes the current video block to code a subsequent video block.

[0703] B13. A method for video processing, comprising: for a conversion between a current video block of a video and a bitstream representation of the video, determining that a height or a width of a current picture that includes the current video block is different from a height or a width of a collocated reference picture associated with the current video block; and based on the determination, performing an upsampling operation or a downsampling operation on a buffer that stores one or more motion vectors of the collocated reference picture.

[0704] B14. A method for video processing, comprising: based on dimensions of a current picture that includes a current video block of a video and dimensions of a collocated picture associated with the current video block, deriving information of an optional temporal motion vector prediction (ATMVP) process applied to the current video block; and using temporal motion vectors to perform the conversion between the current video block and a bitstream representation of the video.

[0705] B15. The method of solution B14, wherein the information comprises temporal motion vectors.

[0706] B16. The method of solution B14, wherein the information comprises a motion vector prediction (MVP) for the current video block, and wherein the deriving is further based on dimensions of a reference picture referenced by the MVP.

[0707] B17. A method for video processing, comprising: configuring a bitstream representation of a video for application of an adaptive resolution conversion (ARC) process to a current video block of the video, wherein information related to the ARC process is signaled in the bitstream representation, wherein a current picture comprising the current video block has a first resolution, and wherein the ARC process comprises resampling a portion of the current video block at a second resolution different from the first resolution; and based on the configuration, performing a conversion between the current video block and the bitstream representation of the current video block.

[0708] B18. The method according to solution B17, wherein the information related to the ARC process comprises parameters for one or more upsampling or downsampling filtering methods.

[0709] B19. The method according to solution B17, wherein the information related to the ARC process comprises a horizontal scaling factor or a vertical scaling factor for scaling a reference picture to enable a resolution change within a coded video sequence.

[0710] B20. The method according to solution B18 or B19, wherein the information is signaled in a decoder parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, or a single video unit.

[0711] B21. The method according to solution B20, wherein the single video unit is a resolution parameter set (RPS) or a conversion parameter set (CPS).

[0712] B22. The method according to solution B19, wherein deriving the horizontal scaling factor or the vertical scaling factor comprises a division operation implemented using one or more tables.

[0713] B23. The method according to any of solutions B1 to B22, wherein the conversion generates the current video block from the bitstream representation.

[0714] B24. The method according to any of solutions B1 to B22, wherein the conversion generates the bitstream representation from the current video block.

[0715] B25. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement a method according to any of solutions B1 to B24.

[0716] B26. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing a method according to any of solutions B1 to B24.

[0717] In some embodiments, the following technical solutions can be implemented:

[0718] C1. A method for video processing, comprising: configuring a bitstream representation of a current video block for an application of an adaptive resolution conversion (ARC) process to the current video block, wherein information related to the ARC process is signaled in the bitstream representation, wherein the current video block has a first resolution, and wherein the ARC process comprises resampling a portion of the current video block at a second resolution different from the first resolution; and performing a conversion between the current video block and the bitstream representation of the current video block based on the configuration.

[0719] C2. The method according to solution CI, wherein the information related to the ARC process comprises a height (H) or a width (W) of a picture containing the current video block.

[0720] C3. The method according to solution CI or C2, wherein the information related to the ARC process is signaled in a supplemental enhancement information (SEI) message different from a decoder parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, and a tile group header.

[0721] C4. The method according to solution CI or C2, wherein the information related to the ARC process is signaled in a single video unit.

[0722] C5. The method according to solution C4, wherein the single video unit is a resolution parameter set (RPS) or a conversion parameter set (CPS).

[0723] C6. The method according to any of solutions CI to C5, wherein the information related to the ARC process is encoded with a fixed length code or a unary code.

[0724] C7. The method according to any of solutions CI to C5, wherein the information related to the ARC process is encoded with an exponential Golomb code of order K, wherein K is an integer greater than zero.

[0725] C8. The method according to solution CI or C2, wherein the information related to the ARC process is signaled in a decoder parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, or a tile group header.

[0726] C9. The method according to solution CI, wherein the information related to the ARC process comprises a ratio of a height and a width of a picture containing the current video block.

[0727] C10. The method of solution CI, wherein the information related to ARC processing comprises a plurality of ratios of different heights and different widths of the picture comprising the current video block.

[0728] C11. The method of solution C10, wherein an index corresponding to an allowed ratio of the plurality of ratios is signaled in a picture parameter set (PPS).

[0729] C12. The method of solution C10, wherein any one of the plurality of ratios is different from any other of the plurality of ratios.

[0730] C13. The method of solution C2, wherein TW min ≤ W ≤ TW max , and wherein TW min and TW max are positive integers.

[0731] C14. The method of solution C13, wherein TW min and TW max are signaled in a bitstream representation of the current video block.

[0732] C15. The method of solution C2, wherein TH min ≤ H ≤ TH max , and TH mi and TH max are positive integers.

[0733] C16. The method of solution C13, wherein TH mi and TH max are signaled in a bitstream representation of the current video block.

[0734] C17. The method of solution CI, wherein the picture comprising the current video block has a height H = h x Y and a width W = w x X, wherein w, h, W, H, X, and Y are positive integers, wherein X and Y are pre-defined integers, and wherein the information related to ARC processing comprises w and h.

[0735] C18. The method of solution C17, wherein X = Y = 4.

[0736] C19. The method of solution CI, wherein the information related to ARC processing comprises at least one of (i) a difference between the first width and the second width, (ii) a difference between the first height and the second height, (iii) a ratio between the first width and the second width, or (iv) a ratio between the first height and the second height.

[0737] C20. The method of solution C19, wherein the information is encoded with a unary code, a truncated unary code, or a fixed-length code.

[0738] C21. The method of solution C1, wherein the bitstream representation further comprises at least one of a syntax element indicating a coding tree unit (CTU) size, a syntax element indicating a minimum coding unit (CU) size, a syntax element indicating a maximum or minimum transform block (TB) size, a syntax element indicating a maximum depth of a partitioning process applicable to the current video block, or a syntax element configured to be tied to a particular picture resolution.

[0739] C22. A method for video processing, comprising: making a decision, for a prediction of a current video block, on whether to use a reference picture of a temporal neighboring block of the current video block; and performing a conversion between the current video block and a bitstream representation of the current video block based on the decision and a reference picture of the current video block.

[0740] C23. The method of solution C22, wherein the reference picture of the current video block has a resolution that is the same as a resolution of the current video block, wherein the reference picture of the temporal neighboring block has a resolution that is different from the resolution of the current video block, and wherein the prediction of the current video block does not use motion information associated with the temporal neighboring block.

[0741] C24. The method of solution C22, wherein the reference picture of the current video block has a resolution that is different from a resolution of the current video block, wherein the reference picture of the temporal neighboring block has a resolution that is different from the resolution of the current video block, and wherein the prediction of the current video block does not use motion information associated with the temporal neighboring block.

[0742] C25. A method for video processing, comprising: making a decision, based on at least one dimension of a reference picture of a current video block, on designating the reference picture as a collocated reference picture; and performing a conversion between the current video block and a bitstream representation of the current video block based on the decision.

[0743] C26. The method of solution C25, wherein the at least one dimension of the reference picture is different from a corresponding dimension of a current picture that includes the current video block, and wherein the reference picture is not designated as the collocated reference picture.

[0744] C27. A method for video processing, comprising: identifying, for a prediction of a current video block, a collocated block based on a comparison of a dimension of a collocated reference picture associated with the collocated block and a dimension of a current picture that includes the current video block; and performing the prediction for the current video block based on the identification.

[0745] C28. The method according to solution C27, wherein the prediction comprises a temporal motion vector prediction process or an alternative temporal motion vector prediction (ATMVP) process.

[0746] C29. The method according to solution C27 or C28, wherein a dimension of the current picture is W0xH0, wherein a dimension of the collocated reference picture is W1xH1, and wherein a position or a size of the collocated block is based on at least one of W0, H0, W1, or H1.

[0747] C30. The method according to solution C29, wherein a derivation of a temporal motion vector in the ATMVP process is based on at least one of W0, H0, W1, or H1.

[0748] C31. The method according to solution C29, wherein a derivation of a motion vector prediction for the current video block is based on at least one of W0, H0, W1, or H1.

[0749] C32. The method according to solution C1, wherein the information related to the ARC process comprises parameters for one or more up-sampling or down-sampling filtering methods.

[0750] C33. The method according to solution C1, wherein the information related to the ARC process comprises a horizontal scaling factor or a vertical scaling factor for scaling a reference picture to achieve a resolution change within a coded video sequence (CVS).

[0751] C34. The method according to solution C32 or C33, wherein the information is signaled in a decoder parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, or a single video unit.

[0752] C35. The method according to solution C34, wherein the single video unit is a resolution parameter set (RPS) or a conversion parameter set (CPS).

[0753] C36. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement a method according to any of solutions C1 to C35.

[0754] C37. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for executing a method according to any of solutions C1 to C35.

[0755] From the foregoing, it will be appreciated that specific embodiments of the presently disclosed technology have been described herein for purposes of illustration, but well-known modifications can be made by persons skilled in the art. Accordingly, the presently disclosed technology is not limited except as by the appended claims.

[0756] Implementations of the subject matter and the functional operations described in this patent document can be implemented in various systems, digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. The subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The terms "data processing apparatus" and "data processing device" encompass all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a combination of one or more of them, or a combination of one or more of them.

[0757] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.

[0758] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0759] For example, a processor suitable for the execution of a computer program includes, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0760] It is intended that the description and examples contained herein serve only as exemplifications of the application. FIG. 1 The foregoing description, for purposes of clarity, only describes exemplary embodiments of the application and the application is not limited to the embodiments characteristics, which are described in particular for a particular embodiment. Some features described in the context of separate embodiments can also be implemented within a single embodiment. Conversely, various features described within conjunction with a single embodiment can also be implemented in multiple embodiments separately or in any suitable

[0761] While this patent document contains many specifics, these should not be construed as limitations on the scope of any invention or on the scope of patent protection possible. It is expressly intended that all such specifics in the detailed description section and in the attached claims be only examples of the many aspects of the application. Some features can be separately claimed or they can form the subject of some other claims.

[0762] Also, although operations can be described as being performed in a particular order in the figures, this should not be understood as requiring that the operations be performed in that particular order, or in sequential order, or that all operations be performed. Further, the separation of various system components in the embodiments described by this patent document should not be understood as requiring such separation in all embodiments.

[0763] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for video processing, comprising: determining that (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is the same as a resolution of a current picture that includes the current video block, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is different from the resolution of the current picture; and in response to the determination, performing a conversion between the current video block and a bitstream of the video by disabling motion information of the second temporal neighboring block in a prediction of the first temporal neighboring block.

2. A method for video processing, comprising: determining that (a) a resolution of a first reference picture of a first temporal neighboring block of a current video block of a video is different from a resolution of a current picture that includes the current video block, and (b) a resolution of a second reference picture of a second temporal neighboring block of the current video block is the same as the resolution of the current picture; and in response to the determination, performing a conversion between the current video block and a bitstream of the video by disabling motion information of the second temporal neighboring block in a prediction of the first temporal neighboring block.

3. The method of claims 1 or 2, further comprising: determining, for the current video block of a video, that a resolution of a reference picture that includes a video block associated with the current video block is different from a resolution of a current picture that includes the current video block; and in response to the determination, performing a conversion between the current video block and a bitstream of the video by disabling a prediction process that is based on the video block in the reference picture.

4. The method of claims 1 or 2, further comprising: making a decision, based on at least one dimension of a reference picture of a current video block, as to whether the reference picture is allowed to be used as a collocated reference picture of the current video block of a current picture; and performing a conversion between the current video block of a video and a bitstream of the video based on the decision.

5. The method of claim 4, wherein, the at least one dimension of the reference picture is different from a corresponding dimension of a current picture that includes the current video block, and wherein the reference picture is not designated as the collocated reference picture.

6. The method of claims 1 or 2, further comprising: identifying, for a prediction of a current video block of a video, a collocated block based on a determination that a dimension of a collocated reference picture that includes the collocated block is the same as a dimension of a current picture that includes the current video block; and performing a conversion between the current video block and a bitstream of the video using the collocated block.

7. The method of claim 6, wherein, the prediction includes temporal motion vector prediction (TMVP) processing or alternative temporal motion vector prediction (ATMVP) processing.

8. The method of claim 7, wherein, the dimension of the current picture is W0xH0, wherein the dimension of the collocated reference picture is W1xH1, wherein a position or a size of the collocated block is based on at least one of W0, H0, W1, or H1, and wherein W0, H0, W1, and H1 are positive integers.

9. The method of claim 8, wherein, derivation of a temporal motion vector in the ATMVP processing is based on at least one of W0, H0, W1, or H1.

10. The method of claim 8, wherein, Derivation of motion vector prediction for the current video block is based on at least one of W0, H0, W1, or H1.

11. The method of claim 1 or 2, further comprising: for the current video block of a video, determining that a reference picture associated with the current video block has a resolution different from a resolution of a current picture that includes the current video block; and as part of a conversion between the current video block and a bitstream of the video, performing an upsampling operation or a downsampling operation on one or more reference samples of the reference picture, and motion information of the current video block or coding information of the current video block.

12. The method of claim 11, further comprising: using information related to the upsampling operation or the downsampling operation in a frame different from a current frame that includes the current video block to code a subsequent video block.

13. The method of claim 1 or 2, further comprising: for a conversion between the current video block of a video and a bitstream of the video, determining that a height or a width of a current picture that includes the current video block is different from a height or a width of a collocated reference picture associated with the current video block; and based on the determination, performing an upsampling operation or a downsampling operation on a buffer that stores one or more motion vectors of the collocated reference picture.

14. The method of claim 1 or 2, further comprising: based on dimensions of a current picture that includes a current video block of a video and dimensions of a collocated picture associated with the current video block, deriving information for an adaptive temporal motion vector prediction (ATMVP) process applied to the current video block; and using a temporal motion vector to perform the conversion between the current video block and a bitstream of the video.

15. The method according to claim 14, wherein the information includes the temporal motion vector.

16. The method of claim 14, wherein, the information includes a motion vector predictor (MVP) for the current video block, and wherein the derivation is further based on dimensions of a reference picture referenced by the MVP.

17. The method of claim 1 or 2, further comprising: for application of an adaptive resolution conversion (ARC) process to a current video block of a video, configuring a bitstream of the video, wherein information related to the ARC process is signaled in the bitstream, wherein a current picture that includes the current video block has a first resolution, and wherein the ARC process includes resampling a portion of the current video block at a second resolution different from the first resolution; and based on the configuration, performing the conversion between the current video block and the bitstream of the current video block.

18. The method of claim 17, wherein, the information related to the ARC process includes parameters for one or more upsampling or downsampling filtering methods.

19. The method of claim 17, wherein, the information related to the ARC process includes a horizontal scaling factor or a vertical scaling factor for scaling a reference picture to achieve a resolution change within a coded video sequence.

20. The method of claim 18, wherein, The information is signaled in a decoder parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a tile group header, or a single video unit.

21. The method of claim 20, wherein, The single video unit is a resolution parameter set (RPS) or a conversion parameter set (CPS).

22. The method of claim 19, wherein, Deriving the horizontal scaling factor or the vertical scaling factor includes a division operation implemented using one or more tables.

23. The method of claim 1 or 2, wherein, The conversion generates the current video block from the bitstream.

24. The method of claim 1 or 2, wherein, The conversion generates the bitstream from the current video block.

25. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to implement the method of any of claims 1-24.

26. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method of any of claims 1-24.