High-Level Syntax for Video Coding Tools

ARC enables adaptive resolution changes in video coding standards, addressing inefficiencies by resampling reference pictures for seamless switching, thus improving user experience and bandwidth efficiency.

JP7807501B2Active Publication Date: 2026-01-27DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024158878
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-10-12
Filing Date
2024-09-13
Publication Date
2026-01-27
Estimated Expiration
2040-10-12

AI Technical Summary

Technical Problem

Existing video coding standards lack the ability to adaptively change resolution without introducing IDR or IRAP pictures, leading to inefficiencies in bandwidth usage, latency, and poor user experience, especially in applications like video conferencing and streaming.

Method used

Adaptive Resolution Change (ARC) technology allows for seamless resolution adjustments by resampling reference pictures, enabling efficient switching between representations with different spatial resolutions using open GOP structures and reference picture resampling.

Benefits of technology

ARC reduces latency and improves user experience by allowing fast and seamless resolution changes, enhancing bandwidth efficiency and reducing decoding complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007807501000052
    Figure 0007807501000052
  • Figure 0007807501000053
    Figure 0007807501000053
  • Figure 0007807501000054
    Figure 0007807501000054
Patent Text Reader

Abstract

To provide an improved method for processing video data.SOLUTION: A method includes determining whether a coding tool is disabled for a picture on the basis of a first syntax element for converting between a video domain of a picture of a video and a bitstream of the video, and performing the conversion on the basis of the determination, and the first syntax element indicating whether the coding tool is disabled for the picture is signaled in a picture header.SELECTED DRAWING: Figure 23
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a divisional application of International Patent Application No. PCT / CN2020 / 120287 filed on October 12, 2020, which claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 110905 filed on October 12, 2019. All of the foregoing applications are incorporated by reference in their entirety into this patent application.

[0002] This patent document is a video coding It relates to techniques, devices and systems. [Background technology]

[0003] Despite advances in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth for digital video usage is expected to continue to increase. Summary of the Invention [Means for solving the problem]

[0004] Digital Video coding , especially video coding The present invention describes an apparatus, system, and method relating to adaptive loop filtering for an existing video signal. coding Standards (e.g., High Efficiency Video coding (HEVC) and future video coding Standards (e.g., general-purpose video coding (VVC)) or codec.

[0005] video codingThe standard has evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding The standard is for time prediction and transformation coding Hybrid video coding Based on the structure. Future video beyond HEVC coding To explore the technology, VCEG and MPEG jointly established the Joint Video Engineering Testing (JVET) in 2015. Since then, many new methods have been adopted by JVET and included in the reference software named JEM (Joint Exploration Model). In April 2018, JVET (Joint Video Expert Team) was launched between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, which aims to achieve a 50% bitrate reduction compared to HEVC.

[0006] In one exemplary aspect, the disclosed technology is used to provide a method for video processing, the method comprising: code For conversion between the quantized representation, code used to represent video regions in a representation coding determining a validity state of a tool and performing the transformation according to the determination; coding A first flag is included in the picture header to indicate the enabled state of the tool.

[0007] In yet another exemplary aspect, the above-described methods are embodied in the form of processor-executable code and stored on a computer-readable program medium.

[0008] In yet another exemplary aspect, an apparatus configured to perform or carry out the above-described method is disclosed. The apparatus may include a processor programmed to implement the method.

[0009] In yet another representative aspect, a video decoder device may implement the methods described herein.

[0010] These and other aspects and features of the disclosed technology are described in more detail in the drawings, specification and claims. [Brief explanation of the drawings]

[0011] [Figure 1] Figure 1 shows an example of adaptive streaming of two representations of the same content coded at different resolutions. [Figure 2] FIG. 2 shows an example of adaptive streaming of two representations of the same content coded at different resolutions. [Figure 3] FIG. 3 shows an example of an open GOP prediction structure for the two representations. [Figure 4] FIG. 4 shows an example of representation switching at an open GOP position. [Figure 5] FIG. 5 shows an example of the decoding process for a RASL picture by using a resampled reference picture from another bitstream as reference. [Figure 6] 6A-6C show an example of MCTS-based RWMR viewport-dependent 360° streaming. [Figure 7] FIG. 7 shows examples of collocated subpicture representations for different IRAP intervals and different sizes. [Figure 8] FIG. 8 shows an example of segments received when a change in viewing orientation results in a change in resolution. [Figure 9] FIG. 9 shows an example of a change in viewing orientation slightly upward and towards the right cubic plane compared to FIG. [Figure 10] FIG. 10 shows an example implementation in which sub-picture representations for two sub-picture positions are shown. [Figure 11] Figure 11 shows an implementation of the ARC encoder. [Figure 12] Figure 12 shows the implementation of the ARC decoder. [Figure 13] FIG. 13 shows an example of tile-based resampling for ARC. [Figure 14] FIG. 14 shows an example of adaptive resolution change. [Figure 15] FIG. 15 shows an example of ATMVP motion prediction for a CU. [Figure 16] 16A and 16B show examples of a simplified four-parameter affine motion model and a simplified six-parameter affine motion model, respectively. [Figure 17] FIG. 17 shows an example of affine MVF per sub-block. [Figure 18] 18A and 18B show examples of a four-parameter affine model and a six-parameter affine model, respectively. [Figure 19] FIG. 19 shows the MVP (Motion Vector Difference) versus AF_INTER for the inherited affine candidates. [Figure 20] FIG. 20 shows the MVP for AF_INTER for the constructed affine candidates. [Figure 21] 1A and 21B show the derivation of the five neighboring blocks and the CPMV predictor, respectively. [Figure 22] FIG. 22 shows examples of candidate positions for the affine merge mode. [Figure 23] FIG. 23 shows a flowchart of an example method for video processing in accordance with the disclosed technology. [Figure 24A] FIG. 24A is a block diagram of an example hardware platform for implementing the visual media decoding or encoding techniques described herein. [Figure 24B] FIG. 24B is a block diagram of an example hardware platform for implementing the visual media decoding or encoding techniques described herein. [Figure 25] FIG. 25 is a block diagram illustrating an example video coding system. [Figure 26] FIG. 26 is a block diagram illustrating an encoder according to some embodiments of the disclosed technology. [Figure 27] FIG. 27 is a block diagram illustrating a decoder according to some embodiments of the disclosed technology. [Figure 28] FIG. 28 shows a flowchart of an exemplary method of video processing according to some embodiments of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION

[0012] The techniques and apparatus disclosed in this application provide a coding tool with adaptive resolution change. AVC and HEVC do not have the ability to change resolution without the need to introduce IDR or Intra Random Access Point (IRAP) pictures. Such a capability can be called Adaptive Resolution Change (ARC). Use cases or application scenarios that benefit from the ARC feature include:

[0013] Rate Adaptation in Video Telephony and Conferencing: To adapt coded video to changing network conditions, when network conditions worsen and available bandwidth decreases, the encoder can adapt by encoding pictures with smaller resolution. Currently, picture resolution changes can only be made after IRAP pictures, but this presents several problems. A reasonable-quality IRAP picture is much larger than an inter-coded picture and is correspondingly more complex to decode. This takes time and resources. This becomes problematic when a resolution change is required by the decoder for load reasons. It also breaks low-latency buffering requirements, forcing audio resynchronization and increasing the end-to-end delay of the stream, at least temporarily. This can result in a poor user experience.

[0014] Active speaker change in multi-party video conference: In a multi-party video conference, it is common for the active speaker to display a larger video size than the videos of the remaining conference participants. When the active speaker changes, the picture resolution of each participant may also need to be adjusted. The need for the ARC function becomes more important when such changes in active speakers occur frequently.

[0015] Fast Start in Streaming: For streaming applications, it is common to buffer up to a certain length of decoded pictures before the application starts displaying. Starting the bitstream with a smaller resolution allows the application to have enough pictures in the buffer to start displaying sooner.

[0016] Adaptive stream switching in streaming: The DASH (Dynamic Adaptive Streaming over HTTP) standard includes a feature named @mediaStreamStructureId, which enables switching between different representations at open GOP random access points with non-decodable leading pictures, such as a CRA picture with an associated RASL picture in HEVC. If two different representations of the same video have different bitrates but the same spatial resolution, and they have the same @mediaStreamStructureId value, switching between the two representations with associated RASL and CRA pictures can be performed, and switching at the CRA picture allows the associated RASL picture to be decoded with acceptable quality, allowing seamless switching. In ARC, the @mediaStreamStructureId feature can also be used to switch between DASH representations with different spatial resolutions.

[0017] ARC is also known as dynamic resolution conversion.

[0018] ARC can also be seen as a special case of reference picture resampling (RPR) such as in H.263 Annex P.

[0019] 1.1 Reference Picture Resampling in H.263 Annex P This mode describes an algorithm for warping a reference picture before its use for prediction. This can be useful for resampling a reference picture that has a different source format than the picture being predicted. It can also be used for global motion estimation or rotational motion estimation by warping the shape, size, and position of the reference picture. The syntax includes the warping parameters and resampling algorithm to be used. The simplest level of operation for the reference picture resampling mode is implicit factor of 4 resampling, since it only requires applying an FIR filter to the upsampling and downsampling process. In this case, no additional signaling overhead is required, since its use is understood when the size of the new picture (indicated in the picture header) is different from the size of the previous picture.

[0020] 1.2 Contribution of ARC to VVC 1.2.1 JVET-M0135 The preliminary design of the ARC described below, taken in part from JCTVC-F158, is suggested to be a placeholder intended only to spark discussion.

[0021] 2.2.1.1 Description of basic tools The basic tool constraints for supporting ARC are: The spatial resolution may differ from the nominal resolution by a factor of 0.5 applied to both dimensions. The spatial resolution may be increased or decreased, resulting in magnifications of 0.5 and 2.0. -The aspect ratio and chroma format of the video format will not be changed. - The cropping area scales proportionally to the spatial resolution. -Reference pictures are simply rescaled if necessary, and inter prediction is applied as usual.

[0022] 2.2.1.2 Scaling Behavior It is proposed to use simple zero-phase separable down- and up-scaling filters, although these filters are only for prediction, and the decoder may use more advanced scaling for the output.

[0023] The following 1:2 downscaling filter is used, with zero phase and 5 taps: (-1, 9, 16, 9, -1) / 32

[0024] The downsampling points are at even sample positions and are co-sited. The same filter is used for luma and chroma.

[0025] For 2:1 upsampling, additional samples at odd grid positions are generated using half-pel motion compensation interpolation filter coefficients in the latest VVC WD.

[0026] The combined up- and down-sampling does not change the position or phase of the chroma sampling points.

[0027] 2.2.1.2 Resolution Description in Parameter Sets The signaling of picture resolution in the SPS is changed as follows, where deletions are marked with double brackets (e.g., [[a]] means deletion of the letter "a").

[0028] [Table 1]

[0029] [[pic_width_in_luma_samples is Decoded Specifies the picture width in luma samples. pic_width_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY. pic_height_in_luma_samples is the height of each DecodedSpecifies the picture height in luma samples. pic_height_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0030] num_pic_size_in_luma_samples_minus1 plus 1 is the number of picture sizes (width and height), code The pixel size is specified in units of luminance samples that may be present in the encoded video sequence.

[0031] pic_width_in_luma_samples[i] is Decoded The width of the ith picture is code pic_width_in_luma_samples[i] is specified in units of luminance samples that may be present in the encoded video sequence. pic_width_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0032] pic_height_in_luma_samples[i] is Decoded The height of the ith picture is code pic_height_in_luma_samples[i] is specified in units of luminance samples that may be present in the encoded video sequence. pic_height_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0033] [Table 2]

[0034] pic_size_idx specifies the index to the ith picture size in the sequence parameter set. The width of a picture referencing a picture parameter set is pic_width_in_luma_samples[pic_size_idx] in luma samples. Similarly, the height of a picture referencing a picture parameter set is pic_height_in_luma_samples[pic_size_idx] in luma samples.

[0035] 1.2.2 JVET-M0259 1.2.2.1 Background: Subpicture The term subpicture track is defined in OMAF (Omnidirectional Media Format) as follows: A track has a spatial relationship with other tracks and represents a spatial subset of the original video content, and is used by content creators to encoding The image is divided into spatial subsets before the subpicture track is generated. A subpicture track for HEVC can be constructed by rewriting the parameter set and slice segment headers for the motion constrained tile set so that it becomes a self-contained HEVC bitstream. A subpicture representation can be defined as a DASH representation that preserves the subpicture track.

[0036] JVET-M0261, which uses the term subpicture as the spatial division unit for VVC, is summarized as follows: 1. A photo is divided into subpictures, tile groups and tiles. 2. A subpicture is a set of rectangles in tile groups starting with the tile group with tile_group_address=0. 3. Each sub-picture may refer to its own PPS and therefore may have its own tile division. 4. Subpictures Decoding It is treated like a picture in the process. 5. Subpicture Decoding The reference picture for Decoded It is generated by extracting a region from a reference picture in the picture buffer that is co-located with the current subpicture. Decoded sub-pictures, i.e., inter-prediction is performed between sub-pictures of the same size and same position within a picture. 6. A tile group is a set of tiles in a tile raster scan of a subpicture.

[0037] In this contribution, the term subpicture can be understood as defined in JVETM0261. However, tracks encapsulating subpicture sequences as defined in JVETM0261 have very similar properties to subpicture tracks as defined in OMAF. The following examples apply to both cases.

[0038] 1.2.2.2 Use case 1.2.2.2.1 Adaptive Resolution Change in Streaming Requirements for Adaptive Streaming Support Section 5.13 of MPEG N17074 ("Requirements for Support of Adaptive Streaming") contains the following requirement for VVC: For adaptive streaming services that offer multiple representations of the same content, each with different characteristics (e.g., spatial resolution or sample bit depth), the standard shall support fast representation switching. The standard shall enable the use of efficient prediction structures (e.g., so-called open picture groups) without compromising the ability to fast and seamlessly switch representations between representations with different characteristics, such as different spatial resolutions.

[0039] Example of an open GOP prediction structure with representation switching Content generation for adaptive bitrate streaming involves the generation of different representations that can have different spatial resolutions. The client requests segments from the representations and can therefore decide at what resolution and bitrate it receives the content. At the client, the segments of the different representations are concatenated, Decoding The client must be able to achieve seamless playback with one decoder instance. A closed GOP structure (starting with an IDR picture) is traditionally used, as shown in Figure 1. Figure 1 shows the different resolutions code 1 illustrates adaptive streaming of two representations of the same content.

[0040] The open GOP prediction structure (starting with a CRA picture) allows for better compression performance than the respective closed GOP prediction structure. For example, an average bitrate reduction of 5.6% was obtained for IRAP pictures with an interval of 24 pictures in terms of the luma Bjontegaard delta bitrate. For convenience, the simulation conditions and results of [2] are summarized in Section YY.

[0041] The open GOP prediction structure has been reported to reduce subjectively visible pumping in quality.

[0042] The challenge of using open GOP in streaming is that after switching representations, RASL pictures need to be remapped to the correct reference pictures. Decoding This representation issue is difficult to resolve at different resolutions. code This is shown in Figure 2, which illustrates adaptive streaming of two representations of the same encoded content, where the segments use either a closed GOP or an open GOP prediction structure.

[0043] A segment that starts with a CRA picture contains an RASL picture with at least one reference picture in the previous segment. This is illustrated in Figure 3, which shows the open GOP prediction structure of the two representations. In Figure 3, picture 0 in both bitstreams is in the previous segment and is used as a reference for predicting the RASL picture.

[0044] The representation switching marked with a dashed rectangle in Figure 2 is shown below in Figure 4, which shows representation switching at an open GOP position. The reference picture for the RASL picture ("Picture 0") is Decoding As a result, the RASL picture is Decoding and there are gaps in the video playback.

[0045] However, if we use resampled reference pictures to generate RASL images, DecodingIt has been found to be subjectively acceptable, see Section 4. Re-sampling "Picture 0" and calling it a RASL picture Decoding Fig. 5 shows the use of a resampled reference picture from another bitstream as a reference for the RASL picture. Decoding Show the process.

[0046] 2.2.2.2.2 Viewport Changes in Region-Wise Mixed Resolution (RWMR) 360° Video Streaming Background: HEVC-based RWMR Streaming RWMR360° streaming provides an improved effective viewport spatial resolution. The tiles covering the viewport are equivalent to "4K". Decoding The scheme derived from the 6K (6144 x 3072) ERP picture or equivalent CMP resolution shown in Figure 6, which has the capability (HEVC Level 5.1), is included in OMAF Sections D.6.3 and D.6.4 and is also adopted in the VR Industry Forum guidelines. Such a resolution is claimed to be suitable for head-mounted displays using quad HD (2560 x 1440) display panels.

[0047] encoding The content is presented in two spatial resolutions with cube face sizes of 1536x1536 and 768x768 respectively. Encoding In both bitstreams, a 6x4 tile grid is used, with a motion constrained tile set (MCTS) at each tile position. code It has been turned into

[0048] Encapsulation Each MCTS sequence is encapsulated as a subpicture track and made available as a DASH subpicture representation.

[0049] Selecting MCTS to be streamed12 MCTSs are selected from the high-resolution bitstream, and 12 complementary MCTSs are extracted from the low-resolution bitstream, so that a hemisphere (180° x 180°) of the streamed content comes from the high-resolution bitstream.

[0050] Merging MCTS into the bitstream to be decoded : The received MCTS for a single time instance is a 1920x4608 image conforming to HEVC Level 5.1. code Another option for the merged picture is to have four tile columns of 768 luma samples wide, two tile columns of 384 luma samples wide, and three tile columns of 768 luma samples high, resulting in a picture of 3840 x 2304 luma samples.

[0051] Figure 6 shows an example of MCTS-based RWMR viewport-dependent 360° streaming. Figure 6A shows code 6A shows an example of a merged bitstream, FIG. 6B shows an example of an MCTS selected for streaming, and FIG. 6C shows an example of pictures merged from the MCTS.

[0052] Background: Some representations of different IRAP intervals for viewport-dependent 360° streaming When the viewing orientation changes in HEVC-based viewport-dependent 360° streaming, the new selection of subpicture representation takes effect at the next IRAP-aligned segment boundary. Decoding For the subpicture representation code The selected subpicture representations are merged into the resulting picture, thus aligning the VCL NAL unit types with all selected subpicture representations.

[0053] To provide a trade-off between response time in response to changes in viewing orientation and rate-distortion performance when viewing direction is stable, multiple versions of the content are presented at different IRAP intervals. code This can be achieved by encodingA set of ordered subpicture representations for is shown in Figure 7 and discussed in detail in Section 3 of H. Chen, H. Yang, and J. Chen, "Another List for Subblock Merging Candidates" (JVET-L0368, October 2018).

[0054] FIG. 7 shows examples of ordered sub-picture representations for different IRAP intervals and different sizes.

[0055] Figure 8 shows an example where a sub-picture location is initially selected to be received at a lower resolution (384x384). A change in viewing orientation results in a new selection where the sub-picture location is received at a higher resolution (768x768). In the example of Figure 8, the segment received when the change in viewing orientation causes a change in resolution is the beginning of segment 4. In this example, segment 4 is received from a sub-picture representation with a short IRAP interval because a change in viewing orientation occurs. After that, the viewing orientation stabilizes, and from segment 5 onwards, a version with a longer IRAP interval can be used.

[0056] The drawback of updating all subpicture positions In a typical viewing situation, the viewing orientation shifts gradually, causing the resolution to change at only a subset of sub-picture locations in RWMR viewport-dependent streaming. Figure 9 shows a change in viewing orientation slightly upward and toward the right cube face from Figure 6. The cube face partition with a different resolution than the previous one is indicated by "C." It can be seen that the resolution has changed for six of the 24 cube face partitions. However, as mentioned above, segments starting with an IRAP picture need to be received for all 24 cube face partitions in response to a change in viewing orientation. Updating all sub-picture locations for segments starting with an IRAP picture is inefficient in terms of streaming rate-distortion performance.

[0057] In addition, it is desirable to be able to use an open GOP prediction structure in the sub-picture representation of RWMR 360° streaming to improve rate-distortion performance and avoid the visible picture quality pumping caused by closed GOP prediction structures.

[0058] Proposed design example We propose the following design goals: 1. The VVC design allows sub-pictures derived from random access pictures and other sub-pictures derived from non-random access pictures to be combined in the same VVC-compliant code It should be possible to merge the image into a composite picture. 2. The VVC design should allow for the merging of subpicture representations into a single VVC bitstream while allowing the use of open GOP prediction structures in the subpicture representations, without compromising the ability for fast and seamless representation switching between subpicture representations of different properties, such as different spatial resolutions.

[0059] An example of the design goal can be seen in Figure 10, where sub-picture representations for two sub-picture locations are shown. For both sub-picture locations, a separate version of the content is provided for each combination of the two resolutions and the two random access intervals. codeSome segments start with an open GOP prediction structure. A change in viewing orientation causes the resolution of subpicture position 1 to switch at the beginning of segment 4. Because segment 4 starts with a CRA picture related to a RASL picture, the reference picture for the RASL picture in segment 3 needs to be resampled. Note that this resampling applies to subpicture position 1, but the decoded subpictures in some other subpicture positions are not resampled. In this example, the change in viewing orientation does not cause a change in resolution for subpicture position 2, so the decoded subpicture in subpicture position 2 is not resampled. In the first picture of segment 4, the segment at subpicture position 1 contains a subpicture derived from a CRA picture, while the segment at subpicture position 2 contains a subpicture derived from a non-random access picture. code It has been suggested that merging these sub-pictures into a composite picture is possible in VVC.

[0060] 2.2.2.2.3 Adaptive resolution change in videoconferencing JCTVC-F158 proposes adaptive resolution scaling primarily for video conferencing. The following subsections are copied from JCTVC-F158 and present use cases where adaptive resolution scaling is claimed to be useful:

[0061] Seamless network adaptation and error tolerance Applications such as video conferencing and streaming over packet networks can be difficult, especially if the bit rate is too high and data is lost. Encoded It is often necessary for a stream to adapt to changing network conditions. Such applications usually have a return channel that allows the encoder to detect errors and make adjustments. Encoders have two main tools at their disposal: bitrate reduction, either temporal or spatial, and resolution change. Temporal resolution change can be achieved using a hierarchical prediction structure. codingHowever, for best quality, a change in spatial resolution is required on the part of a well-designed encoder for video communication.

[0062] When changing spatial resolution within AVC, an IDR frame must be sent and the stream reset. This poses a significant problem: IDR frames of reasonable quality are much larger than interpictures and correspondingly To decode This makes the process more complex. This takes time and resources. This becomes problematic if a resolution change is required by the decoder for load reasons. Also, the low latency buffer requirement may be violated, forcing audio resynchronization and increasing the end-to-end delay of the stream, at least temporarily. This results in a poor user experience.

[0063] To minimize these delays, IDRs are usually transmitted at a lower quality, using a similar number of bits to P frames, and it takes a significant amount of time to return to full quality for a given resolution. To make the delay small enough, the quality can actually be so low that there is often a visible blur before the image "refocuses." In effect, intraframes are almost useless in compressed situations; they are merely a way to restart the stream.

[0064] Therefore, there is a need for a method in HEVC that allows resolution changes with minimal impact on the subjective experience, especially in difficult network conditions.

[0065] Fast Start It would be useful to have a "fast start" mode in which the first frame is sent at reduced resolution and then the resolution is increased over the next few frames to reduce latency and reach normal quality more quickly without initially introducing unacceptable image blur.

[0066] Conference “compose” Videoconferencing often features a feature where the speaker is shown full screen and other participants are shown in a smaller resolution window. To efficiently support this, a smaller picture is often sent at a lower resolution. This resolution increases when a participant becomes a speaker and goes full screen. Sending an intraframe at this point would result in an unpleasant pause in the video stream. This effect can be very noticeable and unpleasant when there is a rapid change of speakers.

[0067] 2.2.2.3 Proposed Design Goals The following high-level design choices are proposed for VVC Version 1:

[0068] 1. It is proposed to include a reference picture resampling process in VVC Version 1 for the following use cases: - Use efficient prediction structures (for example so-called open picture groups) in adaptive streaming, without compromising the ability to perform fast and seamless representation switching between representations with different characteristics such as different spatial resolutions. Adapting low-latency interactive video content to resolution changes caused by network conditions and applications without significant delay or latency changes.

[0069] 2. A sub-picture derived from a random access picture and another sub-picture derived from a non-random access picture can be combined in the same VVC-compliant code A VVC design is proposed that allows merging of multiple images into a single image, which is asserted to enable efficient handling of viewing orientation changes in mixed-quality and mixed-resolution viewport-adaptive 360° streaming.

[0070] 3. It is proposed to include a sub-picture-wise resampling process in VVC Version 1. This is asserted to enable efficient handling of viewing orientation changes in mixed-quality and mixed-resolution viewport-adaptive 360° streaming.

[0071] 2.2.3. JVET-N0048 The use cases and design goals for Adaptive Resolution Change (ARC) are discussed in detail in JVET-M0259, summarized below:

[0072] 1. Real-time communication The following use cases for adaptive resolution change were originally included in JCTVC-F158: a. Seamless network adaptation and error resilience (through dynamic adaptive resolution changes) b. Fast Start (gradual increase in resolution when starting or resetting a session) c. Meeting "Compose" (higher resolution given to the speaker)

[0073] 2. Adaptive Streaming Section 5.13 ("Support for adaptive streaming") of MPEG N17074 includes the following requirement for VVC: For adaptive streaming services that offer multiple representations of the same content, each with different characteristics (e.g., spatial resolution or sample bit depth), the standard shall support fast representation switching. The standard shall allow the use of efficient prediction structures (e.g., so-called open picture groups) without compromising the ability to fast and seamlessly switch representations between representations with different characteristics, such as different spatial resolutions.

[0074] JVET-M0259 discusses how to meet this requirement by resampling the reference pictures of the leading picture.

[0075] 3. 360° Viewport-Dependent Streaming JVET-M0259 specifies the independent identification of the reference picture of the leading picture. code It is discussed how to address this use case by resampling the filtered picture region.

[0076] This contribution is an adaptive resolution algorithm that is determined to meet all of the above use cases and design goals. coding We propose an approach: 360° viewport-dependent streaming and conference "compose" use cases are addressed by this proposal together with JVET-N0045 (which proposes an independent subpicture layer).

[0077] Proposed specification text signaling

[0078] [Table 3]

[0079] sps_max_rpr specifies the maximum number of active reference pictures in reference picture list 0 or 1 for tile groups in a CVS whose pic_width_in_luma_samples and pic_height_in_luma_samples are not equal to the pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively.

[0080] [Table 4]

[0081] [Table 5]

[0082] max_width_in_luma_samples specifies that it is a bitstream conformance requirement that pic_width_in_luma_samples in any active PPS for any picture in a CVS where this SPS is active be less than or equal to max_width_in_luma_samples.

[0083] max_height_in_luma_samples specifies that it is a bitstream conformance requirement that pic_height_in_luma_samples in any active PPS for any picture in a CVS where this SPS is active be less than or equal to max_height_in_luma_samples.

[0084] High Level Decoding process For the current image CurrPic Decoding The process works as follows. 1. NAL Unit Decoding is specified in Section 8.2. 2. The process in Section 8.3 uses the following syntax elements on the tile group header layer and above: Decoding Define the process.

[0085] - The picture order count variables and functions are derived as specified in section 8.3.1. This only needs to be called for the first group of picture tiles.

[0086] - for each tile group of a non-IDR picture Decoding At the start of the process, the reference picture list configuration specified in Section 8.3.2 Decoding The process is called to derive reference picture list 0 (RefPicList[0]) and reference picture list 1 (RefPicList[1]).

[0087] - for reference picture marking in clause 8.3.3 Decoding A process can be invoked to mark a reference picture as "unused for reference" or "used for long-term reference", which only needs to be invoked for the first group of picture tiles.

[0088] For each active reference picture in RefPicList[0] and RefPicList[1] whose pic_width_in_luma_samples or pic_height_in_luma_samples is not equal to CurrPic's pic_width_in_luma_samples or pic_height_in_luma_samples, respectively, the following applies:

[0089] - The resampling process of the XYZ term is called [Ed.(MH): details of the calling parameters to be added] and the output has the same reference picture markings and picture order count as the input.

[0090] The reference pictures used as input to the resampling process are marked as "unused for reference".

[0091] CTU (coding tree units), for scaling, transformation, in-loop filtering, etc. Decoding The invocation of a process can be further discussed.

[0092] All tile groups in the current picture Decoding After that, the current Decoded The picture is marked as "used for short-term reference."

[0093] Resampling Process The SHVC resampling process (section H8.1.4.2 of HEVC) is proposed with the following additions:

[0094] If sps_ref_wrapaund_enabled_flag=0, the sample values ​​tempArray[n] (n=0..7) are derived as follows:

[0095]

number

[0096] Otherwise, the sample values ​​tempArray[n] (n=0..7) are derived as follows:

[0097]

number

[0098] If sps_ref_wrapaund_enabled_flag=0, the sample values ​​tempArray[n] (n=0..3) are derived as follows:

[0099]

number

[0100] Otherwise, the sample values ​​tempArray[n] (n=0..3) are derived as follows:

[0101]

number

[0102] 2.2.4. JVET-N0052 The concept of adaptive resolution change in video compression standards has been around since at least 1996, particularly in H.263+ related proposals for reference picture resampling (RPR, Annex P) and reduced resolution update (Annex Q). It has recently gained some attention, first proposed by Cisco during the JCT-VC era, then in the context of VP9 (which is now reasonably widely deployed), and more recently in the context of VVC. ARC is a standard for code This reduces the number of samples that need to be sampled, allowing the resulting reference picture to be upsampled to a higher resolution if desired.

[0103] The ARC of particular interest is considered in two scenarios.

[0104] 1) Intra frames such as IDR pictures code Intrapictures are often much larger than interpictures. Regardless of the reason, intrapictures code Downsampling pictures with the intention of reducing the quality may provide better input for future prediction, which is also clearly advantageous from a rate control point of view, at least in low latency applications.

[0105] 2) When operating a codec near its breaking point, as at least some cable and satellite operators routinely do, ARC can detect non-intra scene transitions without hard transition points. code This can even be useful for filtered pictures.

[0106] 3) Looking perhaps a bit too much forward: Is the concept of fixed resolution generally justifiable? With the move away from CRTs and the proliferation of scaling engines in rendering devices, rendering and coding Hard-binding to resolution is a thing of the past. There is also available research suggesting that when there is a lot of activity going on in a video sequence, most people are unable to focus on the details (presumably associated with high resolution), even if that activity is happening spatially elsewhere. If this is correct and generally accepted, fine-grained resolution variation could be a better rate control mechanism than adaptive QP. This point is currently under debate. Removing the notion of fixed-resolution bitstreams has myriad system layer and implementation implications that are well known (at least at the level of their existence, if not their detailed nature).

[0107] Technically, ARC can be implemented as reference picture resampling. The implementation of reference picture resampling involves two main aspects: the resampling filter and the signaling of resampling information in the bitstream. This application focuses on the latter, and touches on the former to the extent that implementation experience exists. Further research into appropriate filter designs is encouraged, and Tencent will carefully consider and, where appropriate, support any suggestions that substantially improve upon the proposed design provided.

[0108] Overview of Tencent's ARC Implementation 11 and 12 show the implementation of Tencent's ARC encoder and decoder, respectively. The implementation of the disclosed technology allows changing the width and height of a picture at a picture granularity regardless of the picture type. In the encoder, the input image data is Encoding The first input picture is downsampled to the selected picture size for Encoding After that, Decoding The displayed picture Decoded The resulting picture is then downsampled to a different sampling ratio and stored as an interpicture. Encoding If so, the reference pictures in the DPB are up / downscaled according to the spatial ratio between the reference picture size and the current picture size. Decoding The extracted picture is stored in the DPB without resampling. However, the reference picture in the DPB, if used for motion compensation, is currently Decoding The image is up / downscaled relative to the spatial ratio between the picture being edited and the reference. Decoding The resulting picture, when bumped out for display, is upsampled to the original picture size or the desired output picture size. In the motion estimation / compensation process, the motion vectors are scaled relative to the picture size ratio and the picture order count difference.

[0109] ARC Parameter Signaling In this specification, the term ARC parameters is used to refer to any combination of parameters required for ARC to function. In the simplest case, it is a zoom factor or an index into a table with a defined zoom factor. This can be a target resolution (e.g., granularity in samples or maximum CU size) or an index into a table providing the target resolution, as proposed in JVET-M0135. It can also include filter parameters (down to the filter coefficients) or a filter selector for the up / downsampling filter in use.

[0110] From the beginning, the implementation proposed here is to allow, at least conceptually, different ARC parameters for different parts of a picture. As per the current VVC draft, it is proposed that a suitable syntax structure may be a rectangular tile group (TG). Use of scan order TG may be restricted to using ARC only for full pictures, or to the extent that scan order TG is included in a rectangular TG. This can be easily specified in the bitstream constraints.

[0111] Since different TGs may have different ARC parameters, the appropriate location for the ARC parameters is within the TG header or within a parameter set with a scope of TG, which may be referenced by the TG header (adaptive parameter set in the current VVC draft) or a more detailed reference (index) to a table in a higher parameter set. Of these three options, at this point, it is best to use the TG header to reference a table entry containing the ARC parameters. code It was proposed to make the table into a table in the SPS and in the DPS (this time). code The zoom factor is set directly in the TG header without using a parameter setting. code Using PPS for reference, as proposed in JVET-M0135, is contrasted with the case where per-tile group signaling of ARC parameters is a design criterion.

[0112] For the table entries themselves, the following options are available:

[0113] -Downsampling factor for either one of both dimensions or independently for the X and Y dimensions coding This is mostly a (HW-)implementation discussion, and some may prefer results where the zoom factor in the X dimension is fairly flexible, but the Y dimension is fixed at 1 or has very few options. It is suggested that syntax is the wrong place to express such constraints, and that if they are desirable, constraints expressed as requirements for conformance are preferable, i.e., keeping the syntax flexible.

[0114] -Target resolution coding It is proposed below to: There may be more or less complex constraints on these resolutions relative to the current resolution, possibly expressed in bitstream adaptation requirements.

[0115] - Downsampling per tile group is preferred to allow picture synthesis / extraction, but is not important from a signaling point of view. If a group makes the unwise decision to allow ARC only at picture granularity, they can include a bitstream conformance requirement that all TGs use the same ARC parameters.

[0116] - Control information for ARC. In our following design, it includes the reference picture size.

[0117] -Do you need flexibility in filter design? Do you need more than a few code points? If the answer is yes, put them in APS? Some implementations suggest that if the downsample filter changes and the ALF stays, the bitstream will need to eat overhead.

[0118] For now, in order to keep the proposed technology consistent and simple (as much as possible), we suggest the following: -Fixed filter design -Target resolution in the table in the SPS for bitstream constraints TBD -Min / Max target resolution in DPS to facilitate cap exchange / negotiation The resulting syntax may look like this:

[0119] [Table 6]

[0120] max_pic_width_in_luma_samples is in luma samples in the bitstream Decoding Specifies the maximum width of the encoded picture. max_pic_width_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY. The value of dec_pic_width_in_luma_samples[i] cannot be greater than the value of max_pic_width_in_luma_samples.

[0121] max_pic_height_in_luma_samples is in luma samples in the bitstream Decoding Specifies the maximum height of the encoded picture. max_pic_height_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY. The value of dec_pic_height_in_luma_samples[i] cannot be greater than the value of max_pic_height_in_luma_samples.

[0122] [Table 7]

[0123] adaptive_pic_resolution_change_flag=0 specifies the output picture size (output_pic_width_in_luma_samples, output_pic_height_in_luma_samples), Decoded The number of picture sizes (num_dec_pic_size_in_luma_samples_minus1) and at least one Decoded Specifies that the picture size (dec_pic_width_in_luma_samples[i], dec_pic_height_in_luma_samples[i]) is present in the SPS. The reference picture size (reference_pic_width_in_luma_samples, reference_pic_height_in_luma_samples) is present and is subject to the value of reference_pic_size_present_flag.

[0124] output_pic_width_in_luma_samples specifies the width of the output picture in units of luma samples. output_pic_width_in_luma_samples shall not be equal to 0.

[0125] output_pic_height_in_luma_samples specifies the height of the output picture in units of luma samples. output_pic_height_in_luma_samples shall not be equal to 0.

[0126] reference_pic_size_present_flag=1 specifies that reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples are present.

[0127] reference_pic_width_in_luma_samples specifies the width of the reference picture in units of luma samples. output_pic_width_in_luma_samples shall not be equal to 0. If not present, the value of reference_pic_width_in_luma_samples is inferred to be equal to dec_pic_width_in_luma_samples[i].

[0128] reference_pic_height_in_luma_samples specifies the height of the reference picture in units of luma samples. output_pic_height_in_luma_samples shall not be equal to 0. If not present, the value of reference_pic_height_in_luma_samples is inferred to be equal to dec_pic_height_in_luma_samples[i].

[0129] NOTE 1 - The size of the output picture shall be equal to the values ​​of output_pic_width_in_luma_samples and output_pic_height_in_luma_samples. The size of the reference picture shall be equal to the values ​​of reference_pic_width_in_luma_samples and reference_pic_height_in_luma_samples if a reference picture is used for motion compensation.

[0130] num_dec_pic_size_in_luma_samples_minus1+1 is code in units of luma samples of the encoded video sequence Decoded Specifies the number of picture sizes (dec_pic_width_in_luma_samples[i], dec_pic_height_in_luma_samples[i]).

[0131] dec_pic_width_in_luma_samples[i] codein units of luma samples of the encoded video sequence Decoded Specifies the ith width of the picture size. dec_pic_width_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0132] dec_pic_height_in_luma_samples[i] is code in units of luma samples of the encoded video sequence Decoded Specifies the i-th height of the picture size. dec_pic_height_in_luma_samples[i] must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0133] Note 2-ith Decoded Picture size (dec_pic_width_in_luma_samples[i], dec_pic_height_in_luma_samples[i]) is code in the encoded video sequence Decoded Picture Decoded It may be equal to the picture size.

[0134] [Table 8]

[0135] dec_pic_size_idx is Decoded The picture width is equal to pic_width_in_luma_samples[dec_pic_size_idx], Decoded Specifies that the picture height is equal to pic_height _in_luma_samples[dec_pic_size_idx].

[0136] filter The proposed design conceptually consists of a downsampling filter from the original picture to the input picture, an up / downsampling filter that rescales the reference picture for motion estimation / compensation, and DecodedIt includes four different filter sets: picture-to-output-picture upsampling filters. The first and last ones can be left as a non-standard matter. Within the scope of the specification, the up / downsampling filters need to be explicitly signaled or predefined in the appropriate parameter sets.

[0137] Our implementation uses the SHVC downsampling filter (SHM ver. 12.4), a 12-tap 2D separable filter, for downsampling to resize the reference picture to be used for motion compensation. In the current implementation, only dyadic sampling is supported. Therefore, the phase of the downsampling filter is set to zero by default. For upsampling, a 16-phase 8-tap interpolation filter is used to shift the phase and align the luma and chroma pixel positions to their original positions.

[0138] Tables 9 and 10 show the 8-tap filter coefficients fL[p,x] (p=0..15 and x=0..7) used in the luma upsampling process and the 4-tap filter coefficients fC[p,x] (p=0..15 and x=0..3) used in the chroma upsampling process.

[0139] Table 11 shows the 12-tap filter coefficients for the downsampling process. The same filter coefficients are used for both luma and chroma in the downsampling.

[0140] [Table 9]

[0141] [Table 10]

[0142] [Table 11]

[0143] It is expected that (possibly significant) subjective and objective gains can be expected when using filters that adapt to content and / or scaling factors.

[0144] Tile group boundary discussion As is likely true for much of the work related to tile groups, our implementation of tile group (TG)-based ARC is not completely finished. Our preference is to revisit the implementation when the discussion of spatial organization and extraction of multiple sub-pictures into a composite picture in the compressed domain has at least provided a starting point. However, this does not prevent us from extrapolating the results to some extent and adapting our signaling design accordingly.

[0145] For now, the tile group header is the right place for things like dec_pic_size_idx, as proposed above, for the reasons already stated. A single ue(v) codepoint, dec_pic_size_idx, is conditionally present in the tile group header and is used to indicate the ARC parameters used. To match implementations that are per-picture ARC only, the spec space allows for only a single tile group. coding or given code The condition for bitstream compliance is that all TG headers (if present) of a coded picture have the same value of dec_pic_size_idx.

[0146] The parameter dec_pic_size_idx can be moved to any of the headers that start a subpicture, which can still be a tile group header.

[0147] Beyond these syntactic considerations, some additional work is required to enable tile-group or sub-picture based ARC. Perhaps the most challenging part is how to deal with the issue of unnecessary samples in pictures where the sub-pictures have been resampled to a smaller size.

[0148] Figure 13 shows an example of tile-group-based resampling for ARC. Consider the picture on the right, which is composed of four sub-pictures (possibly represented as four rectangular tile groups in the bitstream syntax). The bottom right TG on the left is subsampled to half size. We need to discuss what to do with samples outside the relevant region, denoted "half".

[0149] Many (most? all?) of my previous videos coding The standards have in common that they do not support spatial extraction of parts of a picture in the compressed domain. This means that each sample of a picture is represented by one or more syntax elements, and each syntax element affects at least one sample. To maintain this, it may be necessary to somehow add an area around the samples covered by the downsampled TG, labeled "half". Annex P of H.263+ solves this problem by padding. In fact, the sample values ​​of the padded samples may be signaled in the bitstream (within certain strict limits).

[0150] Although this may constitute a significant departure from the previous assumption, an alternative that may be necessary to support sub-bitstream extraction (and construction) based on rectangular portions of a picture is that each sample of the reconstructed picture is code This may be a relaxation of the current understanding that all blocks in a given picture need to be represented by something in the given picture (even if that something is only a skipped block).

[0151] Implementation Considerations, System Implications and Profiles / Levels We propose basic ARCs to be included in the "Baseline / Main" profile. Sub-profiling may be used to remove them if not required for a particular application scenario. Certain limitations may be tolerated. In this regard, it should be noted that certain H.263+ profiles and "Preferred Modes" (previously profiles) included Annex P, which was used only as an "implicit factor of 4", i.e., binomial downsampling in both dimensions. This was sufficient to support fast start for videoconferencing (quickly getting past I-frames).

[0152] This design allows all filtering to be done "on the fly" with no or negligible memory bandwidth increase, so there seems to be no need to move ARC to exotic profiles.

[0153] Complex tables, etc., may not be meaningfully used for capability exchange, as discussed in Marrakech with JVET-M0135. The number of options simply grows to allow meaningful cross-vendor interop, assuming offer-answer and similar handshakes of limited depth. Realistically, to meaningfully support ARC in capability exchange scenarios, most interop points must fall back to just a few: for example, no ARC, ARC with an implicit factor of 4, or full ARC. Alternatively, it is possible to specify required support for all ARCs and leave bitstream complexity limitations to higher-level SDOs. This is a strategic discussion that should be made at some point anyway (beyond those already discussed in the context of sub-profiling and flags).

[0154] Regarding levels, the basic design principle is that, as a condition for bitstream conformance, even if upsampling is signaled in the bitstream, the sample count of the upsampled picture must conform to the level of the bitstream, and all samples must be upsampled. code Note that this is not the case in H263+, where certain samples may not exist.

[0155] 2.2.5. JVET-N0118 The following aspects are proposed:

[0156] 1. A list of picture resolutions is signaled in the SPS, and an index into the list is signaled in the PPS to specify the size of each individual picture.

[0157] 2. For the picture to be output, Decoded The picture is cropped (if necessary) for output, i.e. the resampled picture is not for output but is only for inter prediction reference.

[0158] 3. Support resampling ratios of 1.5x and 2x. Do not support arbitrary resampling ratios. Consider the need for one or more other resampling ratios.

[0159] 4. Between picture-level resampling and block-level resampling, proponents prefer block-level resampling.

[0160] However, if picture-level resampling is chosen, the following aspects are proposed:

[0161] i. When a reference picture is resampled, both the resampled version of the reference picture and the original resampled version are stored in the DPB, and so both affect the fullness of the DPB.

[0162] ii. A resampled reference picture is marked as "not used for reference" if the corresponding non-resampled reference picture is marked as "not used for reference".

[0163] iii. The RPL signaling syntax remains unchanged, whereas the RPL construction process is modified as follows: If a reference picture needs to be included in an RPL entry and there is no version of the reference picture in the DPB that has the same resolution as the current picture, the picture resampling process is invoked and a resampled version of that reference picture is included in the RPL entry.

[0164] iv. The number of resampled reference pictures present in the DPB should be limited, e.g., to 2 or less.

[0165] b. Otherwise (if block-level resampling is selected), the following is suggested:

[0166] i. To limit the worst-case decoder complexity, it is proposed to prohibit bidirectional prediction of blocks from reference pictures with a different resolution than the current picture.

[0167] ii. Another option is that if resampling and quarter-pixel interpolation need to be done, the two filters are combined and the operation is applied at once.

[0168] 5. Regardless of whether a picture-based or block-based resampling approach is chosen, it is proposed that temporal motion vector scaling be applied as needed.

[0169] 2.2.5.1. Implementation The ARC software was run on VTM-4.0.1 with the following modifications:

[0170] -The list of supported resolutions is signaled in the SPS.

[0171] -Spatial resolution signaling has been moved from SPS to PPS.

[0172] A picture-based resampling scheme was implemented to resample the reference pictures. Decoding After being reconstructed, the reconstructed picture may be resampled to a different spatial resolution. Both the original reconstructed picture and the resampled reconstructed picture are stored in the DPB and can be used by future pictures. Decoding are available for sequential reference.

[0173] The implemented resampling filter is based on the filter tested in JCTVC-H0234 as follows:

[0174] -Upsampling filter: 4-tap + / - 1 / 4-phase DCTIF with taps of (-4, 54, 16, -2) / 64 -Downsampling filter: h11 filter with taps (1, 0, -3, 0, 10, 16, 10, 0, -3, 0, 1) / 32 When constructing the reference picture list for the current picture (i.e., L0 and L1), only reference pictures with the same resolution as the current picture are used, although reference pictures may be available in both their original size or in a resampled size.

[0175] TMVP and ATVMP may be enabled, but the originals of the current picture and reference pictures coding If the resolutions are different, TMVP and ATMVP are disabled for that reference picture.

[0176] - For simplicity and ease of implementation of the Starting Point software, when outputting a picture the decoder outputs the highest available resolution.

[0177] Picture size and picture output signaling 1. In the bitstream code List of spatial resolutions for image-encoded pictures Currently, all CVS code The resulting pictures have the same resolution. Therefore, it is straightforward to signal only one resolution (i.e., picture width and height) in the SPS. Support for ARC requires signaling a list of picture resolutions instead of one resolution. It is proposed to signal this list in the SPS and to signal an index of the list in the PPS to specify the size of each individual picture.

[0178] 2. Picture output For the picture to be output, Decoded It is proposed that the picture is cropped (if necessary) and output, i.e. the resampled picture is not for output but only for inter-prediction reference. The ARC resampling filter needs to be designed to optimize the use of the resampled picture for inter-prediction; such a filter may not be optimal for picture output / display purposes, whereas video terminal devices usually have optimized output zooming / scaling functionality already implemented.

[0179] 2.2.5.3. About resampling Decoded Picture resampling can be either picture-based or block-based. In the final ARC design in VVC, block-based resampling is preferred over picture-based resampling. We discuss these two approaches and recommend that JVET decide which of these two to specify for ARC support in VVC.

[0180] Picture-Based Resampling In picture-based resampling for ARC, a picture is resampled only once for a particular resolution and then stored in the DPB, whereas a non-resampled version of the same picture is also maintained in the DPB.

[0181] The use of picture-based resampling for ARC has two problems: 1) an additional DPB buffer is required to store the resampled reference pictures, and 2) additional memory bandwidth is required due to the increased number of operations to read and write reference picture data from and to the DPB.

[0182] Keeping only one version of a reference picture in a DPB is not a good idea for picture-based resampling. If only the unresampled version is stored, the reference picture may need to be resampled multiple times because multiple pictures may refer to the same reference picture. On the other hand, if a reference picture is resampled and only the resampled version is kept, inverse resampling must be applied when the reference picture needs to be output, because, as mentioned above, it is better to output the unresampled picture. This is problematic because the resampling process is not a lossless operation. If you take picture A, downsample it, and then upsample it to get A' with the same resolution as A, A and A' are not the same. A' may contain less information than A because some high-frequency information was lost during the downsampling and upsampling processes.

[0183] To address the issue of additional DPB buffer and memory bandwidth, if the ARC design in VVC uses picture-based resampling, it is proposed that the following applies.

[0184] 1. When a reference picture is resampled, both the resampled version of the reference picture and the original resampled version are stored in the DPB, so both affect the fullness of the DPB.

[0185] 2. A resampled reference picture is marked as "not used for reference" if the corresponding non-resampled reference picture is marked as "not used for reference".

[0186] 3. The Reference Picture List (RPL) for each tile group contains a reference picture with the same resolution as the current picture. The RPL signaling syntax does not need to be changed, but the RPL construction process is modified as follows to ensure what was stated in the previous sentence: if a reference picture needs to be included in an RPL entry and a version of that reference picture with the same resolution as the current picture is not already available, a picture resampling process is invoked and a resampled version of that reference picture is included.

[0187] 4. The number of resampled reference pictures that can be present in a DPB should be limited, for example to two or less.

[0188] Furthermore, to enable the use of temporal MVs (e.g., merge mode and ATMVP) when the temporal MVs come from a reference frame at a different resolution than the current frame, we propose to scale the temporal MVs to the current resolution as needed.

[0189] Block-based ARC resampling In block-based resampling for ARC, reference blocks are resampled as needed and the resampled pictures are not stored in the DPB.

[0190] The main problem here is the additional decoder complexity, since a block in a reference picture can be referenced multiple times by multiple blocks in other pictures and by multiple blocks in multiple pictures.

[0191] When a block in a reference picture is referenced by a block in the current picture and the resolutions of the reference picture and the current picture are different, the reference block is resampled by invoking an interpolation filter so that the reference block has integer pixel resolution. If the motion vector is in quarter-pixel units, the interpolation process is invoked again to obtain a reference block resampled at quarter-pixel resolution. Thus, for each motion compensation operation for the current block from a reference block with a different resolution, at most two interpolation filtering operations are required instead of one. Without ARC support, at most one interpolation filtering operation (i.e., generating a reference block at quarter-pixel resolution) is required.

[0192] To limit the worst-case complexity, it is proposed that the following applies if the ARC design in VVC uses block-based resampling.

[0193] Bidirectional prediction of blocks from reference pictures with a different resolution than the current picture is prohibited.

[0194] More precisely, the constraint is as follows: if a current block blkA in a current picture picA refers to a reference block blkB in a reference picture picB, then block blkA is a unidirectionally predicted block if picA and picB have different resolutions.

[0195] This constraint allows the block to DecodingThe worst-case number of interpolation operations required to achieve this is limited to two. If a block references a block from a picture of a different resolution, the number of interpolation operations required is two, as mentioned above. This is because if a block references a reference block from a picture of the same resolution and is treated as a bidirectionally predicted block, code is the same as when the pixel is interpolated, because the number of interpolation operations is also two (i.e., one for each reference block to obtain a resolution of 1 / 4 pixel).

[0196] To simplify the implementation, if the ARC design in VVC uses block-based resampling, another variation is proposed in which the following applies.

[0197] -If the reference frame and the current frame have different resolutions, the corresponding position of each pixel of the predictor is calculated first, and then interpolation is applied only once. That is, two interpolation operations (one for resampling and one for quarter-pixel interpolation) are combined into only one interpolation operation. The sub-pel interpolation filter in the current VVC can be reused, but in this case, the granularity of interpolation should be enlarged, but the number of interpolation operations is reduced from two to one.

[0198] To enable the use of temporal MVs (e.g., merge mode and ATMVP) when the temporal MVs come from a reference frame at a different resolution than the current frame, we propose to scale the temporal MVs to the current resolution as needed.

[0199] Resampling Ratio To begin the discussion of ARC, JVET-M0135[1] proposed to consider only a resampling ratio of 2x (meaning 2x2 for upsampling and 1 / 2x1 / 2 for downsampling) as a starting point for ARC. Further discussion on this topic after the Marrakech meeting showed that supporting only a 2x resampling ratio is very limiting, since in some cases it can be beneficial to have a smaller difference between the resampled and non-resampled resolution.

[0200] Although it may be desirable to support arbitrary resampling ratios, this may be difficult to do because the number of resampling filters that would need to be defined and implemented would be too large, placing a heavy burden on the decoder implementation.

[0201] It is proposed that resampling ratios greater than 1 but small numbers (at least 1.5x and 2x resampling ratios) should be supported, and that arbitrary resampling ratios are not supported.

[0202] 2.2.5.4 Maximum DPB Buffer Size and Buffer Fullness In the ARC, DPBs are generated at different spatial resolutions within the same CVS. Decoded For DPB management and related aspects, Decoded Counting DPB size and fullness in units of pictures no longer works.

[0203] Below is a discussion of certain aspects that need to be addressed if ARC is supported, and possible solutions in the final VVC specification.

[0204] 1. PicSizeInSamplesY (i.e., PicSizeInSamplesY = pic_width_in_luma_samples * Instead of using the value of MinPicSizeInSamplesY (pic_height_in_luma_samples) to derive MaxDpSize (i.e., the maximum number of reference pictures that can be present in a DPB), the derivation of MaxDpbSize is based on the value of MinPicSizeInSamplesY, which is defined as follows:

[0205] MinPicSizeInSampleY=(Minimum picture resolution width in bitstream)*(Minimum resolution height in bitstream) The derivation of MaxDpbSize is modified (based on the HEVC formula) as follows:

[0206]

number

[0207] 2. Each Decoded A picture is associated with a value called PictureSizeUnit. Decoded An integer value that specifies how large the picture size is relative to MinPicSizeInSampleY. The definition of PictureSizeUnit depends on which resampling ratios are supported by ARC in VVC.

[0208] For example, if ARC only supports a resampling ratio of 2, then PictureSizeUnit is defined as follows:

[0209] - have a minimum resolution in the bitstream Decoded A picture is associated with a PictureSizeUnit of 1.

[0210] -Has a minimum bitstream resolution of 2x2 Decoded A picture relates to a PictureSizeUnit of 4 (ie 1x4).

[0211] In another example, if the ARC supports both resampling ratios of 1.5 and 2, then PictureSizeUnit is defined as follows:

[0212] - have a minimum resolution in the bitstream Decoded The picture is associated with a PictureSizeUnit of 4.

[0213] -Has a minimum bitstream resolution of 1.5x1.5 Decoded The picture relates to a PictureSizeUnit of 9 (ie 2.25 x 4).

[0214] -Has a minimum bitstream resolution of 2x2 Decoded A picture relates to 16 PictureSizeUnits (ie 4x4).

[0215] For other resampling ratios supported by ARC, the same principles as in the example above should be used to determine the value of PictureSizeUnit for each picture size.

[0216] 3. Let the variable MinPictureSizeUnit be the smallest possible value of PictureSizeUnit, i.e. if the ARC only supports resampling ratio 2, MinPictureSizeUnit is 1, if the ARC supports resampling ratios 1.5 and 2, MinPictureSizeUnit is 4, and the same principle for determining the value of MinPictureSizeUnit is used.

[0217] 4. The value range of sps_max_dec_pic_buffering_minus1[i] is specified as 0 to (MinPictureSizeUnit * (MaxDpbSize - 1)). The variable MinPictureSizeUnit is the smallest possible value of PictureSizeUnit.

[0218] 5. DPB fullness behavior is defined based on PictureSizeUnit as follows:

[0219] -HRD Decoding It is initialized with unit 0, and both the CPB and DPB are set to empty (DPB fullness is set to 0).

[0220] - If the DPB is flushed (i.e. all pictures are removed from the DPB), DPB fullness is set to 0.

[0221] - If a picture is removed from the DPB, the DPB fullness is decremented by the value of the PictureSizeUnit associated with the removed picture.

[0222] - If a picture is inserted into the DPB, the DPB fullness is incremented by the value of the PictureSizeUnit associated with the inserted picture.

[0223] 2.2.5.5 Resampling Filters In the software implementation, the resampling filters implemented are simply taken from the previously available filters described in JCTVC-H0234 [3]. Other resampling filters should be tested and used if they offer better performance and / or lower complexity. It is suggested to test various resampling filters to negotiate the trade-off between complexity and performance. Such testing can be done in CE.

[0224] 2.2.5.6 Other Required Modifications to Existing Tools To support ARC, some modifications and / or additional operations may be required for some existing coding tools. For example, ARC software-implemented picture-based resampling may require the originals of the current picture and reference pictures for simplicity. coding If the resolutions are different, TMVP and ATMVP are disabled.

[0225] 2.2.6 JVET-N0279 "Future Video coding According to the "Requirements for Standards," "For adaptive streaming services that offer multiple representations of the same content, each with different characteristics (e.g., spatial resolution or sample bit depth), the standard shall support fast representation switching." In real-time video communication, code Varying the resolution within a coded video sequence not only allows video data to seamlessly adapt to dynamic channel conditions or user preferences, but also can eliminate the beat effect caused by I-pictures. A hypothetical example of adaptive resolution change is shown in Figure 14, where the current picture is predicted from reference pictures of different sizes.

[0226] This contribution proposes a high-level syntax to signal adaptive resolution changes and modifications to the current motion-compensated prediction process in VTM. These modifications are limited to motion vector scaling and sub-pel position derivation, without any changes to the existing motion-compensated interpolator. This allows the existing motion-compensated interpolator to be reused, and does not require new processing blocks to support adaptive resolution changes, which would incur additional costs.

[0227] 2.2.6.1 Adaptive Resolution Change Signaling

[0228] [Table 12]

[0229] [[pic_width_in_luma_samples]] is the width of each Decoded Specifies the picture width in luma samples. pic_width_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY. pic_height_in_luma_samples is the height of each Decoded Specifies the picture height in luma samples. pic_height_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0230] max_pic_width_in_luma_samples refers to the SPS Decoded Specifies the maximum picture width in luma samples. max_pic_width_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0231] max_pic_height_in_luma_samples refers to the SPS Decoded Specifies the maximum picture height in luma samples. max_pic_height_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY.

[0232] [Table 13]

[0233] pic_size_difference_from_max_flag=1 specifies that the PPS signals a picture width or height that is different from max_pic_width_in_luma_samples and max_pic_height_in_luma_sample of the referenced SPS. pic_size_different_from_max_flag=0 specifies that pic_width_in_luma_samples and pic_height_in_luma_samples are the same as max_pic_width_in_luma_samples and max_pic_height_in_luma_sample of the referenced SPS.

[0234] pic_width_in_luma_samples is the width of each DecodedSpecifies the picture width in luma samples. pic_width_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY. If pic_width_in_luma_samples is not present, it is inferred to be equal to max_pic_width_in_luma_samples.

[0235] pic_height_in_luma_samples is the height of each Decoded Specifies the picture height in luma samples. pic_height_in_luma_samples must not be equal to 0 and must be an integer multiple of MinCbSizeY. If pic_height_in_luma_samples is not present, it is inferred to be equal to max_pic_height_in_luma_samples.

[0236] The horizontal and vertical scaling ratios for all active reference pictures must be in the range of 1 / 8 to 2 for bitstream conformance. The scaling ratios are defined as follows:

[0237] horizontal_scaling_ratio=((reference_pic_width_in_luma_samples<<14)+(pic_width_in_luma_samples / 2)) / pic_width_in_luma_samples vertical_scaling_ratio=((reference_pic_height_in_luma_samples<<14)+(pic_height_in_luma_samples / 2)) / pic_height_in_luma_samples

[0238] [Table 14]

[0239] Reference Picture Scaling Process When there is a resolution change in a CVS, a picture may have a different size than one or more of its reference pictures. This proposal normalizes all motion vectors to the grid of the current picture instead of the grid of their corresponding reference pictures. This is claimed to be beneficial in maintaining design consistency and making resolution changes transparent to the motion vector prediction process. Otherwise, adjacent motion vectors pointing to reference pictures of different sizes cannot be directly used for spatial motion vector prediction due to different scaling.

[0240] When a resolution change occurs, both the motion vectors and reference blocks need to be scaled during motion compensated prediction. The scaling range is limited to [1 / 8, 2], i.e., upscaling is limited to 1:8 and downscaling is limited to 2:1. Note that upscaling refers to the case where the reference picture is smaller than the current picture, while downscaling refers to the case where the reference picture is larger than the current picture. The following section describes the scaling process in more detail.

[0241] Luma Block The scaling factors and their fixed point representations are defined as follows:

[0242]

number

[0243]

number

[0244] 1. Map the top left corner pixel of the current block to the reference picture.

[0245] 2. Use horizontal and vertical step sizes to specify the reference positions of other pixels in the current block.

[0246] If the coordinates of the upper left corner pixel of the current block are (x, y), the sub-pel position (x^', y^') in the reference picture pointed to by the motion vector (mvX, mvY) in units of 1 / 16 pixels is specified as follows:

[0247] The horizontal position of the reference picture is as follows:

[0248]

number

[0249] x' is further reduced to retain only 10 bits.

[0250]

number

[0251] Similarly, the vertical position of the reference picture is:

number

[0252] y' is further reduced.

[0253]

number

[0254] At this point, the reference position of the pixel in the upper left corner of the current block is at (x^', y^'). Other reference sub-pixel positions are calculated relative to (x^', y^') with horizontal and vertical step sizes. These step sizes are derived from the horizontal and vertical scaling factors above with an accuracy of 1 / 1024 pixels as follows:

[0255]

number

[0256]

number

[0257]

number

[0258]

number

[0259] For sub-pel interpolation, x' i , y' j must be split into a full-pel and a partial-pel part.

[0260] The full pel part for addressing the reference block is equal to:

[0261]

number

[0262]

number

[0263] The fractional pels that can be used to select an interpolation filter are equal to:

[0264]

number

[0265]

number

[0266] Once the full-pel and fractional-pel locations in the reference picture are identified, the existing motion compensation can be used without any additional modifications: the full-pel locations are used to obtain reference block patches from the reference picture, and the fractional-pel locations are used to select the appropriate interpolation filter.

[0267] Chroma Block If the chroma format is 4:2:0, the chroma motion vectors have an accuracy of 1 / 32 pixel. The scaling process of chroma motion vectors and chroma reference blocks is almost the same as that of luma blocks, except for the chroma format-related adjustments.

[0268] If the coordinates of the top left corner pixel of the current chroma block are (xc, yc), the initial horizontal and vertical positions of the reference chroma picture are as follows:

[0269]

number

[0270]

number

[0271] Here, mvX and mvY are the original luma motion vectors, but now they should be checked to 1 / 32 pixel accuracy.

[0272] To maintain accuracy to 1 / 1024 pixel, xc' and yc' are further scaled down.

[0273]

number

[0274]

number

[0275] Compared to the associated luma equation, the above shift to the right increases by 1 bit.

[0276] The step size used is the same as for luma. For a chroma pixel at (i,j) relative to the top-left pixel, the horizontal and vertical coordinates of its reference pixel are derived below:

[0277]

number

[0278]

number

[0279] For sub-pel interpolation, x c ' i , y c ' j is also divided into a full pel portion and a partial pel portion.

[0280] The full pel part for addressing the reference block is equal to:

[0281]

number

[0282]

number

[0283] The fractional pel portion used to select the interpolation filter is equal to:

[0284]

number

[0285]

number

[0286] other coding Interacting with the tool Some coding Due to the increased complexity and memory bandwidth associated with the interaction of tools and reference picture scaling, it is recommended to add the following constraints to the VVC specification:

[0287] If tile_group_temporal_mvp_enabled_flag=1, the current picture and its collocated picture shall have the same size.

[0288] If resolution changes are allowed within a sequence, the decoder's motion vector refinement should be turned off.

[0289] If resolution changes are allowed within a sequence, sps_bdof_enabled_flag=0.

[0290] 2.3 In JVET-N0415 coding Tree Block (CTB)-based Adaptive Loop Filter (ALF) Slice-level time filters VTM4 introduced adaptive parameter sets (APS). Each APS contains a set of signaling ALF filters, and up to 32 APSs are supported. Slice-level temporal filters were tested in the proposal. Tile groups can reuse ALF information from APSs to reduce overhead. APSs are updated as a first-in, first-out (FIFO) buffer.

[0291] CTB-based ALF For the luma component, when ALF is applied to the luma CTB, a choice of 16 fixed filter sets, 5 temporal filter sets, or 1 signal filter set is indicated. Only the filter set index is signaled. Only one of the 25 new filter sets can be signaled per slice. If a new set is signaled for a slice, all luma CTBs in the same slice share that set. The fixed filter set can be used to predict a new slice-level filter set and can be used as a filter set candidate for the luma CTB. The total number of filters is 64.

[0292] For chroma components, if an ALF is applied to the chroma CTB, if a new filter is signaled for the slice, the CTB uses the new filter; otherwise, the latest temporal chroma filter that satisfies the temporal scalability constraints is applied.

[0293] As a slice-level temporal filter, the APS is updated as a first-in, first-out (FIFO) buffer.

[0294] 2.4 Alternative Temporal Motion Vector Prediction (also known as Sub-Block Based Temporal Merge Candidates in VVC) In the alternative temporal motion vector prediction (ATMVP) method, motion vector temporal motion vector prediction (TMVP) is modified by retrieving multiple sets of motion information (including motion vectors and reference indexes) from blocks smaller than the current CU. As shown in Figure 14, a sub-CU is a square NxN block (N is set to 8 by default).

[0295] ATMVP predicts motion vectors for sub-CUs within a CU in two steps. The first step is to identify the corresponding block in a reference picture using a so-called temporal vector. This reference picture is called the motion source picture. The second step is to divide the current CU into sub-CUs and obtain a motion vector from the block corresponding to each sub-CU, along with the reference index for each sub-CU, as shown in Figure 15, which shows an example of ATMVP motion prediction for a CU.

[0296] In the first step, the reference picture and corresponding blocks are determined by the motion information of the spatially neighboring blocks of the current CU. To avoid the repeated scanning process of neighboring blocks, a merge candidate from block A0 (left block) in the merge candidate list of the current CU is used. The first available motion vector from block A0, which refers to the collocated reference picture, is set as the temporal vector. In this way, in ATMVP, corresponding blocks can be identified more accurately than in TMVP. The corresponding block (sometimes called the collocated block) is always located in the bottom right or center position with respect to the current CU.

[0297] In the second step, the corresponding block of the sub-CU is identified by a time vector in the motion source picture by adding the time vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) is used to derive the motion information of the sub-CU. After the motion information of the corresponding NxN block is identified, it is converted into the motion vector and reference index of the current sub-CU, and motion scaling and other procedures are applied, similar to TMVP in HEVC.

[0298] 2.5 Affine Motion Estimation In HEVC, only a translation motion model is applied for motion compensation prediction (MCP). The real world may have many types of motion, such as zoom in / out, rotation, projective motion, and other irregular motions. In VVC, simplified affine transformation motion compensation prediction is applied using a 4-parameter affine model and a 6-parameter affine model. Figures 16a and 16b show the simplified 4-parameter affine model and the 6-parameter affine model, respectively. As shown in Figures 16a and 16b, the affine motion field of a block is described by two control point motion vectors (CPMVs) for the 4-parameter affine model and three CPMVs for the 6-parameter affine model.

[0299] The motion vector field (MVF) of a block is described by the four-parameter affine model of equation (1) (where the four parameters are defined as variables a, b, e, and f) and the six-parameter affine model of equation (2) (where the six parameters are defined as a, b, c, d, e, and f), respectively, as follows:

[0300]

number

[0301] where (mvh0,mvh0) is the motion vector of the control point in the upper left corner, (mvh1,mvh1) is the motion vector of the control point in the upper right corner, and (mvh2,mvh2) is the motion vector of the control point in the lower left corner. All three motion vectors are called control point motion vectors (CPMVs), (x,y) represent the coordinates of the representative point relative to the upper left sample in the current block, and (mvh(x,y),mvv(x,y)) is the motion vector derived for the sample located at (x,y). The CP motion vectors may be signaled (similar to affine AMVP mode) or derived on the fly (similar to affine mode). w and h are the width and height of the current block. In practice, the division is performed by right shift with rounding operation. In VTM, the representative point is defined to be the center position of a sub-block, for example, if the coordinates of the upper-left corner of the sub-block relative to the upper-left sample in the current block are (xs, ys), the coordinates of the representative point are defined to be (xs+2, ys+2). For each sub-block (i.e., 4x4 in VTM), the representative point is used to derive a motion vector for the entire sub-block.

[0302] To further simplify motion compensation prediction, sub-block-based affine transformation prediction is applied. To derive a motion vector for each M×N sub-block (M and N are both set to 4 in the current VVC), the motion vector of the center sample of each sub-block shown in Figure 17 can be calculated according to equations (25) and (26) and rounded to 1 / 16 fractional precision. Then, a motion compensation interpolation filter for 1 / 16 pel is applied to generate a prediction for each sub-block according to the derived motion vector. The interpolation filter for 1 / 16 pel is introduced by the affine mode.

[0303] After the MCP, the high-precision motion vectors for each sub-block are rounded and saved to the same precision as the normal motion vectors.

[0304] 2.5.1 Signaling Affine Predictions Similar to the translational model, there are also two modes for signaling side information with the affine model: AFFINE_INTER mode and AFFINE_MERGE mode.

[0305] 2.5.2 AF_INTER mode For CUs whose width and height are both greater than 8, the AF_INTER mode may be applied. An affine flag at the CU level is signaled in the bitstream to indicate whether the AF_INTER mode is used.

[0306] In this mode, for each reference picture list (List0 or List1), an affine AMVP candidate list is constructed with three types of affine motion predictors in the following order, with each candidate containing an estimated CPMV for the current block: The difference between the best CPMV found at the encoder side (e.g., mv0, mv1, mv2 in Figures 18a and 18b) and the estimated CPMV is signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is further signaled.

[0307] 1) Genetic Affine Motion Predictor The test order is similar to that of spatial MVP in the HEVC AMVP list structure. First, the left genetic affine motion predictor is derived from the first block in {A1, A0} that has the same reference picture as the current block and is affine coded. Second, the top genetic affine motion predictor is derived from the first block in {B1, B0, B2} that has the same reference picture as the current block and is affine coded. Five blocks A1, A0, B1, B0, B2 are shown in Figure 19.

[0308] If it is found that the neighboring blocks are coded in affine mode, the CPMV of the coding unit covering the neighboring blocks is used to derive the predictor of the CPMV of the current block. For example, if A1 is coded in non-affine mode and A0 is coded in 4-parameter affine mode, the left side genetic affine MV predictor will be derived from A0. In this case, for the top left CPMV in Figure 18b, MV0 N and MV1 for the upper right CPMV. N The CPMV of the CU covering A0, represented by C , MV1 C , MV2 C is utilized to derive the estimated CPMV of the current block, denoted by

[0309] 2) Constructed affine motion predictor The constructed affine motion predictor consists of control point motion vectors (CPMVs) derived from neighboring inter-coded blocks that have the same reference picture, as shown in Figure 20. If the current affine motion model is a 4-parameter affine, the number of CPMVs is 2; otherwise, if the current affine motion model is a 6-parameter affine, the number of CPMVs is 3. Top-left CPMV [Outside 1] TIFF0007807501000045.tif10127 (hereafter referred to as bar mv0) is derived from the MV of the first block in the inter-coded group {A, B, C} that has the same reference picture as the current block. [Outside 2] TIFF0007807501000046.tif10127 (hereafter bar mv1) is derived from the MV of the first block in the inter-coded group {D, E} that has the same reference picture as the current block. [Outside 3] TIFF0007807501000047.tif10127 (hereafter referred to as mv2) is derived from the MV of the first block in the inter-coded group {F, G} that has the same reference picture as the current block.

[0310] -If the current affine motion model is a four-parameter affine, the constructed affine motion predictor is inserted into the candidate list only if both bar mv0 and bar mv1 are determined, i.e., bar mv0 and bar mv1 are used as the estimated CPMVs for the top-left position (having coordinates (x0, y0)) and top-right position (having coordinates (x1, y1)) of the current block.

[0311] - If the current affine motion model is a 6-parameter affine, the constructed affine motion predictor is inserted into the candidate list only if bar mv0, bar mv1, and bar mv2 are all determined, i.e., bar mv0, bar mv1, and bar mv2 are used as estimated CPMVs for the top-left position (having coordinates (x0, y0)), top-right position (having coordinates (x1, y1)), and bottom-right position (having coordinates (x2, y2)) of the current block.

[0312] When inserting the constructed affine motion predictor into the candidate list, no pruning process is applied.

[0313] 3) Regular AMVP motion predictor The following is applied until the number of affine motion predictors reaches a maximum value.

[0314] i. Derive affine motion predictors by setting all CPMVs equal to mv2 if available.

[0315] ii. Derive affine motion predictors by setting all CPMVs equal to mv1 if available.

[0316] iii. Derive affine motion predictors by setting all CPMVs equal to mv0 if available.

[0317] iv. Derive affine motion predictors by setting all CPMVs equal to HEVC TMVP if available.

[0318] v. Derive the affine motion predictor by setting all CPMVs to zero MV.

[0319] It should be noted that [Outside 4] TIFF0007807501000048.tif10127 (hereafter, bar mv i ) is a point that has already been derived in the constructed affine motion.

[0320] In AF_INTER mode, two or three control points are required when the four or six parameter affine modes are used, so two or three MVDs need to be coded for those control points, as shown in Figures 18a and 18b. JVET-K0337 proposes to derive the MVs as follows: mvd1 and mvd2 are predicted from mvd0.

[0321]

number

[0322] Here, bar mv i , mvd i and MV iare the predicted motion vector, motion vector differential, and motion vector of the top-left pixel (i=0), top-right pixel (i=1), or bottom-left pixel (i=2), respectively, as shown in Figure 18(b). Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two components separately. That is, newMV = mvA + mvB, and the two components of newMV are set to (xA + xB) and (yA + yB), respectively.

[0323] 2.5.2.1 AF_MERGE Mode When a CU is applied in AF_MERGE mode, it obtains the first block coded in affine mode from valid neighboring reconstructed blocks. Figure 21 shows candidates for AF_MERGE. The selection order of candidate blocks is from left to right, top, top right, bottom left, top left, as shown in Figure 21a (represented by A, B, C, D, E, respectively). For example, if the neighboring bottom left block is coded in affine mode, as represented by A0 in Figure 21b, the control point (CP) motion vectors mv0 of the top left, top right, and bottom left corners of the neighboring CU / PU containing block A are N , mv1 N and mv2 N Then, the upper left corner / upper right / lower left motion vector mv0 on the current CU / PU is fetched. C , mv1 C and mv2 C (Only used for 6-parameter affine models) is mv0 N , mv1 N and mv2 NIn VTM-2.0, if the current block is affine coded, the sub-block located in the upper left corner (e.g., a 4x4 block in VTM) stores mv0, and the sub-block located in the upper right corner stores mv1. If the current block is coded using a 6-parameter affine model, the sub-block located in the lower left corner stores mv2; otherwise (using a 4-parameter affine model), the LB stores mv2'. The other sub-blocks store the MV used for MC.

[0324] The MVF of the current CU is generated after the CPMVs of the current CU, mv0C, mv1C, and mv2C, are derived according to the simplified affine motion model, Equations 25 and 26. To identify whether the current CU is coded in AF_MERGE mode, an affine flag is signaled in the bitstream if there is at least one neighboring block coded in affine mode.

[0325] In JVET-L0142 and JVET-L0632, the affine merge candidate list is constructed using the following steps:

[0326] 1) Insertion of genetic affine candidates An inherited affine candidate means that the candidate is derived from the affine motion model of its valid neighboring affine-coded blocks. Up to two inherited affine candidates are derived from the affine motion models of neighboring blocks and inserted into the candidate list. For the left predictor, the scan order is {A0, A1}, and for the upper predictor, the scan order is {B0, B1, B2}.

[0327] 2) Insertion of constructed affine candidates If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (e.g., 5), constructed affine candidates are inserted into the candidate list. Constructed affine candidates means that the candidates are constructed by combining neighboring motion information of each control point.

[0328] a) The motion information of a control point is first derived from the designated spatial and temporal neighborhoods shown in Figure 22. CPk (k=1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial positions for predicting CPk (k=1, 2, 3). T is the temporal position for predicting CP4.

[0329] The coordinates of CP1, CP2, CP3 and CP4 are (0,0), (W,0), (H,0) and (W,H), respectively, where W and H are the width and height of the current block.

[0330] The motion information for each control point is obtained according to the following priority: -For CP1, the check priority is B2 → B3 → A2. B2 is used if it is available. Otherwise, if B3 is available, B3 is used. If both B2 and B3 are unavailable, A2 is used. If all three candidates are unavailable, the motion information of CP1 cannot be obtained.

[0331] For -CP2, the check priority is B1 → B0.

[0332] -For CP3, the check priority is A1 → A0.

[0333] For -CP4, T is used.

[0334] b) Second, a combination of control points is used to construct affine merge candidates.

[0335] Motion information of three control points is required to construct a six-parameter affine candidate. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4} will be transformed into a six-parameter motion model represented by the top-left, top-right, and bottom-left control points.

[0336] - Motion vectors of two control points are required to construct a 4-parameter affine candidate. The two control points can be selected from one of the following two combinations ({CP1, CP2}, {CP1, CP3}). The two combinations will be transformed into a 4-parameter motion model represented by the top-left and top-right control points.

[0337] -The constructed affine candidate combinations are inserted into the candidate list in the following order: {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}.

[0338] i. For each combination, the reference indexes in list X for each CP are checked, and if they are all the same, then this combination has a valid CPMV in list X. If the combination does not have valid CPMVs in both list 0 and list 1, then this combination is marked invalid. Otherwise, it is valid and the CPMVs are placed in the sub-block merge list.

[0339] 3) Padding with zero-affine motion vector candidates If the number of candidates in the affine merge candidate list is less than five, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.

[0340] More specifically, for the subblock merging candidate list, the four-parameter merging candidate has MV set to (0, 0) and prediction direction set to unidirectional prediction (for P slices) and bidirectional prediction (for B slices) from list 0.

[0341] Shortcomings of existing implementations When applied to VVC, ARC may have the following problems:

[0342] ALF, Luma Mapping Chroma Scaling (LMCS), Decoder-Side Motion Vector Refinement (DMVR), Bidirectional Optical Flow (BDOF), Affine Prediction, Triangular Prediction Mode (TPM), Symmetric Motion Vector Difference (SMVD), Merged Motion Vector Difference (MMVD), Inter-Intra Prediction (also known as Combined Inter-Picture and Intra-Picture Merging (CIIP) in VVC), Localized Illumination Compression (LIC), and History-Based Motion Vector Prediction (HMVP). coding It is unclear whether the tool will be applied to VVC.

[0343] Adaptive resolution conversion tools coding Example methods for Embodiments of the disclosed technology overcome the shortcomings of existing implementations. The following examples of the disclosed technology are discussed to facilitate understanding of the disclosed technology and should be construed in a manner that limits the disclosed technology. Unless otherwise specified, various features described in these examples can be combined.

[0344] In the following discussion, SatShift(x, n) is defined as follows:

[0345]

number

[0346] Shift(x, n) is defined as Shift(x, n)=(x+offset0)>>n.

[0347] In one example, offset0 and / or offset1 are (1<<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.

[0348] Another example is offset0=offset1=((1<<n)> >1)-1 or ((1<<(n-1)))-1.

[0349] Clip3(min, max, x) is defined as follows:

[0350]

number

[0351] Floor(x) is defined as the largest integer less than or equal to x.

[0352] Ceil(x) is the smallest integer greater than or equal to x.

[0353] Log2(x) is defined as the base-2 logarithm of x.

[0354] Some aspects of the implementation of the disclosed technology are listed below.

[0355] 1. MV offset in MMVD / SMVD and / or De The derivation of the refined motion vector in the coder-side derivation process may depend on the resolution of the reference picture associated with the current block and the resolution of the current picture. For example, a second MV offset referencing a second reference picture may be scaled from a first MV offset referencing a first reference picture, and the scaling factor may depend on the resolution of the first and second reference pictures.

[0356] 2. The motion candidate list building process can be built according to the reference picture resolution associated with the spatial / temporal / historical motion candidates. a. In one example, a merge candidate that refers to a reference picture with a higher resolution may have a higher priority than a merge candidate that refers to a reference picture with a lower resolution. In the discussion, when W0 < W1 and H0 < H1, the resolution W0 * H0 is lower than the resolution W1 * H1. b. For example, in the merge candidate list, a merge candidate that refers to a reference picture with a higher resolution may be placed before a merge candidate that refers to a reference picture with a lower resolution. c. For example, a motion vector that refers to a reference picture having a resolution lower than that of the current picture cannot be put into the merge candidate list. d. In one example, whether and / or how to update the history buffer (reference table) Decoded may depend on the reference picture resolution related to the motion candidate. i. In one example, Decoded if one of the reference pictures related to the motion candidate has a different resolution, such a motion candidate is not permitted to update the history buffer.

[0357] 3. It is proposed that pictures should be filtered with ALF parameters related to the corresponding dimensions. a. In one example, the ALF parameters signaled in a video unit such as APS can be associated with one or more image dimensions. b. In one example, a video unit such as the APS signaling ALF parameter can be associated with one or more image dimensions. c. For example, an image can apply only the ALF parameters signaled by a video unit such as APS associated with the same dimensions. d. The display of the resolution / index / resolution of PPS can be signaled with ALF APS. e. It is restricted that the ALF parameters may be inherited / predicted only from those used for images with the same resolution.

[0358] 4. It is proposed that the ALF parameters associated with a first corresponding dimension may inherit or be predicted from the ALF parameters associated with a second corresponding dimension. In one example, the first corresponding dimension must be the same as the second corresponding dimension. b. In one example, the first corresponding dimension may be different from the second corresponding dimension.

[0359] 5. It is proposed that the samples in the picture should be reshaped with the LMCS parameters related to the corresponding dimensions. In one example, LMCS parameters signaled in a video unit such as an APS may relate to one or more picture dimensions. b. In one example, a video unit such as an APS that signals LMCS parameters may be associated with one or more picture dimensions. c. For example, a picture may only apply LMCS parameters signaled in a video unit such as an APS that is associated with the same dimensions. d. PPS resolution / index / resolution indication may be signaled in the LMCS APS. e.LMCS parameters are restricted to only be inherited / predicted from those used for pictures with the same resolution.

[0360] 6. It is proposed that the LMCS parameters associated with a first corresponding dimension may inherit or be predicted from the LMCS parameters associated with a second corresponding dimension. In one example, the first corresponding dimension must be the same as the second corresponding dimension. b. In one example, the first corresponding dimension may be different from the second corresponding dimension.

[0361] 7. TPM (Triangular Prediction Mode) / GEO (Inter Prediction with Geometric Partitioning) or other methods that can divide a block into two or more sub-partitions codingWhether and / or how the tool is enabled depends on the associated reference picture information of the two or more sub-partitions. In one example, it may depend on the resolution of one of the two reference pictures and the resolution of the current picture. In one example, if at least one of the two reference pictures is associated with a different resolution compared to the current picture, such coding The tool is disabled. b. In one example, it may depend on whether the resolution of the two reference pictures is the same. In one example, if two reference pictures are associated with different resolutions, such coding The tool can be disabled. ii. In one example, if two reference pictures both relate to different resolutions compared to the current picture, such coding The tool is disabled. iii. Alternatively, if both of the two reference pictures are associated with a different resolution compared to the current picture, but the two reference pictures have the same resolution, such coding The tool can still be disabled. iv. Alternatively, if at least one of the reference pictures has a different resolution than the current picture, and the reference pictures have different resolutions, coding Tool X can be disabled. c. Alternatively, it may further depend on whether the two reference pictures are the same reference picture. d. Alternatively, or in addition, it may depend on whether the two reference pictures are in the same reference list. e. Or such coding The tool can always be disabled in case of RPR (reference picture resampling is enabled in the slice / picture header / sequence parameter set).

[0362] 8. For a block, if the block references at least one reference picture that has different dimensions than the current picture: codingIt is suggested that tool X may be disabled. a. In one example, coding Information about tool X does not need to be signaled. b. In one example, the motion information for such blocks may not be inserted into the HMVP table. c. Or, to a block coding When tool X is applied, a block cannot refer to a reference picture with different dimensions than the current picture. i. In one example, merge candidates that refer to reference pictures with different dimensions than the current picture may be skipped or not included in the merge candidate list. ii. In one example, reference indices corresponding to reference pictures with different dimensions than the current picture may be skipped or not allowed to be signaled. d. Or, coding Tool X may be applied after scaling two reference blocks or pictures according to the resolution of the current picture and the resolution of the reference picture. e. Or, coding Tool X may be applied after scaling the two MVs or MVDs according to the resolution of the current picture and the resolution of the reference picture. f. In one example, a block (e.g., a block with multiple hypotheses from the same reference picture list with different reference pictures or different MVs, or a bidirectional Encoding For a block (or a block with multiple hypotheses from different reference picture lists), coding Whether tool X is disabled or enabled may depend on the resolution of the reference picture associated with the reference picture list and / or the current reference picture. i. In one example, coding Tool X may be disabled for one reference picture list, but enabled for the other reference picture list. ii. In one example, codingTool X may be disabled for one reference picture but enabled for the other reference picture, where the two reference pictures may be from different reference picture lists or the same reference picture list. iii. In one example, for each reference picture list L, coding The enabling / disabling of the tool is determined regardless of the reference pictures in other reference picture lists different from list L. 1. In one example, it can be determined by the reference pictures in list L and the current picture. 2. In one example, if the associated reference picture of list L is different from the current picture, the tool may be disabled for list L. iv. Or coding The enabling / disabling of the tool is determined by the resolution of all reference pictures and / or the current picture. 1. In one example, if at least one of the reference pictures has a different resolution than the current picture, coding Tool X can be disabled. 2. In one example, if at least one of the reference pictures has a different resolution than the current picture, but the reference pictures have the same resolution, coding Tool X can still be enabled. 3. In one example, if at least one of the reference pictures has a different resolution than the current picture and the reference pictures have different resolutions, coding Tool X can be disabled. g. coding Tool X can be any of the following: iii.DMVR iv.BDOF v. Affine prediction vi. Triangle Prediction Mode vii.SMVD viii.MMVD ix. Inter-intra prediction in VVC x.LIC xi.HMVP xii. Multiple Transformation Set (MTS) xiii. Sub-Block Transform (SBT) xiv.PROF and / or other Decoding Side Movement / Prediction Refinement Method xv.LFNST (Low Frequency Non-Squaring Transform) xvi. Filtering method (e.g. deblocking filter / SAO / ALF) xvii.GEO / TPM / Cross-Component ALF

[0363] 9. The reference picture list of a picture contains no more than K different resolutions. In one example, K=2.

[0364] 10. N consecutive ( Decoding No more than K different resolutions are allowed for pictures (order or display order). In one example, N=3 and K=3. b. In one example, N=10 and K=3. c. In one example, no more than K different resolutions are allowed in a GOP. d. In one example, no more than K different resolutions are allowed between two pictures with the same specific temporal layer ID (denoted as tid). i. For example, K=3, tid=0.

[0365] 11. Resolution changes are only allowed for intra pictures.

[0366] 12. If one or two reference pictures of a block have a different resolution than the current picture, Decoding In the process, bi-prediction can be converted to uni-prediction. In one example, predictions from list X that have corresponding reference pictures with a different resolution than the current picture may be discarded.

[0367] 13. Whether to enable or disable inter prediction from reference pictures of different resolutions may depend on the motion vector accuracy and / or the resolution ratio. In one example, if the motion vector scaled according to the resolution ratio points to an integer position, inter prediction may still be applied. b. In one example, if the motion vector scaled according to the resolution ratio points to a sub-pel position (eg, 1 / 4 pel) that is allowed without ARC, inter prediction may still be applied. c. Alternatively, if both reference pictures have a different resolution than the current picture, bidirectional prediction may not be allowed. d. Alternatively, bidirectional prediction may be enabled when one reference picture has a different resolution than the current picture and the other has the same resolution. e. Alternatively, if the reference picture has a different resolution than the current picture and the block dimensions meet certain conditions, unidirectional prediction may be prohibited for that block.

[0368] 14. Ko A first flag (eg, pic_disable_X_flag) indicating whether loading tool X is disabled may be signaled in the picture header. a. For slices / tiles / bricks / subpictures / other video units smaller than a picture Ruko Whether the loading tool is enabled can be controlled by this flag in the picture header and / or the slice type. b. In one example, if the first flag is true, ,Ko Loading tool X is disabled. i. Or if the first flag is false ,Ko Reading Tool X is enabled. ii. In one example, it is enabled / disabled for all samples in the picture. c. In one example, the signaling of the first flag may depend on a syntax element or multiple syntax elements in the SPS / VPS / DPS / PPS. In one example, the signaling of flags is RukoThis may depend on the valid flags of loading tool X. ii. Alternatively, in addition, a second flag (eg, sps_X_slice_present_flag) may be signaled in the SPS to indicate the presence of the first flag in the picture header. 1) Alternatively, and additionally, a second flag may be conditionally signaled if coding tool X is enabled for the sequence (eg, if sps_X_enabled_flag is true). 2) Alternatively and additionally, only the second flag indicates the presence of the first flag, which may be signaled in the picture header. d. In one example, the first and / or second flags are coded with one bit. e .Ko Loading tool X can be: i. In one example ,Ko Reading tool X is PROF. ii. In one example ,Ko Reading Tool X is a DMVR. iii. In one example ,Ko Loading Tool X is BDOF iv. In one example ,Ko The loading tool X is a cross-component ALF. v. In one example ,Ko Reading Tool X is GEO. VI. In one example ,Ko Loading tool X is a TPM. vii. In one example ,Ko Loading Tool X is MTS.

[0369] 15. Whether a block can refer to a reference picture with different dimensions than the current picture may depend on the width (WB) and / or height (HB) of the block and / or the block prediction mode (i.e., bidirectional prediction or short-directional prediction). In one example, a block can refer to a reference picture with different dimensions than the current picture if WB>=T1 and HB>=T2, eg, if T1=T2=8. b. In one example, a block can refer to a reference picture with different dimensions than the current picture if WB*HB>=T, eg, T1=64. c. In one example, a block can refer to a reference picture with different dimensions than the current picture if Min(WB,HB)>=T, eg, if T=8. d. In one example, a block can reference a reference picture with different dimensions than the current picture if Max(WB,HB)>=T, eg, if T=8. e. In one example, a block can refer to a reference picture with different dimensions than the current picture if WB<=T1 and HB<=T2, eg, if T1=T2=62. f. In one example, a block can reference a reference picture with different dimensions than the current picture if WB*HB<=T, eg, T1=4096. g. In one example, a block can reference a reference picture with different dimensions than the current picture if Min(WB,HB)<=T, eg, if T=64. h. In one example, a block can reference a reference picture with different dimensions than the current picture if Max(WB,HB)<=T, eg, if T=64. i. Alternatively, a block is prohibited from referencing a reference picture with different dimensions than the current picture if WB<=T1 and / or HB<=T2, e.g., if T1=T2=8. j. Alternatively, a block is prohibited from referencing a reference picture with different dimensions than the current picture if WB<=T1 and / or HB<=T2, e.g., if T1=T2=8.

[0370] 23 shows a flowchart of an example method for video processing. Referring to FIG. 23, method 2300 includes, at step 2310, deriving a motion vector offset based on a resolution of a reference picture associated with the current video block and a resolution of the current picture during conversion between the current video block and a bitstream representation of the current video block. Method 2300 further includes, at step 2320, performing the conversion using the motion vector offset.

[0371] In some implementations, deriving the motion vector offset includes deriving a first motion vector offset that references a first reference picture, and deriving a second motion vector offset that references a second reference picture based on the first motion vector offset. In some implementations, the method further includes performing a reference picture resolution-based motion candidate list building process for the current video block associated with spatial, temporal, or historical motion candidates. In some implementations, whether or how to update the lookup table is determined by: Decorated The motion candidate resolution is determined by a reference picture resolution associated with the motion candidate. In some implementations, the method further includes performing a filtering operation on the current picture using adaptive loop filter (ALF) parameters associated with the corresponding dimensions. In some implementations, the ALF parameters include a first ALF parameter associated with the first corresponding dimension and a second ALF parameter associated with the second corresponding dimension, where the second ALF parameter is inherited or predicted from the first ALF parameter.

[0372] In some implementations, the method further includes reshaping the samples in the current picture using LMCS (Luma Mapping Chroma Scaling) parameters associated with the corresponding dimensions. In some implementations, the LMCS parameters include a first LMCS parameter associated with the first corresponding dimension and a second LMCS parameter associated with the second corresponding dimension, where the second LMCS parameter is inherited or predicted from the first LMCS parameter. In some implementations, the ALF parameters or LMCS parameters signaled in the video unit are associated with one or more picture dimensions. In some implementations, the method includes reshaping the samples in the current picture when the current video block references at least one reference picture having different dimensions than the current picture. coding In some implementations, the method further includes skipping or omitting merge candidates that reference a reference picture having dimensions different from dimensions of the current picture. In some implementations, the method includes, after scaling the two reference blocks or the two reference pictures based on the resolution of the reference picture and the resolution of the current picture, or after scaling the two MVs or MVDs based on the resolution of the reference picture and the resolution of the current picture, coding In some implementations, the current picture includes no more than K different resolutions, where K is a natural number. In some implementations, K different resolutions are allowed for N consecutive pictures, where N is a natural number.

[0373] In some implementations, the method further includes applying a resolution change to the current picture, which is an intra picture. In some implementations, the method further includes converting bidirectional prediction to unidirectional prediction when one or two reference pictures of the current video block have a different resolution than the resolution of the current picture. In some implementations, the method further includes enabling or disabling inter-prediction from a reference picture of a different resolution depending on at least one of a resolution ratio between the current block dimension and the reference block dimension or motion vector precision. If the current block dimension is W*H, the reference block dimension is W1*H1, and the resolution ratio may refer to W1 / W, H1 / H, (W1*H1) / (W*H), max(W1 / W, H1 / H), or min(W1 / W, H1 / H), etc. In some implementations, the method further includes applying bidirectional prediction depending on whether both reference pictures or one reference picture has a different resolution than the current picture. In some implementations, whether the current video block references a reference picture having different dimensions than the current picture depends on at least one of a size of the current video block or a block prediction mode. In some implementations, performing the transformation includes generating the current video block from a bitstream representation. In some implementations, performing the transformation includes generating the bitstream representation from the current video block.

[0374] FIG. 24A is a block diagram of a video processing device 2400. The device 2400 may be used to implement one or more of the methods described herein. The device 2400 may be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 2400 may include one or more processors 2402, one or more memories 2404, and video processing hardware 2406. The processor 2402 may be configured to implement one or more methods described herein (including, but not limited to, method 2300). The memory(s) 2404 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 2406 may be a hardware circuit and may be used to implement some of the techniques described herein. In some implementations, the hardware 2406 may be fully or partially within the processor 2401, for example, as a graphics processor.

[0375] 24B is a block diagram illustrating an example video processing system 2410 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 24100. System 2410 may include an input 2412 for receiving video content. The video content may be received in raw or uncompressed form, e.g., 8- or 10-bit multi-component pixel values, or may be compressed or Encoding The input 2412 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0376] The system 2410 may implement various coding or encoding methods described herein. coding Component 2414 may be included. codingThe component 2414 reduces the average bit rate of the video from the input 2412 to the output of the coding component 2414 to reduce the code Therefore, the coding techniques are sometimes called video compression or video transcoding techniques. coding The output of component 2414 may be stored or transmitted over a communication connection, as represented by component 2416. A stored or communicated bitstream (or code The decompressed (compressed) representation may be used by component 2418 to generate displayable video or pixel values ​​that are sent to display interface 2420. The process of generating user-viewable video from the bitstream representation is sometimes called video decompression. Furthermore, certain video processing operations are referred to as " coding " It is called an action or tool, coding The tool or action is used in the encoder, coding Invert the result of the corresponding Decoding It will be understood that the tools or operations will be performed at the decoder.

[0377] Examples of peripheral bus interfaces or display interfaces include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI®) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be embodied in a variety of electronic devices such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0378] In some implementations, video coding The method may be performed using an apparatus implemented on a hardware platform such as those described with respect to Figure 24A or Figure 24B.

[0379] FIG. 25 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0380] 25, video coding system 100 may include a source device 110 and a destination device 120. Source device 110 generates encoded video data and may be referred to as a video coding device. Destination device 120 decodes the encoded video data generated by source device 110 and may be referred to as a video decoding device.

[0381] The source device 110 may include a video source 112, a video encoder 114, and an input / output interface .

[0382] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers and / or computer graphics systems for generating video data, or a combination of such sources. The video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. En Coe Dosa The encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. The encoded video data may also be stored on storage medium / server 130b for access by destination device 120.

[0383] The destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0384] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to interact with an external display device.

[0385] Video encoder 114 and video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.

[0386] FIG. 26 is a block diagram illustrating an example of a video encoder 200, which may be video encoder 114 in system 100 shown in FIG.

[0387] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 26, video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0388] The functional components of the video encoder 200 include a partitioning unit 201, a prediction unit 202, which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding and unit 214.

[0389] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0390] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but are depicted separately in the example of FIG. 26 for illustrative purposes.

[0391] Partition unit 201 may divide a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support a variety of video block sizes.

[0392] The mode select unit 203 selects one of the coding modes (intra or inter) based on, for example, an error result, and provides the resulting intra- or inter-coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct an encoded block for use as a reference picture. In some examples, the mode select unit 203 may select a combination of intra- and inter-prediction (CIIP) modes, in which prediction is based on an inter-prediction signal and an intra-prediction signal. The mode select unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer-pixel precision) in the case of inter-prediction.

[0393] To perform inter prediction on the current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may identify a prediction video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0394] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice, for example.

[0395] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating a reference picture in list 0 or list 1, including the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predictive video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0396] In other examples, motion estimation unit 204 may perform bidirectional prediction for the current video block, where motion estimation unit 204 may search reference pictures in list 0 for a reference video block for the current video block and search reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 204 may then generate reference indexes indicating the reference pictures in lists 0 and 1, each including a reference video block and a motion vector indicating a spatial displacement between the reference video block and the current video block. Motion estimation unit 204 may output the reference index and the motion vector for the current video block as motion information for the current video block. Motion compensation unit 205 may generate a predictive video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0397] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.

[0398] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Rather, motion estimation unit 204 may signal the motion information of the current video block with reference to motion information of other video blocks. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0399] In one example, motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0400] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0401] As mentioned above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0402] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate predictive data for the current video block based on decoded samples of other video blocks within the same picture. The predictive data for the current video block may include a predicted video block and various syntax elements.

[0403] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0404] In other examples, for example in skip mode, residual data for the current video block may not exist and residual generation unit 207 may not perform the subtraction operation.

[0405] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0406] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0407] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block related to the current block for storage in buffer 213.

[0408] After reconstruction unit 212 reconstructs the video blocks, a loop filter operation may be performed to reduce video blocking artifacts in the video blocks.

[0409] Entropy encoding Unit 214 may receive data from other functional components of video encoder 200. coding When unit 214 receives the data, it calculates the entropy coding The unit 214 includes one or more entropy coding Doing the work and entropy Encoded Generate data and entropy Encoded A bitstream containing the data may be output.

[0410] FIG. 27 is a block diagram illustrating an example of a video decoder 300, which may be video decoder 114 in system 100 shown in FIG.

[0411] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 27, video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0412] In the example of FIG. 27, the video decoder 300 Decoding 26. The video decoder 300 may include a decoding path that is generally the reverse of the encoding path described with respect to the video encoder 200 (e.g., FIG. 26).

[0413] Entropy Decoding Unit 301 is Encoded The bitstream can be read. Encoded Bitstream is entropy code encoded video data (e.g., Encoded block). Entropy DecodingUnit 301 is Entropy code Converted video data Decoding And entropy Decoded From the video data, the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, implementing AMVP and merge mode.

[0414] The motion compensation unit 302 generates motion-compensated blocks and may optionally perform interpolation based on a compensation filter. Identifiers for interpolation filters to be used with sub-pixel accuracy may be included in syntax elements.

[0415] The motion compensation unit 302 calculates the motion of the video block. encoding The motion compensation unit 302 may calculate interpolated values ​​for the sub-integer pixels of the reference block using an interpolation filter used by video encoder 200 during the prediction. The motion compensation unit 302 may determine the interpolation filter used by video encoder 200 according to received syntax information and use the interpolation filter to generate the predictive block.

[0416] The motion compensation unit 302 Encoded Frames and / or slices of a video sequence Encoding the size of the blocks used to Encoded Partition information describing how each macroblock of a picture in a video sequence is divided; Encoding Indicates the mode, each interface Encoded one or more reference frames (and reference frame lists) for the block, and Encoded Video sequence Decoding The suffix "syntax" may be used to determine other information for the suffix "syntax".

[0417] The intra prediction unit 303 may form a prediction block from spatially neighboring blocks, for example, using an intra prediction mode received in the bitstream. Decoding By Unit 301 Decoding The inverse transform unit 305 applies an inverse transform to the quantized video block coefficients.

[0418] The reconstruction unit 306 sums the residual block with the corresponding prediction block generated by the motion compensation unit 205 or the intra prediction unit 303 to obtain Decoded Optionally, to remove block artifacts, Decoded A deblocking filter may also be applied to filter the blocks. Decoded The video blocks are stored in a buffer 307, which provides reference blocks for subsequent motion compensation.

[0419] Some embodiments of the disclosed techniques include making a decision or determination to enable a video processing tool or mode. In one example, if a video processing tool or mode is enabled, an encoder uses or implements the tool or mode in processing blocks of video, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, conversion of blocks of video to a bitstream representation of video uses the video processing tool or mode if enabled based on the decision or determination. In another example, if a video processing tool or mode is enabled, a decoder processes the bitstream with the knowledge that the bitstream has been modified based on the video processing tool or mode. That is, conversion of a bitstream representation of video to blocks of video occurs using the video processing tool or mode that is enabled based on the decision or determination.

[0420] Some embodiments of the disclosed techniques include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder does not use the tool or mode in converting blocks of video into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, a decoder processes the bitstream with the knowledge that the bitstream has not been modified using the video processing tool or mode that was disabled based on the decision or determination.

[0421] In this application, the term "video processing" means video encoding ,video Decoding , video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation or vice versa. The bitstream representation of a current video block may correspond to bits that are either co-located or distributed at different locations within the bitstream, e.g., as defined by syntax. For example, a macroblock may be converted and code in terms of the error residual values ​​and using bits in the header and other fields in the bitstream. Encoding It can be done.

[0422] The first set of clauses below may be implemented in some embodiments.

[0423] 1. A method for video processing, comprising: during a conversion between a current video block and a bitstream representation of the current video block, deriving a motion vector offset based on a resolution of a reference picture associated with the current video block and a resolution of the current picture; and performing the conversion using the motion vector offset.

[0424] 2. The method of claim 1, wherein deriving the motion vector offset includes deriving a first motion vector offset that references a first reference picture, and deriving a second motion vector offset that references a second reference picture based on the first motion vector offset.

[0425] 3. The method of claim 1, further comprising performing a motion candidate list building process for the current video block based on reference picture resolutions associated with spatial, temporal or historical motion candidates.

[0426] 4. Whether and how to update the lookup table Decoded 10. The method of claim 1, wherein the motion candidate is determined by a reference picture resolution associated with the motion candidate.

[0427] 5. The method of claim 1, further comprising performing a filtering operation on the current picture using adaptive loop filter (ALF) parameters associated with the corresponding dimension.

[0428] 6. The method of claim 5, wherein the ALF parameters include a first ALF parameter associated with a first corresponding dimension and a second ALF parameter associated with a second corresponding dimension, the second ALF parameter being inherited or predicted from the first ALF parameter.

[0429] 7. The method of claim 1, further comprising reshaping samples in the current picture using LMCS (Luma Mapping Chroma Scaling) parameters associated with corresponding dimensions.

[0430] 8. The method of claim 7, wherein the LMCS parameters include a first LMCS parameter associated with a first corresponding dimension and a second LMCS parameter associated with a second corresponding dimension, the second LMCS parameter being inherited or predicted from the first LMCS parameter.

[0431] 9. The method of clause 5 or 7, wherein the ALF or LMCS parameters signaled in the video unit are related to one or more picture dimensions.

[0432] 10. If the current video block references at least one reference picture that has different dimensions than the current picture, coding 10. The method of claim 1, further comprising disabling the tool.

[0433] 11. The method of clause 10, further comprising skipping or omitting merge candidates that refer to reference pictures having dimensions different from the dimensions of the current picture.

[0434] 12. After scaling two reference blocks or two reference pictures according to the resolution of the reference picture and the resolution of the current picture, or after scaling two MVs or MVDs (motion vector differences) according to the resolution of the reference picture and the resolution of the current picture, coding 10. The method of claim 1, further comprising applying a tool.

[0435] 13. The method according to claim 1, wherein the current picture includes no more than K different resolutions, where K is a natural number.

[0436] 14. The method of clause 13, wherein the K different resolutions are allowed for N consecutive pictures, where N is a natural number.

[0437] 15. The method of claim 1, further comprising applying a resolution change to a current picture that is an intra picture.

[0438] 16. The method of claim 1, further comprising converting bidirectional prediction to unidirectional prediction when one or two reference pictures of the current video block have a different resolution than the resolution of the current picture.

[0439] 17. The method of claim 1, further comprising enabling or disabling inter-prediction from reference pictures of different resolutions depending on at least one of motion vector accuracy or resolution between the current block dimension and the reference block dimension.

[0440] 18. The method of claim 1, further comprising applying bidirectional prediction depending on whether both or one of the reference pictures has a different resolution than the current picture.

[0441] 19. The method of clause 1, wherein whether the current video block references a reference picture having different dimensions than the current picture depends on at least one of a size of the current video block or a block prediction mode.

[0442] 20. The method of claim 1, wherein performing the conversion includes generating the current video block from the bitstream representation.

[0443] 21. The method of claim 1, wherein performing the transformation includes generating the bitstream representation from the current video block.

[0444] 22. An apparatus in a video system comprising a processor and non-transitory memory having instructions, which, when executed by the processor, cause the processor to perform a method according to any one of clauses 1 to 21.

[0445] 23. A computer program product stored on a non-transitory computer readable medium, comprising program code for performing the method of any one of clauses 1 to 21.

[0446] The second set of sections describes certain features and aspects of the techniques disclosed in the previous section.

[0447] 1. A method for video processing (e.g., method 2800 shown in FIG. 28), comprising: code For conversion between the quantized representation, code used to represent video regions in a representation coding determining 2810 an enabled state of a tool and performing the transformation according to the determination; coding The method includes a first flag included in the picture header to indicate the enabled state of the tool.

[0448] 2. The validity state is determined by the first flag and / or the slice type of the video. coding 2. The method of claim 1, wherein the is invalid or valid for the video region.

[0449] 3. The determining step is performed when the first flag is true. coding 10. The method of claim 1, wherein the tool is determined to be invalid.

[0450] 4. The determining step includes determining whether the first flag is false. coding 10. The method of claim 1, wherein the tool is determined to be invalid.

[0451] 5. The determining step includes determining whether the first flag is false. coding 10. The method of claim 1, wherein the tool is determined to be valid.

[0452] 6. The determining step determines whether the first flag is true. coding 10. The method of claim 1, wherein the tool is determined to be valid.

[0453] 7. The foregoing coding 2. The method according to claim 1, wherein a tool is disabled or enabled for all samples in the picture.

[0454] 8. The method according to claim 1, wherein the signaling of the first flag is determined by one or more syntax elements in a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS) or a picture parameter set (PPS) associated with the video domain.

[0455] 9. The signaling of the first flag is coding 9. The method of claim 8, which is determined by an indication of the tool's validity.

[0456] 10. The method according to clause 8, wherein a second flag indicating the presence of the first flag in the picture header is signaled in the sequence parameter set (SPS).

[0457] 11. A second flag indicating the presence of the first flag in the picture header coding 9. The method of clause 8, wherein a signal is generated when a tool is enabled for the sequence of video.

[0458] 12. The method according to clause 8, wherein if the second flag indicating the presence of the first flag indicates the presence of the first flag, the first flag is signaled in the picture header.

[0459] 13. The first flag and / or the second flag are 1 bit. code 13. The method according to any one of claims 1 to 12, wherein the

[0460] 14. The foregoing coding 14. The method according to any of clauses 1 to 13, wherein the tool is Prediction Refinement Using Optical Flow (PROF), in which one or more initial affine predictions are refined based on optical flow calculations.

[0461] 15. The above coding14. The method according to any of clauses 1 to 13, wherein the tool is a decoder-side motion vector refinement (DMVR) tool, in which motion information is refined using a predictive block.

[0462] 16. The above coding 14. The method according to any of clauses 1 to 13, wherein the tool is a bidirectional optical flow (BDOF) tool, where one or more initial predictions are refined using optical flow calculations.

[0463] 17. The above coding 14. The method of any of clauses 1 to 13, wherein the tool is a cross-component adaptive loop filter (CCALF) in which a linear filter is applied to chroma samples based on luma information.

[0464] 18. The above coding 14. A method according to any one of claims 1 to 13, wherein the tool is a geometrical partitioning (GEO) in which predicted samples are generated using weighted values, the predicted samples being based on a partitioning of the video region according to non-horizontal or non-vertical lines.

[0465] 19. The above coding A method according to any one of clauses 1 to 13, wherein the tool is a triangular prediction mode (TPM) in which prediction samples are generated using weighted values, and the prediction samples are based on dividing the video area into two triangular partitions.

[0466] 20. The conversion converts the video to a bitstream representation. encoding 20. The method of any one of claims 1 to 19, comprising:

[0467] 21. The transforming step transforms the bitstream representation to generate the video. Decoding 20. The method of any one of claims 1 to 19, comprising:

[0468] 22. A video processing device including a processor configured to implement the methods of one or more of clauses 1 to 21.

[0469] 23. A computer readable medium storing program code that, when executed, causes a processor to perform the method of one or more of clauses 1 to 21.

[0470] 24. Produced according to any of the methods described above code A computer-readable medium storing a quantized or bitstream representation of the image.

[0471] From the foregoing, it will be understood that, although specific embodiments of the disclosed technology have been described herein for purposes of illustration, various modifications can be made without departing from the scope of the disclosed technology. Accordingly, the disclosed technology is not limited except as by the appended claims.

[0472] Implementations of the subject matter and functional operations described herein can be embodied in various systems, digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or one or more combinations thereof. Implementations of the subject matter described herein can be embodied as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by or to control the operation of a data processing apparatus. A computer-readable medium can be a machine-readable storage device, a machine-readable storage carrier, a memory device, a composition of matter bearing a machine-readable propagated signal, or one or more combinations thereof. The term "data processing unit" or "data processing apparatus" encompasses all apparatus, devices, and machines that process data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.

[0473] A computer program (also known as a program, software, software application, script, or code) can be written in any programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subprograms, or portions of code), or in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communications network.

[0474] The processes and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as a Field Programmable Gate Array (FPGA) or an Application-Specific Integrated Circuit (ASIC).

[0475] Processors suitable for executing a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices, e.g., magnetic, optical, or magnetic disks, for storing data, or be operatively coupled to receive data from and / or transfer data to such one or more mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all types of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and memory may be enhanced by, or incorporated in, dedicated logic circuitry.

[0476] This specification, together with the drawings, are intended to be considered exemplary only. By exemplary, I mean example. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless expressly stated otherwise. Additionally, the use of "or" is intended to include "and / or" unless expressly stated otherwise.

[0477] While this specification contains numerous details, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as operating in a particular combination, and even initially claimed as such, one or more features from a claimed combination may in some cases be deleted from that combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0478] Similarly, although operations are depicted in the figures in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the depicted operations be performed, to achieve desired results. Further, the separation of various system components in the embodiments described herein should not be understood as requiring such separation in all embodiments.

[0479] Only a few implementations and examples have been described; other implementations, enhancements, and variations may be made based on what is described and illustrated herein.

Claims

1. 1. A method for processing video data, comprising: determining whether a coding tool is disabled for a picture based on a first syntax element for converting between a video domain of a picture of a video and a bitstream of the video; performing the conversion based on the determination; and Including, the first syntax element indicating whether the coding tool is disabled for the picture is signaled in a picture header; A method, wherein the coding tool is disabled or enabled for all samples in the picture.

2. The method of claim 1 , wherein whether the coding tool is disabled or enabled for the sub-picture video unit depends on the first syntax element and / or a slice type of the video.

3. The method of claim 1 or 2, wherein if the first syntax element is true, the coding tool is disabled for the picture.

4. The method of claim 1 , wherein if the first syntax element is false, the coding tool is valid for the picture.

5. 5. The method of claim 1, wherein whether the first syntax element is signaled in the picture header depends on one or more syntax elements in a sequence parameter set (SPS) associated with the video domain.

6. 6. The method of claim 5, wherein the one or more syntax elements include a second syntax element that indicates the presence of the first syntax element in the picture header and a third syntax element that indicates whether the coding tool is valid for the sequence of video.

7. The method of claim 6 , wherein the second syntax element is conditionally signaled in the SPS based on the third syntax element.

8. 8. The method of claim 6 or 7, wherein the first syntax element is signaled in the picture header if the second syntax element is true, and the second syntax element is signaled within the SPS if the coding tool is valid for the sequence of video.

9. 9. The method according to claim 6, wherein the first syntax element and / or the second syntax element is coded with one bit.

10. The method of claim 1 , wherein the coding tools include prediction refinement using optical flow (PROF).

11. The method of claim 1 , wherein the coding tools include decoder-side motion vector refinement (DMVR).

12. The method of claim 1 , wherein the coding tools include bidirectional optical flow (BDOF).

13. The method of claim 1 , wherein the converting comprises encoding the video into the bitstream.

14. The method of claim 1 , wherein the converting comprises decoding the video from the bitstream.

15. 1. An apparatus for processing video data, comprising: a processor; and non-transitory memory having instructions stored thereon, the instructions, when executed by the processor, causing the processor to: determining whether a coding tool is disabled for a picture based on a first syntax element for converting between a video domain of a picture of a video and a bitstream of the video; performing the conversion based on the determination; and Let them do this, the first syntax element indicating whether the coding tool is disabled for the picture is signaled in a picture header; The apparatus, wherein the coding tool is disabled or enabled for all samples in the picture.

16. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing a processor to: determining whether a coding tool is disabled for a picture based on a first syntax element for converting between a video domain of a picture of a video and a bitstream of the video; performing the conversion based on the determination; and Let them do this, the first syntax element indicating whether the coding tool is disabled for the picture is signaled in a picture header; The coding tool is disabled or enabled for all samples in the picture.

17. A method for storing a bitstream of video, the method comprising: determining, for a video region of a picture of the video, whether a coding tool is disabled for the picture based on a first syntax element; generating a bitstream of the video based on the determination; and storing the bitstream on a non-transitory computer-readable storage medium; Including, the first syntax element indicating whether the coding tool is disabled for the picture is signaled in a picture header; A method, wherein the coding tool is disabled or enabled for all samples in the picture.