Handling Decoder-Side Motion Vector Refinement (DMVR) Coding Tools for Reference Picture Resampling in Video Coding

By enabling selective DMVR based on resolution matching, the method optimizes resource usage and enhances user experience in video coding by addressing inefficiencies due to differing spatial resolutions.

JP2026053450APending Publication Date: 2026-03-25HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies when dealing with differing spatial resolutions between current and reference pictures, leading to increased resource usage and reduced user experience.

Method used

The implementation of decoder-side motion vector refinement (DMVR) that allows selective enabling and disabling based on resolution matching between current and reference pictures, optimizing resource usage and improving coding efficiency.

Benefits of technology

This approach reduces processor, memory, and network resource consumption, enhancing the user experience by improving coding efficiency and maintaining image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053450000001_ABST
    Figure 2026053450000001_ABST
Patent Text Reader

Abstract

This invention provides a decoding method using decoder-side motion vector refinement (DMVR). [Solution] The decoding method includes the steps of: determining by a video decoder whether the resolution of the current picture to be decoded is the same as the resolution of a reference picture identified by a reference picture list associated with the current picture; enabling decoder-side motion vector refinement (DMVR) by the video decoder for the current block of the current picture if it is determined that they are the same; disabling DMVR by the video decoder for the current block of the current picture if it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; and refining the motion vector corresponding to the current block by the video decoder using DMVR if the DMVR flag is enabled for the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Cross-Reference to Related Applications This patent application claims priority to U.S. Provisional Patent Application No. 62 / 848,410, filed May 15, 2019, by Jianle Chen, et al., entitled "Handling of Decoder-Side Motion Vector Refinement (DMVR) Coding Tools for Reference Picture Resampling in Video Coding", which is incorporated herein by reference in its entirety.

[0002] Technical Field Generally, the present disclosure describes techniques for supporting decoder-side motion vector refinement (DMVR) in video coding. More specifically, the present disclosure enables DMVR for reference picture resampling, but allows DMVR to be disabled for blocks or samples when the spatial resolution of the current and reference pictures is different.

[0003] Background The amount of video data required to depict even relatively short videos can be quite large, which can pose difficulties when the data is streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Also, since memory resources can be limited, the size of the video can be a problem when the video is stored on a storage device. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much in the way of image quality are desirable due to limited network resources and the ever-increasing requirements for higher video quality. [Overview of the Initiative]

[0004] The first aspect relates to a method for decoding a coded video bitstream performed by a video decoder. The method includes the steps of: determining by the video decoder whether the resolution of the current picture to be decoded is the same as the resolution of the reference pictures identified by a list of reference pictures associated with the current picture; enabling decoder-side motion vector refinement (DMVR) for the current block of the current picture if it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures; disabling DMVR for the current block of the current picture if it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; and, if DMVR is enabled for the current block, using DMVR to refine the motion vector corresponding to the current block by the video decoder.

[0005] The method provides a technique that allows for the selective disabling of DMVR instead of requiring the DMVR to be disabled for the entire CVS when Reference Picture Resampling (RPR) is enabled if the spatial resolution of the current picture differs from that of the reference picture. This ability to selectively disable DMVR improves coding efficiency. Consequently, the use of processor, memory, and / or network resources can be reduced in both the encoder and decoder. Therefore, the coder / decoder (also referred to as "codec") in video coding is improved compared to current codecs. In practical terms, the improved video coding process provides a better user experience when video is transmitted, received, and / or viewed.

[0006] Optionally, in any of the above embodiments, another implementation embodiment provides: the step of enabling the DMVR includes the step of setting the DMVR flag to a first value, and the step of disabling the DMVR includes the step of setting the DMVR flag to a second value.

[0007] Optionally, in any of the above embodiments, another implementation provides a step of generating a reference picture for the current picture based on a reference picture list according to a bidirectional interpretation mode.

[0008] Optionally, in any of the above embodiments, another implementation provides a step of selectively enabling and disabling DMVR for blocks within multiple pictures, depending on whether the resolution of each picture differs from or is the same as the resolution of the reference picture associated with the picture.

[0009] Optionally, in any of the above embodiments, another implementation provides a step of enabling reference picture resampling (RPR) for the entire coded video sequence (CVS) containing the current picture when DMVR is disabled.

[0010] Optionally, in any of the above embodiments, another implementation provides the following: the resolution of the current picture is placed in the parameter set of the coded video bitstream, and the current block is taken from a slice of the current picture.

[0011] Optionally, in any of the above embodiments, another implementation provides the step of displaying the image generated using the current block on the display of an electronic device.

[0012] A second aspect relates to a method for encoding a video bitstream performed by a video encoder. The method includes the steps of: determining by the video encoder whether the resolution of the current picture to be encoded is the same as the resolution of the reference pictures identified in a list of reference pictures associated with the current picture; enabling decoder-side motion vector refinement (DMVR) for the current block of the current picture if it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures; disabling DMVR for the current block of the current picture if it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; and refining the motion vector corresponding to the current block using DMVR if DMVR is enabled for the current block.

[0013] The method provides a technique that allows for the selective disabling of DMVR instead of requiring the DMVR to be disabled for the entire CVS when Reference Picture Resampling (RPR) is enabled if the spatial resolution of the current picture differs from that of the reference picture. This ability to selectively disable DMVR improves coding efficiency. Consequently, the use of processor, memory, and / or network resources can be reduced in both the encoder and decoder. Therefore, the coder / decoder (also referred to as "codec") in video coding is improved compared to current codecs. In practical terms, the improved video coding process provides a better user experience when video is transmitted, received, and / or viewed.

[0014] Optionally, in any of the above embodiments, another implementation provides the steps of: determining a motion vector for the current picture based on a reference picture using a video encoder; encoding the current picture based on the motion vector using a video encoder; and decoding the current picture using a virtual reference decoder using a video encoder.

[0015] Optionally, in any of the above embodiments, another implementation embodiment provides: the step of enabling the DMVR includes the step of setting the DMVR flag to a first value, and the step of disabling the DMVR includes the step of setting the DMVR flag to a second value.

[0016] Optionally, in any of the above embodiments, another implementation provides a step of generating a reference picture for the current picture based on a reference picture list according to a bidirectional interpretation mode.

[0017] Optionally, in any of the above embodiments, another implementation provides a step of selectively enabling and disabling DMVR for blocks within multiple pictures, depending on whether the resolution of each picture differs from or is the same as the resolution of the reference picture associated with the picture.

[0018] Optionally, in any of the above embodiments, another implementation provides a step of enabling reference picture resampling (RPR) for the entire coded video sequence (CVS) containing the current picture, even when DMVR is disabled.

[0019] Optionally, in any of the above embodiments, another implementation provides a step of sending the video bitstream containing the current block to a video decoder.

[0020] A third aspect relates to a decoding device. The decoding device includes a receiver configured to receive a coded video bitstream; a memory coupled to the receiver that stores instructions; and a processor coupled to the memory, the processor configured to cause the decoding device to: determine whether the resolution of the current picture to be decoded is the same as the resolution of the reference pictures identified by a list of reference pictures associated with the current picture; enable decoder-side motion vector refinement (DMVR) for the current block of the current picture if it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures; disable DMVR for the current block of the current picture if it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; and, if DMVR is enabled for the current block, use DMVR to refine the motion vector corresponding to the current block.

[0021] The decoding device provides a technique that allows the DMVR to be selectively disabled, instead of requiring the DMVR to be disabled for the entire CVS when Reference Picture Resampling (RPR) is enabled if the spatial resolution of the current picture differs from that of the reference picture. This ability to selectively disable the DMVR improves coding efficiency. Consequently, the use of processor, memory, and / or network resources can be reduced in both the encoder and decoder. Therefore, the coder / decoder (also referred to as "codec") in video coding is improved compared to current codecs. In practical terms, the improved video coding process provides a better user experience when video is transmitted, received, and / or viewed.

[0022] Optionally, in any of the above embodiments, another implementation provides the following: when DMVR is disabled, reference picture resampling (RPR) is enabled for the entire coded video sequence (CVS) containing the current picture.

[0023] Optionally, in any of the above embodiments, another implementation provides a display configured to show an image such that it is generated based on the current block.

[0024] A fourth aspect relates to an encoding device. The encoding device includes a memory containing instructions; a processor coupled to the memory, the processor being configured to execute instructions to the encoding device: the step of determining whether the resolution of the current picture to be encoded is the same as the resolution of the reference pictures identified in a reference picture list associated with the current picture; the step of enabling decoder-side motion vector refinement (DMVR) for the current block of the current picture if it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures; the step of disabling DMVR for the current block of the current picture if it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; the step of using DMVR to refine the motion vector corresponding to the current block if DMVR is enabled for the current block; and the encoding device includes a transmitter coupled to the processor, the transmitter being configured to transmit a video bitstream containing the current video block to a video decoder.

[0025] When the spatial resolution of the current picture is different from that of the reference picture, instead of disabling DMVR for the entire CVS when reference picture resampling (RPR) is enabled, a technique is provided that allows DMVR to be selectively disabled. By having the ability to selectively disable DMVR in this way, the coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Accordingly, the coder / decoder (also referred to as "codec") in video coding is improved compared to the current codec. In practice, the improved video coding process provides a better user experience to the user when the video is transmitted, received, and / or viewed.

[0026] Optionally, in any of the above aspects, another implementation provides the following: Even when DMVR is disabled, reference picture resampling (RPR) is enabled for the entire coded video sequence (CVS) including the current picture.

[0027] Optionally, in any of the above aspects, another implementation provides the following: The memory stores the video bitstream before the transmitter sends the bitstream to the video decoder.

[0028] Aspect 5 is related to a coding device. The coding device includes a receiver configured to receive a picture to be coded or a bitstream to be decoded; a transmitter coupled to the receiver and configured to transmit the bitstream to a decoder or the decoded image to a display; a memory coupled to at least one of the receiver or the transmitter and configured to store instructions; and a processor coupled to the memory and configured to execute the instructions stored in the memory to perform any of the methods disclosed herein.

[0029] The coding device provides a technique that allows DMVR to be selectively disabled instead of having to disable DMVR for the entire CVS when reference picture resampling (RPR) is enabled if the spatial resolution of the current picture is different from the spatial resolution of the reference picture. By having the ability to selectively disable DMVR in this manner, coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Accordingly, the coder / decoder (also referred to as "codec") in video coding is improved compared to the current codec. In practice, the improved video coding process provides a better user experience to the user when the video is transmitted, received, and / or viewed.

[0030] Aspect 6 is related to a system. The system includes an encoder; and a decoder in communication with the encoder, wherein the encoder or the decoder includes a decoding device, an encoding device, or a coding device disclosed herein.

[0031] The system provides a technique that allows for the selective disabling of DMVR instead of requiring the DMVR to be disabled for the entire CVS when Reference Picture Resampling (RPR) is enabled, if the spatial resolution of the current picture differs from that of the reference picture. This ability to selectively disable DMVR improves coding efficiency. Consequently, the use of processor, memory, and / or network resources can be reduced in both the encoder and decoder. Therefore, the coder / decoder (also referred to as "codec") in video coding is improved compared to current codecs. In practical terms, the improved video coding process provides a better user experience when video is transmitted, received, and / or viewed. [Brief explanation of the drawing]

[0032] For a more complete understanding of this disclosure, refer to the following brief descriptions relating to the attached drawings and detailed description. Here, similar reference numbers represent similar parts.

[0033] [Figure 1] This block diagram shows an exemplary coding system that can utilize video coding technology.

[0034] [Figure 2] This block diagram shows an exemplary video encoder capable of implementing video coding technology.

[0035] [Figure 3] This block diagram shows an example of a video decoder capable of implementing video coding technology.

[0036] [Figure 4]This represents the relationship between the IRAP picture and the trailing picture relative to the leading picture, expressed in terms of the decoding order and presentation order.

[0037] [Figure 5] This example demonstrates multi-layer coding for spatial scalability.

[0038] [Figure 6] This is a schematic diagram showing an example of one-way interpretation.

[0039] [Figure 7] This is a schematic diagram showing an example of bidirectional interface prediction.

[0040] [Figure 8] Shows the video bitstream.

[0041] [Figure 9] This demonstrates picture partitioning techniques.

[0042] [Figure 10] This is an embodiment of a method for decoding a coded video bitstream.

[0043] [Figure 11] This is an embodiment of a method for encoding a coded video bitstream.

[0044] [Figure 12] This is a schematic diagram of a video coding device.

[0045] [Figure 13] This is a schematic diagram of an embodiment of the coding means. [Modes for carrying out the invention]

[0046] While exemplary embodiments of one or more embodiments are provided below, it should be understood from the outset that the disclosed systems and / or methods may be carried out using any number of techniques, whether currently known or existing. This disclosure is not limited in any way to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations illustrated and described herein, and can be modified within the scope of the appended claims, along with the entire range of their equivalents.

[0047] As used in this context, resolution refers to the number of pixels in a video file. That is, resolution is the width and height of the projected image, measured in pixels. For example, a video may have a resolution of 1280 (horizontal pixels) × 720 (vertical pixels). This is usually simply written as 1280×720 or abbreviated as 720p. DMVR is a process, algorithm, or coding tool used to refine motion or motion vectors relative to predicted blocks. DMVR allows motion vectors to be discovered based on two motion vectors discovered for bi-prediction, using a bilateral template matching process. DMVR makes it possible to discover weighted combinations of predicted coding units generated by each of the two motion vectors, and the two motion vectors can be refined by replacing them with a new motion vector that optimally points to the combined predicted coding unit. RPR functionality is the ability to change the spatial resolution of a coded picture in the middle of a bitstream without requiring intra-coding of the picture at the resolution change location.

[0048] Figure 1 is a block diagram illustrating an exemplary coding system 10 capable of utilizing the video coding technology described herein. As shown in Figure 1, the coding system 10 includes a source device 12 that provides encoded video data to be subsequently decoded by a destination device 14. In particular, the source device 12 can provide the video data to the destination device 14 via a computer-readable medium 16. The source device 12 and destination device 14 may include any of a wide range of devices, including desktop computers, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, the source device 12 and destination device 14 may be equipped for wireless communication.

[0049] The destination device 14 can receive encoded video data to be decoded via a computer-readable medium 16. The computer-readable medium 16 may include any type of medium or device capable of transferring the encoded video data from the source device 12 to the destination device 14. In one example, the computer-readable medium 16 may include a communication medium that enables the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data can be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may constitute part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may help facilitate communication from the source device 12 to the destination device 14.

[0050] In some examples, the encoded data may be output to a storage device via the output interface 22. Similarly, the encoded data may be accessed from the storage device via the input interface. The storage device may include any variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, digital multipurpose discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data. In another example, the storage device may correspond to a file server or another intermediate storage device capable of storing the encoded video generated by the source device 12. The destination device 14 can access the stored video data from the storage device by streaming or downloading. The file server may be any type of server capable of storing the encoded video data and sending the encoded video data to the destination device 14. Specific examples of file servers include web servers (e.g., websites), File Transfer Protocol (FTP) servers, network-attached storage (NAS) devices, or local disk drives. Destination device 14 can access the encoded video data via any standard data connection, including an internet connection. This may include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber lines (DSL), cable modems, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. Transmission of the encoded video data from the storage device may be via streaming, download, or a combination thereof.

[0051] The technology of this disclosure is not necessarily limited to wireless applications or settings. The technology can be applied to video coding when supporting any variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the coding system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video phone.

[0052] In the example in Figure 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to this disclosure, the video encoder 20 of the source device 12 and / or the video decoder 30 of the destination device 14 can be configured to apply techniques for video coding. In other examples, the source device and destination device may include other components or arrangements. For example, the source device 12 may receive video data from an external video source such as an external camera. Similarly, the destination device 14 may interface with an external display device rather than including an integrated display device.

[0053] The coding system 10 shown in Figure 1 is merely an example. The technique for video coding may be performed by any digital video encoding and / or decoding device. While the technique of this disclosure is generally performed by a video coding device, the technique may also be performed by a video encoder / decoder, typically called a “CODEC”. Furthermore, the technique of this disclosure may also be performed by a video processor. The video encoder and / or decoder may be a graphics processing unit (GPU) or a similar device.

[0054] Source device 12 and destination device 14 are merely examples of coding devices, such that source device 12 generates encoded video data for transmission to destination device 14. In some examples, source device 12 and destination device 14 can operate in a substantially symmetrical manner, such that each of source device 12 and destination device 14 includes video encoding and decoding components. Thus, coding system 10 can support one-way or two-way video transmission between video devices 12, 14 for, for example, video streaming, video playback, video broadcasting, or video phone calls.

[0055] The video source 18 of source device 12 may include a video capture device such as a video camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 18 can generate computer graphics-based data as a source video, or a combination of live video, video stored in an archive, and computer-generated video.

[0056] In some cases, when the video source 18 is a video camera, the source device 12 and destination device 14 may form a so-called camera phone or video phone. However, as stated above, the technology described herein is generally applicable to video coding and is applicable to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video information can then be output to a computer-readable medium 16 via the output interface 22.

[0057] The computer-readable medium 16 may include temporary media such as wireless broadcast or wired network transmission, or storage media (i.e., non-temporary storage media) such as hard disks, flash drives, compact discs, digital video discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) can receive encoded video data from a source device 12 and provide the encoded video data to a destination device 14, for example, via network transmission. Similarly, a computing device in a media manufacturing apparatus, such as a disc stamping machine, can receive encoded video data from a source device 12 and produce a disc containing the encoded video data. Thus, the computer-readable medium 16 can be understood to include one or more computer-readable media of various forms in various examples.

[0058] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 is syntax information defined by the video encoder 20, which is also used by the video decoder 30, and may include syntax elements that describe the features and / or processing of blocks and / or other coded units, such as groups of pictures (GOPs). The display device 32 displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or other types of display devices.

[0059] The video encoder 20 and video decoder 30 can operate in accordance with video coding standards such as the High Efficiency Video Coding (HEVC) standard currently under development and can conform to the HEVC Test Model (HM). Alternatively, the video encoder 20 and video decoder 30 can operate in accordance with other proprietary or industrial standards such as the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) H.264 standard, or alternatively, what is referred to as the Video Expert Group (MPEG)-4, Part 10, Advanced Video Coding (AVC), H.265 / HEVC, or extensions of such standards. However, the technology of this disclosure is not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although not shown in Figure 1, in some embodiments, the video encoder 20 and video decoder 30 may be integrated with an audio encoder and decoder, respectively, and may include a suitable multiplexer-demultiplexer (MUX-DEMUX) unit or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams. Where applicable, the MUX-DEMUX unit may conform to the ITU H.223 Multiplexer Protocol or other protocols such as the User Datagram Protocol (UDP).

[0060] The video encoder 20 and video decoder 30 may each be implemented as any variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If the technology is partially implemented in software, the device may store instructions for the software on a suitable non-temporary computer-readable medium and execute the instructions on hardware using one or more processors to perform the technology of this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated within their respective devices as part of a combined encoder / decoder (CODEC). A device including the video encoder 20 and / or video decoder 30 may include an integrated circuit, a microprocessor, and / or a wireless communication device such as a cellular telephone.

[0061] Figure 2 is a block diagram showing an example of a video encoder 20 capable of implementing video coding techniques. The video encoder 20 can perform intra and intercoding of video blocks within a video slice. Intra coding relies on spatial prediction to reduce or eliminate spatial redundancy in video within a given video frame or picture. Intercoding relies on temporal prediction to reduce or eliminate temporal redundancy in video within adjacent frames or pictures in a video sequence. Intra mode (I mode) can refer to any of several space-based coding modes. Inter modes, such as one-way prediction (also known as bi-prediction) (P mode) or two-way prediction (also known as bi-prediction) (B mode), can refer to any of several time-based coding modes.

[0062] As shown in Figure 2, the video encoder 20 receives the current video block in the video frame to be encoded. In the example in Figure 2, the video encoder 20 includes a mode selection unit 40, a reference frame memory 64, an adder 50, a transformation unit 52, a quantization unit 54, and an entropy coding unit 56. The mode selection unit 40 then includes a motion compensation unit 44, a motion estimation unit 42, an intra prediction unit 46, and a partition unit 48. For video block reconstruction, the video encoder 20 also includes an inverse quantization unit 58, an inverse transformation unit 60, and an adder 62. A deblocking filter (not shown in Figure 2) may also be included to filter block boundaries and remove blocking artifacts from the reconstructed video. If desired, the deblocking filter typically filters the output of the adder 62. Additional filters (in-loop or post-loop) may also be used in addition to the deblocking filter. Such filters are not shown in the diagram for simplicity, but if desired, they can filter the output of adder 50 (as in-loop filters).

[0063] During the encoding process, the video encoder 20 receives video frames or slices to be coded. The frames or slices may be divided into multiple video blocks. The motion estimation unit 42 and the motion compensation unit 44 perform inter-predictive coding of the received video blocks for one or more blocks within one or more reference frames to provide temporal predictions. Alternatively, the intra-predictive unit 46 may perform intra-predictive coding of the received video blocks in relation to one or more adjacent blocks within the same frame or slice as the block to be coded to provide spatial predictions. The video encoder 20 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0064] Furthermore, the partition unit 48 can partition blocks of video data into sub-blocks based on an evaluation of the previous partitioning scheme in a previous coding pass. For example, the partition unit 48 first partitions a frame or slice into the largest coding unit (LCU), and then partitions each LCU into sub-coding units based on rate-distortion analysis (e.g., rate-distortion optimization). The mode selection unit 40 can further generate a quadtree data structure indicating that the LCUs are partitioned into sub-CUs. A leaf node CU of the quadtree can include one or more prediction units (PUs) and one or more transformation units (TUs).

[0065] This disclosure uses the term “block” to refer to any of CUs, PUs, or TUs in the context of HEVC, or similar data structures in the context of other standards (e.g., macroblocks and their subblocks in H.264 / AVC). A CU includes a coding node, a PU, and a TU associated with the coding node. The size of a CU corresponds to the size of the coding node and is square in shape. The size of a CU can range from 8x8 pixels to the size of a tree block of up to 64x64 pixels or more. Each CU may contain one or more PUs and one or more TUs. Syntax data associated with a CU may, for example, describe partitioning the CU into one or more PUs. The partitioning mode may differ depending on whether the CU is skipped or whether it is direct-mode coding, intra-predictive-mode coding, or inter-predictive-mode coding. A PU may be partitioned to be non-square in shape. Syntax data associated with a CU may also, for example, describe partitioning the CU into one or more TUs following a quadtree. The TU can be square or non-square (e.g., rectangular).

[0066] The mode selection unit 40 selects one of the inter- or intra-coding modes, for example, based on the error result, and provides the resulting intra- or inter-coded block to the adder 50 to generate residual block data and to the adder 62 to reconstruct the coded block as a reference frame. The mode selection unit 40 also provides syntax elements such as motion vectors, intra-mode indicators, partition information, and other such syntactic information to the entropy coding unit 56.

[0067] The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is the process of generating motion vectors that estimate the motion of a video block. The motion vectors can, for example, show the displacement of the PU of a video block in the current video frame or picture relative to a predicted block in a reference frame relative to the current block (or other coded unit) being coded in the current frame. The predicted block is a block that has been found to closely match the block to be coded, in terms of pixel differences that can be determined by absolute difference (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, the video encoder 20 can calculate values ​​for sub-integer pixel positions of a reference picture stored in the reference frame memory 64. For example, the video encoder 20 can interpolate values ​​for 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation unit 42 can perform motion searches for full pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.

[0068] The motion estimation unit 42 calculates motion vectors for video blocks in an intercoded slice relative to the PU by comparing the PU's position with the predicted block's position in a reference picture. The reference picture can be selected from either the first reference picture list (List 0) or the second reference picture list (List 1), each of which identifies one or more reference pictures stored in the reference frame memory 64. The motion estimation unit 42 sends the calculated motion vectors to the entropy coding unit 56 and the motion compensation unit 44.

[0069] Motion compensation performed by the motion compensation unit 44 may include fetching or generating a predicted block based on a motion vector determined by the motion estimation unit 42. In some examples, the motion estimation unit 42 and the motion compensation unit 44 may also be functionally integrated. Upon receiving the motion vector for the current video block's PU, the motion compensation unit 44 can locate the position of the predicted block pointed to by the motion vector in one of the reference picture lists. The adder 50 forms a residual video block and a pixel difference value by subtracting the pixel value of the predicted block from the pixel value of the current video block being coded, as described later. Generally, the motion estimation unit 42 performs motion estimation for the luminous component, and the motion compensation unit 44 uses a motion vector calculated based on the luminous component for both the chroma and luminous components. The mode selection unit 40 may also generate syntax elements related to the video block and video slice for use by the video decoder 30 when decoding the video block of the video slice.

[0070] The intra-prediction unit 46 can intra-predict the current block as an alternative to the inter-prediction performed by the motion estimation unit 42 and motion compensation unit 44 as described above. In particular, the intra-prediction unit 46 can determine the intra-prediction mode to use to encode the current block. In some examples, the intra-prediction unit 46 can encode the current block using various intra-prediction modes, for example, between separate encoding passes, and the intra-prediction unit 46 (or, in some examples, the mode selection unit 40) can select an appropriate intra-prediction mode from the tested modes.

[0071] For example, the intra-prediction unit 46 can calculate rate distortion values ​​for various tested intra-prediction modes using rate distortion analysis and select the intra-prediction mode with the optimal rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block encoded to generate the encoded block, as well as the bit rate (i.e., number of bits) used to generate the encoded block. The intra-prediction unit 46 can calculate ratios from the distortion and rate for various encoded blocks and determine which intra-prediction mode exhibits the best rate distortion value for the block.

[0072] Furthermore, the intra-prediction unit 46 may be configured to code depth blocks of the depth map using a depth modeling mode (DMM). The mode selection unit 40 can determine, for example, using rate-distortion optimization (RDO), whether an available DMM mode yields better coding results than the intra-prediction mode and other DMM modes. The texture image data corresponding to the depth map can be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also be configured to inter-predict depth blocks of the depth map.

[0073] After selecting an intra-prediction mode for a block (for example, a conventional intra-prediction mode or one of several DMM modes), the intra-prediction unit 46 can provide the entropy coding unit 56 with information indicating the selected intra-prediction mode for the block. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode. The video encoder 20 can include in the transmitted bitstream configuration data, which may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also called codeword mapping tables), definitions of coding contexts for various blocks, instructions for the most probable intra-prediction mode, intra-prediction mode index tables, and modified intra-prediction mode index tables to be used for each context.

[0074] The video encoder 20 forms a residual video block by subtracting the predicted data from the mode selection unit 40 from the original video block to be coded. The adder 50 represents one or more components that perform this subtraction operation.

[0075] The transformation processing unit 52 applies a transformation such as a discrete cosine transform (DCT) or a conceptually similar transformation to the residual block to generate a video block containing residual transformation coefficient values. The transformation processing unit 52 may also perform other transformations conceptually similar to the DCT. Wavelet transforms, integer transforms, subband transforms, or other types of transformations may also be used.

[0076] The transformation processing unit 52 applies the transformation to the residual block to generate a block of residual transformation coefficients. The transformation can convert residual information from the pixel value domain to a transformation domain such as the frequency domain. The transformation processing unit 52 can send the resulting transformation coefficients to the quantization unit 54. The quantization unit 54 quantizes the transformation coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized transformation coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0077] After quantization, the entropy coding unit 56 entropy codes the quantized transformation coefficients. For example, the entropy coding unit 56 can perform context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), stochastic interval partitioning entropy (PIPE) coding, or other entropy coding techniques. In the case of context-based entropy coding, the context can be based on adjacent blocks. Following entropy coding by the entropy coding unit 56, the encoded bitstream can be transmitted to another device (e.g., video decoder 30) or stored in an archive for later transmission or retrieval.

[0078] The inverse quantization unit 58 and the inverse transform unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual blocks in the pixel domain, for example, to be used later as reference blocks. The motion compensation unit 44 can compute a reference block by adding the residual blocks to the prediction blocks of one of the frames in the reference frame memory 64. The motion compensation unit 44 can also compute sub-integer pixel values ​​for use in motion estimation by applying one or more interpolation filters to the reconstructed residual blocks. The adder 62 adds the reconstructed residual blocks to the motion-compensated prediction blocks generated by the motion compensation unit 44, generating a reconstructed video block for storage in the reference frame memory 64. The reconstructed video block can be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for intercoding blocks in subsequent video frames.

[0079] Figure 3 is a block diagram showing an example of a video decoder 30 capable of implementing video coding technology. In the example in Figure 3, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra-prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. In some examples, the video decoder 30 can perform a decoding path that is roughly the inverse of the coding path described with respect to the video encoder 20 (Figure 2). The motion compensation unit 72 can generate prediction data based on the motion vector received from the entropy decoding unit 70, while the intra-prediction unit 74 can generate prediction data based on an intra-prediction mode indicator received from the entropy decoding unit 70.

[0080] During the decoding process, the video decoder 30 receives an encoded video bitstream from the video encoder 20, representing the video blocks of the decoded video slice and their associated syntax elements. The entropy decoding unit 70 of the video decoder 30 decodes the bitstream and generates quantized coefficients, motion vectors or intra-predictive mode indicators, and other syntax elements. The entropy decoding unit 70 transfers the motion vectors and other syntax elements to the motion compensation unit 72. The video decoder 30 can receive syntax elements at the video slice level and / or video block level.

[0081] When a video slice is coded as an intra-coded (I) slice, the intra-prediction unit 74 can generate prediction data for the video blocks of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current frame or picture. When a video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, the motion compensation unit 72 generates prediction blocks for the video blocks of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 70. Prediction blocks can be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 can construct the reference frame lists, List 0 and List 1, using default construction techniques based on the reference pictures stored in the reference frame memory 82.

[0082] The motion compensation unit 72 determines the prediction information for the video blocks of the current video slice by analyzing the motion vectors and other syntax elements, and uses this prediction information to generate the prediction blocks of the current video block to be decoded. For example, the motion compensation unit 72 uses some of the received syntax elements to determine the prediction mode used to code the video blocks of the video slice, the interprediction slice type (e.g., B slice, P slice, or GPB slice), one or more configuration pieces of the slice's reference picture list, the motion vector for each intercoded video block of the slice, the interprediction status for each intercoded video block of the slice, and other information for decoding the video blocks in the current video slice.

[0083] The motion compensation unit 72 can also perform interpolation based on an interpolation filter. The motion compensation unit 72 can calculate interpolated values ​​for sub-integer pixels of a reference block using an interpolation filter, such as the one used by the video encoder 20 during the encoding of the video block. In this case, the motion compensation unit 72 can determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to generate the predicted block.

[0084] The texture image data corresponding to the depth map can be stored in the reference frame memory 82. The motion compensation unit 72 can also be configured to interpret depth blocks of the depth map.

[0085] In some embodiments, the video decoder 30 includes a user interface 84. The user interface 84 is configured to receive input from a user of the video decoder 30 (e.g., a network administrator). Through the user interface 84, the user can manage or change settings in the video decoder 30. For example, the user can input or otherwise provide values ​​for parameters (e.g., flags) to control the configuration and / or operation of the video decoder 30 according to the user's preferences. The user interface 84 may be a graphical user interface (GUI) that allows the user to interact with the video decoder 30, for example, through graphical icons, drop-down menus, check boxes, etc. In some cases, the user interface 84 can receive information from the user via a keyboard, mouse, or other peripheral device. In some embodiments, the user can access the user interface 84 via a smartphone, tablet device, personal computer located remotely from the video decoder 30, etc. As used herein, the user interface 84 may be referred to as an external input or external means.

[0086] With the above in mind, video compression techniques perform spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate redundancy inherent in video sequences. In block-based video coding, a video slice (i.e., a video picture or a portion of a video picture) can be partitioned into video blocks, which may be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks in an inter-coded (P or B) slice of a picture can use spatial prediction with respect to reference samples in adjacent blocks within the same picture, or use temporal prediction with respect to reference samples in other reference pictures. Pictures may be referred to as frames, and reference pictures may be referred to as reference frames.

[0087] Spatial or temporal predictions become predicted blocks of the blocks to be coded. Residual data represents the pixel difference between the original blocks to be coded and the predicted blocks. Intercoded blocks are coded according to motion vectors pointing to the reference sample blocks that form the predicted blocks, and the residual data shows the difference between the coded blocks and the predicted blocks. Intracoded blocks are coded according to the intracoded mode and residual data. For further compression, the residual data is transformed from the pixel domain to the transformation domain, resulting in residual transformation coefficients that can be subsequently quantized. The quantized transformation coefficients are initially arranged in a two-dimensional array and can be scanned to generate a one-dimensional vector of transformation coefficients, and entropy coding can be applied to achieve even greater compression.

[0088] Image and video compression is rapidly developing, moving towards various coding standards. Such video coding standards include Advanced Video Coding (AVC), also known as ITU-T H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC).

[0089] There is also a new video coding standard called Versatile Video Coding (VVC), developed by the ITU-T and ISO / IEC Joint Video Expert Team (JVET). The VVC standard includes several working drafts, but in particular, one working draft of VVC, B. Bross, J. Chen, and S. Liu, “Versatile Video Coding (Draft 5),” JVET-N1001-v3, 13th JVET Meeting, March 27, 2019 (VVC Draft 5), is incorporated herein by reference in its entirety.

[0090] The technical descriptions disclosed herein are based on the Multipurpose Video Coding (VVC) standard, which is under development by the ITU-T and ISO / IEC Joint Video Expert Team (JVET). However, the technology is also applicable to other video codec standards.

[0091] Figure 4 shows the relationship between the intra-random access point (IRAP) picture 402 and the trailing picture 406 with respect to the leading picture 404, represented by the decoding order 408 and the presentation order 410. In embodiments, the IRAP picture 402 is referred to as a clean random access (CRA) picture or as an instantaneous decoder refresh (IDR) picture with a random access decodeable (RADL) picture. In HEVC, IDR pictures, CRA pictures, and broken link access (BLA) pictures are all considered IRAP pictures 402. In VVC, at the 12th JVET meeting in October 2018, it was agreed that both IDR and CRA pictures should be considered as IRAP pictures. In embodiments, broken link access (BLA) and stepwise decoder refresh (GDR) pictures can also be considered IRAP pictures. The decoding process for coded video sequences always begins with IRAP.

[0092] As shown in Figure 4, the leading picture 404 (e.g., pictures 2 and 3) follows the IRAP picture 402 in the decoding order 408, but precedes the IRAP picture 402 in the presentation order 410. The trailing picture 406 follows the IRAP picture 402 in both the decoding order 408 and the presentation order 410. Although two leading pictures 404 and one trailing picture 406 are shown in Figure 4, those skilled in the art will recognize that more or fewer leading pictures 404 and / or trailing pictures 406 may be present in the decoding order 408 and trailing order 410 in actual applications.

[0093] The reading picture 404 in Figure 4 is divided into two types: Random Access Skip Reading (RASL) and RADL. If decoding begins with IRAP picture 402 (e.g., picture 1), the RADL picture (e.g., picture 3) can be properly decoded; however, the RASL picture (e.g., picture 2) cannot be properly decoded. Therefore, the RASL picture is discarded. According to the distinction between RADL and RASL pictures, the type of reading picture 404 associated with IRAP picture 402 should be identified as either RADL or RASL for efficient and proper coding. HEVC has a constraint that, if both RASL and RADL pictures exist, with respect to RASL and RADL pictures associated with the same IRAP picture 402, the RASL picture shall precede the RADL picture in presentation order 410.

[0094] IRAP picture 402 provides the following two important functions / benefits. First, the presence of IRAP picture 402 indicates that the decoding process can start from that picture. This function enables a random access function where, as long as IRAP picture 402 is present in that position, the decoding process starts from that position in the bitstream and is not necessarily the beginning of the bitstream. Second, the presence of IRAP picture 402 refreshes the decoding process, and as a result, coded pictures that begin with IRAP picture 402, excluding RASL picture, are coded without referencing any preceding pictures. The presence of IRAP picture 402 in the bitstream therefore prevents any errors that may have occurred during the decoding of coded pictures preceding IRAP picture 402 from propagating to IRAP picture 402 and the pictures that follow IRAP picture 402 in decoding order 408.

[0095] IRAP picture 402 provides important functionality but incurs a penalty to compression efficiency. The presence of IRAP picture 402 causes a surge in bitrate. This penalty to compression efficiency stems from two reasons. Firstly, because IRAP picture 402 is an intra-predicted picture, the picture itself requires relatively more bits for representation compared to other pictures that are inter-predicted pictures (e.g., leading picture 404, trailing picture 406). Secondly, the presence of IRAP picture 402 interrupts the temporal prediction (because the decoder refreshes the decoding process, one of the actions of the decoding process for this is to remove the preceding reference picture in the buffer (DPB) of the decoded picture), and IRAP picture 402 makes coding of the picture that follows IRAP picture 402 in the decoding order 408 inefficient (i.e., requires more bits to represent) because there are fewer reference pictures for their inter-predicted coding.

[0096] Among the picture types considered to be IRAP picture 402, the IDR picture in HEVC has different signaling and derivation compared to other picture types. Some of the differences are as follows:

[0097] In signaling and deriving the Picture Order Count (POC) value of an IDR picture, the most significant bit (MSB) portion of the POC is not derived from the preceding key picture, but is simply set to 0.

[0098] Regarding the signaling information required for reference picture management, the slice header of an IDR picture does not contain the information necessary to be signaled to assist in reference picture management. For other picture types (i.e., CRA, trailing, Temporal Sublayer Access (TSA), etc.), information such as the Reference Picture Set (RPS) or other forms of similar information (e.g., a Reference Picture List) described below is required for the reference picture marking process (i.e., the process of determining the state of reference pictures in the decoded picture buffer (DPB), whether they are used for reference or not). However, in the case of IDR pictures, such information does not need to be signaled because the presence of the IDR indicates that the decoding process will simply mark all reference pictures in the DPB as not to be used for reference.

[0099] In HEVC and VVC, the IRAP picture 402 and the leading picture 404 may each be contained within a single Network Abstraction Layer (NAL) unit. A set of NAL units is sometimes called an access unit. The IRAP picture 402 and the leading picture 404 are given different NAL unit types, and as a result, they can be easily identified by system-level applications. For example, a video splicer needs to understand the coded picture types without having to understand the excessive details of the syntax elements in the coded bitstream, and in particular needs to identify the IRAP picture 402 from non-IRAP pictures and the leading picture 404 from the trailing picture 406, which includes determining the RASL and RADL pictures. The trailing picture 406 is a picture associated with the IRAP picture 402 and follows the IRAP picture 402 in presentation order 410. A picture may follow a specific IRAP picture 402 in decoding order 408 and precede any other IRAP picture 402 in decoding order 408. For this reason, under IRAP picture 402 and reading picture 404, their own NAL unit types are useful in such applications.

[0100] In HEVC, the NAL unit types for IRAP pictures include the following: BLA with Reading Picture (BLA_W_LP): A NAL unit of Broken Link Access (BLA) that may be followed by one or more reading pictures in the decryption order. BLA with RADL (BLA_W_RADL): This is the NAL unit of the BLA picture, and may be followed by one or more RADL pictures in the decoding order, but not RASL pictures. BLA without a reading picture (BLA_N_LP): This is a NAL unit of a BLA picture, and the reading picture does not follow it in the decoding order. DLA with RADL (IDR_W_RADL): This is the NAL unit of the IDR picture, and may be followed by one or more RADL pictures in the decoding order, but not RASL pictures. IDR without a reading picture (IDR_N_LP): This is a NAL unit of the IDR picture, and the reading picture does not follow it in the decoding order. CRA: A Clean Random Access (CRA) picture NAL unit, which may be followed by a leading picture (i.e., a RASL picture, a RADL picture, or both). RADL: NAL unit of RADL picture. RASL: NAL unit for RASL pictures.

[0101] In VVC, the NAL unit types for IRAP picture 402 and reading picture 404 are as follows: IDR with RADL (IDR_W_RADL): This is the NAL unit of the IDR picture, and may be followed by one or more RADL pictures in the decoding order, but not RASL pictures. IDR without a reading picture (IDR_N_LP): This is a NAL unit of the IDR picture, and the reading picture does not follow it in the decoding order. CRA: A Clean Random Access (CRA) picture NAL unit, which may be followed by a leading picture (i.e., a RASL picture, a RADL picture, or both). RADL: NAL unit of RADL picture. RASL: NAL unit for RASL pictures.

[0102] Reference Picture Resampling (RPR) is the ability to change the spatial resolution of a coded picture midway through a bitstream without requiring intra-coding of the picture at the resolution change point. To enable this, the picture needs to be able to reference one or more reference pictures whose spatial resolution differs from that of the current picture, for interpretation purposes. Therefore, resampling of such reference pictures, or parts thereof, is required for the encoding and decoding of the current picture. Hence the name RPR. This feature may also be referred to as Adaptive Resolution Change (ARC) or by other names. There are use cases or application scenarios that benefit from the RPR feature, including the following:

[0103] Rate adaptation in video conferencing and video conferencing. This involves adapting coded video to changing network conditions. When network conditions deteriorate and as a result the available bandwidth decreases, the encoder can adapt by encoding a picture at a lower resolution.

[0104] Changing the Active Speaker in Multi-Party Video Conferencing. In multi-party video conferences, the video size for the active speaker is typically larger or higher than the video size for the remaining conference participants. When the active speaker changes, it may be necessary to adjust the picture resolution for each participant as well. The need for ARC (Automatic Speaker Response) functionality becomes more important when the active speaker changes frequently.

[0105] Faster starts in streaming. In streaming applications, it's common for the application to buffer a predetermined length of the decoded picture before starting to display it. Starting the bitstream at a lower resolution allows the application to have enough picture buffered to begin displaying faster.

[0106] Adaptive Stream Switching in Streaming. The Dynamic Adaptive Streaming (DASH) standard in HTTP includes a feature called @mediaStreamStructureId. This feature enables switching between different representations at open GOP random access points that have undecodeable reading pictures, such as a CRA picture with an associated RASL picture in HEVC. If two different representations of the same video have the same spatial resolution but different bitrates, and they have the same @mediaStreamStructureId value, it is possible to perform switching between the two representations in a CRA picture associated with a RASL picture, and the RASL picture associated with the switching in the CRA picture can be decoded with acceptable quality, thus enabling seamless switching. In ARC, the @mediaStreamStructureId feature can also be used for switching between DASH representations with different spatial resolutions.

[0107] Various methods facilitate fundamental techniques to support RPR / ARC, such as listing picture resolutions and signaling some constraints on resampling of reference pictures in the DPB. Furthermore, at the 14th JVET meeting in Geneva, there were several input contributions suggesting constraints that should be applied to VVCs to support RPR. The proposed constraints include:

[0108] Some tools will disable the coding of blocks within the current picture if they reference a reference picture with a different resolution than the current picture. These tools include:

[0109] Temporal motion vector prediction (TMVP) and advanced TMVP (ATMVP). This was proposed by JVET-N0118.

[0110] Decoder-side motion vector refinement (DMVR). This was proposed by JVET-N0279.

[0111] Bidirectional optical flow (BIO). This was proposed by JVET-N0279.

[0112] Bi-prediction of blocks from a reference picture with a different resolution than the current picture is not allowed. This was proposed by JVET-N0118.

[0113] In the case of motion compensation, sample filtering is applied only once; that is, if resampling and interpolation are required to obtain a finer Pell resolution (e.g., 1 / 4 Pell resolution), the two filters need to be applied in combination only once. This was proposed by JVET-N0118.

[0114] Scalability in video coding is typically supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise (SNR) scalability, and multi-view scalability. When multi-layer coding techniques are used, a picture or part thereof can be coded (1) without using a reference picture, i.e., without using intra-prediction, (2) by referencing a reference picture in the same layer, i.e., by using inter-prediction, or (3) by referencing a reference picture in another layer, i.e., by using inter-layer prediction. A reference picture used for inter-layer prediction of the current picture is called an inter-layer reference picture (ILRP).

[0115] Figure 5 shows an example of multi-layer coding for spatial scalability 500. Picture 502 in layer N has a different resolution (e.g., a lower resolution) than picture 504 in layer N+1. In this embodiment, as described above, layer N is considered to be the base layer and layer N+1 is considered to be the enhancement layer. Picture 502 in layer N and picture 504 in layer N+1 may be coded using inter-prediction (as indicated by the solid arrows). Picture 502 may also be coded using inter-layer prediction (as indicated by the dashed arrows).

[0116] In an RPR scenario, the reference picture can be resampled by selecting a reference picture from a lower layer or by using inter-layer prediction to generate a higher layer reference picture based on a lower layer reference picture.

[0117] The earlier H.26x video coding family provides support for scalability through separate profiles within the profile for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of AVC / H.264 that provides support for spatial, temporal, and qualitative scalability. In SVC, each macroblock (MB) in an EL picture is signaled with a flag that specifies whether the EL MB is predicted using collocated blocks from lower layers. Predictions from collocated blocks may include textures, motion vectors, and / or coding modes. Implementations of SVC cannot directly reuse unmodified H.264 / AVC implementations in their designs. The SVC EL macroblock syntax and decoding process differ from the H.264 / AVC syntax and decoding process.

[0118] Scalable HEVC (SHVC) is an extension of the HEVC / H.265 standard that supports spatial and qualitative scalability, Multiview HEVC (MV-HEVC) is an extension of HEVC / H.265 that supports multiview scalability, and 3D HEVC (3D-HEVC) is an extension of HEVC / H.264 that supports three-dimensional (3D) video coding and is more advanced and efficient than MV-HEVC. Note that temporal scalability is included as an integral part of the single-layer HEVC codec. The design of the multi-layer extensions of HEVC uses the idea that decoded pictures used for inter-layer prediction come from only the same access unit (AU), are treated as long-term reference pictures (LTRPs), and are given a reference index in the reference picture list along with other temporal reference pictures of the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the reference index value to reference the inter-layer reference picture in the reference picture list.

[0119] In particular, both reference picture resampling and spatial scalability features require resampling of the reference picture or a portion thereof. Reference picture resampling can be implemented at either the picture level or the coding block level. However, when RPR is referred to as a coding feature, it is a feature for single-layer coding. Even so, it is possible, or even desirable, from a codec design perspective to use the same resampling filter for both the RPR feature of single-layer coding and the spatial scalability feature of multi-layer coding.

[0120] JVET-N0279 proposed disabling DMVR for RPR. More precisely, it proposed disabling the use of DMVR for the entire coded video sequence (CVS) when RPR is enabled. It has been found that even when the RPR function is enabled, the current picture does not, in many cases, reference a reference picture with a different resolution. Therefore, disabling DMVR for the entire CVS is unnecessarily restrictive and may impair coding efficiency.

[0121] Disclosed here is a technology that allows the DMVR to be selectively disabled, instead of requiring the DMVR to be disabled for the entire CVS when RPR is enabled if the spatial resolution of the current picture differs from the spatial resolution of the reference picture. By providing the ability to selectively disable the DMVR in this manner, coding efficiency can be improved. Consequently, the use of processor, memory, and / or network resources can be reduced in both the encoder and decoder. Thus, the coder / decoder (also referred to as “codec”) in video coding is improved compared to current codecs. In practical terms, the improved video coding process provides users with a better user experience when video is transmitted, received, and / or viewed.

[0122] Figure 6 is a schematic diagram showing an example of a one-way interpretation 600. The one-way interpretation 600 can be used to determine the motion vectors of the encoded and / or decoded blocks generated when partitioning a picture.

[0123] The one-way interpretation 600 uses a reference frame 630 containing a reference block 631 to predict the current block 611 within the current frame 610. The reference frame 630 may be placed temporally after the current frame 610 (e.g., as a subsequent reference frame), as illustrated, but in some examples it may be placed temporally before the current frame 610 (e.g., as a preceding reference frame). The current frame 610 is an exemplary frame / picture to be encoded / decoded at a particular time. The current frame 610 contains an object in the current block 611 that matches an object in the reference block 631 of the reference frame 630. The reference frame 630 is the frame used as a reference for encoding the current frame 610, and the reference block 631 is the block in the reference frame 630 that contains an object that is also included in the current block 611 of the current frame 610.

[0124] The current block 611 is any coding unit being coded / decoded at a particular point in the coding process. The current block 611 may be an entire partitioned block, or a subblock if an affine interpredictive mode is used. The current frame 610 is separated from the reference frame 630 by a time distance (TD) 633. TD 633 represents the amount of time between the current frame 610 and the reference frame 630 in the video sequence and can be measured in frames. Predictive information for the current block 611 can refer to the reference frame 630 and / or reference block 631 by a reference index that indicates the directional and temporal distance between frames. Over the period represented by TD 633, objects in the current block 611 move from their position in the current frame 610 to another position in the reference frame 630 (e.g., the position in reference block 631). For example, an object may move along a motion trajectory 613, which is the direction of the object's motion over time. The motion vector 635 describes the direction and magnitude of the object's motion along the motion trajectory 613 across TD 633. Thus, the encoded motion vector 635, the reference block 631, and the residual including the difference between the current block 611 and the reference block 631 provide sufficient information to reconstruct the current block 611 and position it within the current frame 610.

[0125] Figure 7 is a schematic diagram showing an example of a bidirectional interpretation 700. The bidirectional interpretation 700 can be used to determine the motion vectors of the encoded and / or decoded blocks generated when partitioning a picture.

[0126] The bidirectional interpretation 700 is similar to the unidirectional interpretation 600, but uses a pair of reference frames to predict the current block 711 within the current frame 710. Thus, the current frame 710 and the current block 711 are quite similar to the current frame 610 and the current block 611, respectively. The current frame 710 is temporally positioned between a preceding reference frame 720 that appears before the current frame 710 in the video sequence and a succeeding reference frame 730 that appears after the current frame 710 in the video sequence. The preceding reference frame 720 and the succeeding reference frame 730 are also quite similar to the reference frame 630 in other respects.

[0127] The current block 711 coincides with the preceding reference block 721 in the preceding reference frame 720 and the subsequent reference block 731 in the subsequent reference frame 730. Such a coincidence indicates that, in the course of the video sequence, an object moves from the position of the preceding reference block 721, through the current block 711 along the motion trajectory 713, to the position of the subsequent reference block 731. The current frame 710 is separated from the preceding reference frame 720 by a preceding time distance (TD0) 7233 and from the subsequent reference frame 730 by a subsequent time distance (TD1) 733. TD0 723 indicates the amount of time in frames between the preceding reference frame 720 and the current frame 710 in the video sequence. TD1 733 indicates the amount of time in frames between the current frame 710 and the subsequent reference frame 730 in the video sequence. Therefore, the object moves from the preceding reference block 721 along the motion trajectory 713 to the current block 711 over the period indicated by TD0 723. The object also moves from the current block 711 along the motion trajectory 713 to the succeeding reference block 731 over the period indicated by TD1 733. The prediction information for the current block 711 can refer to the preceding reference frame 720 and / or preceding reference block 721 and the succeeding reference frame 730 and / or succeeding reference block 731 by a pair of reference indices indicating the temporal distance and direction between frames.

[0128] The preceding motion vector (MV0) 725 describes the direction and magnitude of the object's motion along the motion trajectory 713 over TD0 723 (for example, between the preceding reference frame 720 and the current frame 710). The succeeding motion vector (MV1) 735 describes the direction and magnitude of the object's motion along the motion trajectory 713 over TD1 733 (for example, between the current frame 710 and the succeeding reference frame 730). Thus, in the bidirectional interpretation 700, the current block 711 can be coded and reconstructed using the preceding reference block 721 and / or succeeding reference block 731, MV0 725, and MV1 735.

[0129] In this embodiment, interpretation and / or biinterpretation can be performed sample by sample (e.g., pixel by pixel) rather than block by block. That is, motion vectors pointing to each sample in the preceding reference block 721 and / or succeeding reference block 731 can be determined for each sample in the current block 711. In such an embodiment, the motion vectors 725 and 735 shown in Figure 7 represent multiple motion vectors corresponding to multiple samples in the current block 711, the preceding reference block 721, and the succeeding reference block 731.

[0130] In both merge mode and advanced motion vector prediction (AMVP) mode, the candidate list is generated by adding candidate vectors to the candidate list in an order defined by the candidate list determination pattern. Such candidate motion vectors may include motion vectors generated by one-way interpretation 600, two-way interpretation 700, or a combination thereof. Specifically, when such a block is encoded, motion vectors are generated for adjacent blocks. Such motion vectors are added to the candidate list for the current block, and motion vectors for the current block are selected from the candidate list. The motion vectors can then be signaled as indices of the selected motion vectors in the candidate list. The decoder can construct the candidate list using the same process as the encoder and can determine the motion vector selected from the candidate list based on the signaled indices. Thus, the candidate motion vectors include motion vectors generated according to one-way interpretation 600 and / or two-way interpretation 700, depending on which approach is used when such adjacent blocks are encoded.

[0131] Figure 8 shows the video bitstream 800. As used herein, the video bitstream 800 may be referred to as a coded video bitstream, bitstream, or a variation thereof. As shown in Figure 8, the bitstream 800 includes a sequence parameter set 802, a picture parameter set 804, a slice header 806, and picture data 808.

[0132] SPS 802 contains data common to all pictures in a sequence of pictures (SOP). In contrast, PPS 804 contains data common to all pictures. The slice header 806 contains information about the current slice, such as the slice type and which reference picture is used. SPS 802 and PPS 804 may also be commonly referred to as parameter sets. SPS 802, PPS 804, and slice header 806 are types of Network Abstraction Layer (NAL) units. A NAL unit is a syntactic structure that contains an indication of the type of data it follows (e.g., coded video data). NAL units are classified into Video Coding Layer (VCL) and non-VCL NAL units. A VCL NAL unit contains data representing the values ​​of samples within a video picture, while a non-VCL NAL unit contains any relevant additional information, such as parameter sets (essential header data applicable to multiple VCL NAL units) and supplemental enhancement information (timing information and other supplemental data that may enhance the availability of the video signal being decoded, but are not essential for decoding the values ​​of samples within the video picture). Those skilled in the art will understand that a bitstream 800 may contain other parameters and information in a real-world application.

[0133] The image data 808 in Figure 8 includes data related to an image or video to be encoded or decoded. The image data 808 may also be referred to simply as the payload or data carried within the bitstream 800. In embodiments, the image data 808 includes a CVS 814 (or CLVS) containing a plurality of pictures 810. The CVS 814 is a coded video sequence for the entire coded layer video sequence (CLVS) within the video bitstream 800. Notably, if the video bitstream 800 contains a single layer, the CVS and CLVS are the same. The CVS and CLVS differ only when the video bitstream 800 contains multiple layers.

[0134] Each slice of picture 810 may be contained within its own VCL NAL unit 812. The set of VCL NAL units 812 within CVS 814 may be referred to as an access unit.

[0135] Figure 9 shows a partitioning technique 900 for picture 910. Picture 910 may be similar to any of the pictures 810 in Figure 8. As shown, picture 910 may be partitioned into multiple slices 912. A slice is a spatially distinct frame region (e.g., a picture) that is encoded separately from any other region within the same frame. Three slices 912 are shown in Figure 9, but more or fewer slices may be used in the actual application. Each slice 912 may be partitioned into multiple blocks 914. The blocks 914 in Figure 9 may be similar to the current block 711, preceding reference block 721, and succeeding reference block 731 in Figure 7. Blocks 914 may represent CUs. Four blocks 914 are shown in Figure 9, but more or fewer blocks may be used in the actual application.

[0136] Each block 914 may be partitioned into multiple samples 916 (e.g., pixels). In this embodiment, the size of each block 914 is measured in luma samples. Sixteen samples 916 are shown in Figure 9, but more or fewer samples may be used in actual applications.

[0137] Figure 10 shows an embodiment of method 1000 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). Method 1000 may be performed after the decoded bitstream has been received directly or indirectly from a video encoder (e.g., video encoder 20). Method 1000 improves the decoding process by allowing the DMVR to be selectively disabled, instead of requiring the DMVR to be disabled for the entire CVS when RPR is enabled if the spatial resolution of the current picture differs from the spatial resolution of the reference picture. By allowing the DMVR to be selectively disabled in this manner, coding efficiency can be improved. Thus, in practice, codec performance is improved, resulting in a better user experience.

[0138] In block 1002, the video decoder determines whether the resolution of the current picture to be decoded is the same as the resolution of a reference picture identified by a reference picture list. In an embodiment, the video decoder receives a coded video bitstream (e.g., bitstream 800). The coded video bitstream includes a reference picture list, indicates the resolution of the current picture, and indicates a bidirectional interpretation mode. In an embodiment, the reference picture list structure includes a reference picture list. In an embodiment, the reference picture list is used for bidirectional interpretation. In an embodiment, the resolution of the current picture is placed within a parameter set of the coded video bitstream. In an embodiment, the resolution of the reference picture is derived based on the current picture, inferred based on the resolution of the current picture, parsed from the bitstream, or otherwise obtained. In an embodiment, the reference picture of the current picture is generated based on the reference picture list according to the bidirectional interpretation mode.

[0139] In block 1004, the video decoder enables DMVR for the current block of the current picture if it determines that the resolution of the current picture is the same as the resolution of each of the reference pictures. In embodiments, the video decoder enables DMVR by setting the DMVR flag to a first value (e.g., true, 1). In embodiments, DMVR is an optional process even when DMVR is enabled; that is, DMVR does not need to be performed even when DMVR is enabled.

[0140] In block 1006, the video decoder disables the DMVR for the current block of the current picture if the resolution of the current picture is different from any of the resolutions of the reference picture. In an embodiment, the video decoder disables the DMVR by setting the DMVR flag to a second value (e.g., false, zero).

[0141] In block 1008, the video decoder refines the motion vector corresponding to the current block when the DMVR flag is set to a first value. In embodiments, method 1000 further includes the step of selectively enabling and disabling DMVR for other blocks in the current picture, depending on whether the resolution of the current picture is different from or the same as the resolution of the reference picture.

[0142] In embodiments, the method further includes the step of enabling reference picture resampling (RPR) for the entire coded video sequence (CVS) containing the current picture when the DMVR is disabled.

[0143] In one embodiment, the current block is obtained from a slice of the current picture. In another embodiment, the current picture includes multiple slices, and the current block is obtained from a slice among these multiple slices.

[0144] In this embodiment, the image generated based on the current picture is displayed to the user of an electronic device (e.g., a smartphone, tablet, laptop, personal computer, etc.).

[0145] Figure 11 shows an embodiment of method 900 for encoding a video bitstream implemented by a video encoder (e.g., video encoder 20). Method 900 may be performed when a picture (e.g., from video) is encoded into a video bitstream and sent to a video decoder (e.g., video decoder 30). Method 1100 improves the encoding process by allowing the DMVR to be selectively disabled, instead of requiring the DMVR to be disabled for the entire CVS when RPR is enabled if the spatial resolution of the current picture differs from the spatial resolution of the reference picture. By allowing the DMVR to be selectively disabled in this manner, coding efficiency can be improved. Thus, in practice, codec performance is improved, resulting in a better user experience.

[0146] In block 1102, the video encoder determines whether the resolution of the current picture to be encoded is the same as the resolution of a reference picture identified by the reference picture list. In an embodiment, the reference picture list structure includes a reference picture list. In an embodiment, the reference picture list is used for bidirectional interpretation. In an embodiment, the resolution of the current picture is encoded in the parameter set of the video bitstream. In an embodiment, the reference picture of the current picture is generated based on the reference picture list according to the bidirectional interpretation mode.

[0147] In block 1104, the video encoder enables DMVR for the current block of the current picture if it determines that the resolution of the current picture is the same as the resolution of each of the reference pictures. In embodiments, the video encoder enables DMVR by setting the DMVR flag to a first value (e.g., true, 1). In embodiments, DMVR is an optional process even when DMVR is enabled; that is, DMVR does not need to be performed even when DMVR is enabled.

[0148] In the embodiment, the method includes the steps of determining the motion vector of the current picture based on a reference picture, encoding the current picture based on the motion vector, and decoding the current picture using a virtual reference decoder (HRD).

[0149] In block 1106, the video encoder disables DMVR for the current block of the current picture if the resolution of the current picture is different from any of the resolutions of the reference picture. In an embodiment, the video encoder disables DMVR by setting the DMVR flag to a second value (e.g., false, zero).

[0150] In block 1108, the video encoder refines the motion vector corresponding to the current block when the DMVR flag is set to a first value. In embodiments, method 1100 further includes the step of selectively enabling and disabling DMVR for other blocks in the current picture, depending on whether the resolution of the current picture is different from or the same as the resolution of the reference picture.

[0151] In embodiments, the method further includes the step of enabling reference picture resampling (RPR) for the entire coded video sequence (CVS) containing the current picture, even when DMVR is disabled.

[0152] In one embodiment, the current block is obtained from a slice of the current picture. In another embodiment, the current picture includes multiple slices, and the current block is obtained from a slice among these multiple slices.

[0153] In one embodiment, the video encoder generates a video bitstream containing the current block and transmits the video bitstream to the video decoder. In another embodiment, the video encoder stores the video bitstream for transmission to the video decoder.

[0154] In an embodiment, a method for decoding a video bitstream is disclosed. The video bitstream includes at least one picture. Each picture includes a plurality of slices. Each slice of the plurality of slices includes a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists includes a plurality of reference pictures that can be used for interpretation of coding blocks in a slice.

[0155] The method includes the steps of: parsing a parameter set to obtain resolution information of the current picture; obtaining two reference picture lists for the current slice of the current picture; determining the reference picture to decode the current coding block of the current slice; determining the resolution of the reference picture; determining whether decoder-side motion vector refinement (DMVR) is used or enabled for decoding the current coding block based on the resolutions of the current picture and the reference picture; and decoding the current coding block.

[0156] In this embodiment, if the resolution of the current picture and the reference picture are different, DMVR is either not used or disabled for decoding the current child block.

[0157] In an embodiment, a method for decoding a video bitstream is disclosed. The video bitstream includes at least one picture. Each picture includes a plurality of slices. Each slice of the plurality of slices is associated with a header containing a plurality of syntax elements. Each slice of the plurality of slices includes a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists includes a plurality of reference pictures that may be used for interpretation of coding blocks in the current slice.

[0158] The method involves the steps of: parsing a parameter set to obtain a flag that indicates whether a decoder-side motion vector refinement (DMVR) coding tool / technology may be used to decode a picture in the currently coded video sequence; obtaining the current slice in the current picture; and The process includes the step of parsing the slice header associated with the current slice to obtain a flag that indicates whether a Decoder-Side Motion Vector Refinement (DMVR) coding tool / technology may be used to decode a coding block in the current slice, if the value of a flag that indicates whether a DMVR coding tool / technology may be used to decode a picture in the current coded video sequence indicates that DMVR may be used.

[0159] In the embodiment, the DMVR coding tool is not used for decoding the current coding block or is disabled if the value of a flag that specifies whether the DMVR coding tool may be used for decoding the coding block in the current slice indicates that the coding tool may not be used for decoding the current slice.

[0160] In the embodiment, the value of a flag that, if not present, specifies whether a DMVR coding tool may be used to decode a coding block in the current slice is assumed to be the same as the value of a flag that specifies whether a decoder-side motion vector refinement (DMVR) coding tool / technique may be used to decode a picture in the current coded video sequence.

[0161] In embodiments, a method for encoding a video bitstream is disclosed. The video bitstream includes at least one picture. Each picture includes a plurality of slices. Each slice of the plurality of slices is associated with a header containing a plurality of syntax elements. Each slice of the plurality of slices includes a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists consists of a plurality of reference pictures that may be used for interpretation of the current coding block.

[0162] The method includes the steps of: determining whether a decoder-side motion vector refinement (DMVR) coding tool / technology may be used to encode a picture in the currently coded video sequence; parsing a set of parameters to obtain resolution information for each picture bitstream; obtaining two reference picture lists for the current slice in the current picture; parsing the reference picture list for the current slice to obtain an active reference picture that may be used to decode the coding block of the current slice; and imposing a constraint that the DMVR coding tool may not be used to encode a coding block in the current slice if at least one of the following conditions is met: the DMVR coding tool may not be used to encode a picture in the currently coded video sequence; and the resolution of the current picture and at least one reference picture are different.

[0163] A method for decoding a video bitstream is disclosed. The video bitstream includes at least one picture. Each picture includes multiple slices. Each slice of the multiple slices is associated with a header containing multiple syntax elements. Each slice of the multiple slices includes multiple coding blocks and multiple reference picture lists. Each reference picture list of the multiple reference picture lists consists of multiple reference pictures that may be used for interpretation of coding blocks in the current slice.

[0164] The method includes the steps of: parsing a parameter set to obtain a flag indicating whether a decoder-side motion vector refinement (DMVR) coding tool / technology may be used to decode a picture in the currently coded video sequence; and parsing a parameter set to obtain a flag indicating whether a decoder-side motion vector refinement (DMVR) coding tool / technology may be used to decode a picture that references the parameter set, wherein the parameter set is a picture parameter set (PPS).

[0165] In the embodiment, if the value of a flag that specifies whether the DMVR coding tool may be used to decode a picture referencing a PPS indicates that the coding tool may not be used, the DMVR coding tool is not used or is disabled for decoding the current coding block.

[0166] In embodiments, a method for encoding a video bitstream is disclosed. The video bitstream includes at least one picture. Each picture includes a plurality of slices. Each slice of the plurality of slices is associated with a header containing a plurality of syntax elements. Each slice of the plurality of slices includes a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists consists of a plurality of reference pictures that may be used for interpretation of coding blocks in the current slice.

[0167] The method includes the steps of: determining whether a decoder-side motion vector refinement (DMVR) coding tool / technology may be used to encode a picture in the currently coded video sequence; determining whether a decoder-side motion vector refinement (DMVR) coding tool / technology may be used to encode a picture when referring to the current PPS; and imposing the constraint that, if the DMVR coding tool may not be used to encode a picture in the currently coded sequence, the DMVR coding tool may not be used to encode a picture when referring to the current PPS.

[0168] The following syntax and semantics may be used to implement the embodiments disclosed herein. The following description is relative to the base text, which is the latest VVC draft specification. In other words, only the differences are described, and text in the base text not mentioned below applies as appropriate. Text added to the base text is shown in bold, and text deleted is shown in italics.

[0169] The process for creating reference picture lists will be revised as follows:

[0170]

number

[0171] Derivation of a flag that determines whether DMVR is used.

[0172] The decoding process for coding units coded in interpredictive mode includes the following sequence of steps:

[0173] 1. The variable dmvrFlag is set to equal to 0.

[0174] 2. The motion vector components and reference indices of the current coding unit are derived as follows:

[0175]

number

[0176]

number

[0177] dmvrFlag is set to equal to 1 if all of the following conditions are true:

[0178]

number

[0179]

number

[0180]

number

[0181]

number

[0182]

number

[0183]

number

[0184]

number

[0185] cbWidth is 8 or greater.

[0186] cbHeight is 8 or greater.

[0187] cbHeight * cbWidth is 128 or greater.

[0188]

number

[0189]

number

[0190]

number

[0191]

number

[0192]

number

[0193]

number

[0194] Derivation of a flag that determines whether DMVR is used.

[0195] The decoding process for coding units coded in interpredictive mode includes the following sequence of steps:

[0196] 1. The variable dmvrFlag is set to equal to 0.

[0197] 2. The motion vector components and reference indices of the current coding unit are derived as follows:

[0198]

number

[0199]

number

[0200] dmvrFlag is set to equal to 1 if all of the following conditions are true:

[0201]

number

[0202]

number

[0203]

number

[0204]

number

[0205]

number

[0206]

number

[0207]

number

[0208]

number

[0209] cbWidth is 8 or greater.

[0210] cbHeight is 8 or greater.

[0211] cbHeight * cbWidth is 128 or greater.

[0212]

number

[0213]

number

[0214]

number

[0215]

number

[0216]

number

[0217] Derivation of a flag that determines whether DMVR is used.

[0218] The decoding process for coding units coded in interpredictive mode consists of the following sequential steps:

[0219] The variable dmvrFlag is set to equal to 0.

[0220] The motion vector components and reference indices of the current coding unit are derived as follows:

[0221]

number

[0222]

number

[0223] dmvrFlag is set to equal to 1 if all of the following conditions are true:

[0224]

number

[0225]

number

[0226]

number

[0227]

number

[0228]

number

[0229] [Number]

[0230] [Number]

[0231] [Number]

[0232] cbWidth is 8 or more.

[0233] cbHeight is 8 or more.

[0234] cbHeight * cbWidth is 128 or more.

[0235] FIG. 12 is a schematic diagram of a video coding device 1200 (e.g., video encoder 20 or video decoder 30) according to an embodiment of the present disclosure. The video coding device 1200 is suitable for implementing the embodiments disclosed as described herein. The video coding device 1200 includes an input port 1210 and a receiver unit (Rx) 1220 for receiving data, a processor, logic unit, or central processing unit (CPU) 1230 for processing data, a transmitter unit (Tx) 1240 and an output port 1250 for transmitting data, and a memory 1260 for storing data. Also, the video coding device 1200 may also include optical - electrical (OE) components and electrical - optical (EO) components coupled to the input port 1210, the receiver unit 1220, the transmitter unit 1240, and the output port 1250 with respect to the exits and entrances of optical or electrical signals.

[0236] The processor 1230 is implemented by hardware and software. The processor 1230 may be implemented as one or more CPU chips, cores (e.g., a multi-core processor), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1230 communicates with the inlet port 1210, the receiver unit 1220, the transmitter unit 1240, the exit port 1250, and the memory 1260. The processor 1230 includes a coding module 1270. The coding module 1270 implements the embodiments disclosed above. For example, the coding module 1270 implements, processes, prepares, or provides various codec functions. Thus, including the coding module 1270 results in a significant improvement to the functionality of the video coding device 1200 and results in changes to various states of the video coding device 1200. Alternatively, the coding module 1270 is implemented as instructions stored in memory 1260 and executed by processor 1230.

[0237] The video coding device 1200 may also include input and / or output (I / O) devices 1280 for communicating data to and from the user. The I / O devices 1280 may include output devices such as a display for displaying video data and a speaker for outputting audio data. The I / O devices 1280 may also include input devices such as a keyboard, mouse, or trackball, and / or corresponding interfaces for interacting with such output devices.

[0238] Memory 1260 includes one or more disks, tape drives, and solid-state drives, and may be used as an overflow data storage device, storing the program if such a program is selected for execution, and storing instructions and data read during the execution of the program. Memory 1260 may be volatile and / or non-volatile, and may be read-only memory (ROM), random-access memory (RAM), tri-associative memory (TCAM), and / or static random-access memory (SRAM).

[0239] Figure 13 is a schematic diagram of an embodiment of the coding means 1300. In this embodiment, the coding means 1300 is implemented by a video coding device 1302 (e.g., a video encoder 20 or a video decoder 30). The video coding device 1302 includes a receiving means 1301. The receiving means 1301 is configured to receive a picture to be coded or a bitstream to be decoded. The video coding device 1302 includes a transmitting means 1307 coupled to the receiving means 1301. The transmitting means 1307 is configured to transmit a bitstream to a decoder or a decoded picture to a display means (e.g., one of the I / O devices 1280).

[0240] The video coding device 1302 includes a storage means 1303. The storage means 1303 is coupled to at least one of the receiving means 1301 or the transmitting means 1307. The storage means 1303 is configured to store instructions. The video coding device 1302 also includes a processing means 1305. The processing means 1305 is coupled to the storage means 1303. The processing means 1305 is configured to execute instructions stored in the storage means 1303 in order to perform the method disclosed herein.

[0241] It should be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of the steps of such methods should be understood to be merely illustrative. Similarly, additional steps may be included in such methods, and some steps may be omitted or combined in a manner consistent with the various embodiments of the present disclosure.

[0242] While several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of this disclosure. The embodiments provided herein are illustrative and non-limiting, and their intent is not limited to the details given herein. For example, various elements or components may be combined or integrated into other systems, or certain features may be omitted or not implemented.

[0243] Furthermore, technologies, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items illustrated or described as being coupled to one another, directly coupled, or communicating with one another may be indirectly coupled or communicating through some interface, device, or intermediate component, whether by electrical, mechanical, or other means. Other examples of modifications, substitutions, and alternatives are evident to those skilled in the art and can be made without departing from the spirit and scope of the disclosure herein.

Claims

1. Decryption method: Steps include receiving a bitstream containing a parameter set; A step of analyzing the aforementioned parameter set to obtain the resolution of the current picture to be decoded; Steps to obtain the width and height of the current picture; A step of determining whether the current picture satisfies a set of conditions, the set of conditions including at least that the resolution of the current picture to be decoded is the same as the resolution of a reference picture identified by a reference picture list associated with the current picture, the width of the current picture is greater than 8, the height of the current picture is greater than 8, and the product of the width and the height is greater than 128; If the current picture satisfies the set of conditions, enable decoder-side motion vector refinement (DMVR) for the current block of the current picture; If the current picture does not satisfy the set of conditions, the step of disabling the DMVR for the current block of the current picture; and If the DMVR is enabled for the current block, the step is to use the DMVR to refine the motion vector corresponding to the current block; A method that includes this.

2. The method according to claim 1, wherein the step of enabling the DMVR includes setting the DMVR flag to a first value, and the step of disabling the DMVR includes setting the DMVR flag to a second value.

3. The method according to claim 1 or 2, further comprising the step of generating the reference picture for the current picture based on the reference picture list in accordance with a bidirectional interpretation mode.

4. The method according to any one of claims 1-3, further comprising the step of selectively enabling and disabling the DMVR for blocks in a plurality of pictures, depending on whether the resolution of each picture is different from or the same as the resolution of a reference picture associated with the picture.

5. The method according to any one of claims 1-4, further comprising the step of enabling reference picture resampling (RPR) for the entire coded video sequence (CVS) containing the current picture when the DMVR is disabled.

6. The method according to any one of claims 1 to 5, wherein the resolution of the current picture is located within a parameter set of a coded video bitstream, and the current block is obtained from a slice of the current picture.

7. The method according to any one of claims 1 to 6, further comprising the step of displaying an image generated using the current block on a display of an electronic device.

8. An encoding method: Steps to obtain the width and height of the current picture; A step of determining whether the current picture satisfies a set of conditions, the set of conditions including at least that the resolution of the current picture to be encoded is the same as the resolution of a reference picture identified in a reference picture list associated with the current picture, the width of the current picture is greater than 8, the height of the current picture is greater than 8, and the product of the width and the height is greater than 128; If the current picture satisfies the set of conditions, enable decoder-side motion vector refinement (DMVR) for the current block of the current picture; If the current picture does not satisfy the set of conditions, the step of disabling the DMVR for the current block of the current picture; If the DMVR is enabled for the current block, the step of using the DMVR to refine the motion vector corresponding to the current block; and A step of encoding the resolution of the current picture to be encoded into a set of bitstream parameters; A method that includes this.

9. The aforementioned method is: A video encoder determines the motion vector for the current picture based on the reference picture; A step of encoding the current picture using the video encoder based on the motion vector; and A step of decoding the current picture using the video encoder with a virtual reference decoder; The method according to claim 8, further comprising:

10. The method according to claim 8 or 9, wherein the step of enabling the DMVR includes setting the DMVR flag to a first value, and the step of disabling the DMVR includes setting the DMVR flag to a second value.

11. The method according to any one of claims 8-10, further comprising the step of generating the reference picture for the current picture based on the reference picture list in accordance with a bidirectional interprediction mode.

12. The method according to any one of claims 8-11, further comprising the step of selectively enabling and disabling the DMVR for blocks in a plurality of pictures, depending on whether the resolution of each picture is different from or the same as the resolution of a reference picture associated with the picture.

13. The method according to any one of claims 8-12, further comprising the step of enabling reference picture resampling (RPR) for the entire coded video sequence (CVS) containing the current picture, even if the DMVR is disabled.

14. The method according to any one of claims 8-13, further comprising the step of transmitting the video bitstream, including the current block, to a video decoder.

15. It is a decryption device: A receiver configured to receive a coded video bitstream; A memory connected to the receiver, which stores instructions; A processor coupled to the aforementioned memory; A decoding device comprising a processor configured to execute the instructions to cause the decoding device to perform the method according to any one of claims 1 to 7.

16. It is an encoding device: Memory containing instructions; A processor coupled to the memory, wherein the processor is configured to execute the instructions to cause the encoding device to perform the method according to any one of claims 8-14.

17. A method for transmitting a bitstream: A receive step of receiving a bitstream containing a plurality of parameter sets, wherein the plurality of parameter sets contain resolution information and a first flag, the first flag indicating whether decoder-side motion vector refinement (DMVR) based bidirectional interpretation is enabled; the first flag being equal to a first value indicates that the DMVR-based bidirectional interpretation is enabled, and the first flag being equal to a second value indicates that the DMVR-based bidirectional interpretation is disabled; the resolution information is used to determine whether the resolution of the current picture to be decoded is the same as the resolution of a reference picture identified in a reference picture list associated with the current picture, and to determine whether to enable the DMVR for the current block of the current picture if the first flag is equal to the first value; The bitstream is then decoded by the decoder: A step of determining whether the current picture satisfies a set of conditions, the set of conditions including at least that the resolution of the current picture to be decoded is the same as the resolution of a reference picture identified by a reference picture list associated with the current picture, the width of the current picture is greater than 8, the height of the current picture is greater than 8, and the product of the width and the height is greater than 128; If the current picture satisfies the set of conditions, the step of enabling DMVR for the current block of the current picture; If the current picture does not satisfy the set of conditions, the step of disabling the DMVR for the current block of the current picture; and If the DMVR is enabled for the current block, a receiving step is performed using the DMVR to refine the motion vector corresponding to the current block; The steps of storing the bitstream in a storage medium; and A step of transmitting the bitstream stored in the storage medium; A method that includes this.