Handling of bi-directional optical flow (BIO) coding tool for reference picture resampling in video coding

The method addresses the challenge of varying video resolutions by selectively disabling BDOF in video coding, thereby improving coding efficiency and reducing resource usage, leading to enhanced video codec performance and user experience.

JP2025087712AActive Publication Date: 2025-06-10HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025020501
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-05-15
Filing Date
2025-02-12
Publication Date
2025-06-10
Estimated Expiration
2040-05-14

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing video data, particularly when dealing with videos of varying spatial resolutions, which can lead to reduced coding efficiency and increased resource utilization.

Method used

The proposed method enables selective disabling of bidirectional optical flow (BDOF) for video blocks when the spatial resolution of the current picture differs from that of the reference pictures, allowing for improved coding efficiency by reducing processor, memory, and network resource usage.

Benefits of technology

By selectively disabling BDOF based on resolution differences, the method enhances coding efficiency, reduces resource utilization, and improves the performance of video codecs, resulting in a better user experience during video transmission and viewing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087712000001_ABST
    Figure 2025087712000001_ABST
Patent Text Reader

Abstract

To provide a method, encoding device, decoding device, system, and storage medium, for supporting a bi-directional optical flow (BDOF).SOLUTION: A method of decoding implemented by a video decoder includes: determining whether a resolution of a current picture being decoded is the same as resolutions of reference pictures identified by a reference picture list associated with the current picture; enabling, by the video decoder, a BDOF for a current block of the current picture when the resolution of the current picture is determined to be the same as the resolution of each of the reference pictures; disabling, by the video decoder, the BDOF for the current block of the current picture when the resolution of the current picture is determined to be different from a resolution of either of the reference pictures; and refining, by the video decoder, motion vectors corresponding to the current block using the BDOF when the BDOF is enabled for the current block.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 848,409, filed on May 15, 2019, by Jianle Chen et al., and is hereby incorporated by reference.

[0002] Generally, this disclosure describes techniques for supporting bi-directional optical flow (BDOF) in video coding. More specifically, this disclosure enables BDOF for reference picture resampling, but allows BDOF to be disabled for blocks or samples when the spatial resolution of the current picture and the reference picture is different.

Background Art

[0003] Even for relatively short videos, the amount of video data required to depict them can be quite large, which can pose difficulties when streaming data or communicating in other ways over a communication network with limited bandwidth capacity. Therefore, video data is generally compressed before being communicated over today's telecommunications networks. The size of videos can also be a problem when the videos are stored on storage devices because memory resources can be limited. Video compression devices often code video data using software and / or hardware at the source prior to transmission or storage, thereby reducing the amount of data required to represent the digital video image. Then, the compressed data is received at the destination by a video decompression device that decodes the video data. With limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve the compression ratio with little to no sacrifice in image quality are desirable.

Summary of the Invention

[0004] A first aspect relates to a method of decoding a coded video bitstream implemented by a video decoder. The method includes determining, by the video decoder, whether the resolution of the currently decoded picture is the same as the resolution of a reference picture specified by a reference picture list associated with the currently decoded picture; enabling bidirectional optical flow (BDOF) for a current block of the currently decoded picture when it is determined that the resolution of the currently decoded picture is the same as the resolution of each of the reference pictures; disabling BDOF for a current block of the currently decoded picture when it is determined that the resolution of the currently decoded picture is different from the resolution of any of the reference pictures; and refining a motion vector corresponding to the current block using BDOF when BDOF is enabled for the current block.

[0005] The method provides a technique that enables BDOF to be selectively disabled when the spatial resolution of the currently decoded picture is different from the spatial resolution of the reference pictures, instead of disabling BDOF for the entire CVS when reference picture resampling (RPR) is enabled. By having the ability to selectively disable BDOF in this way, the coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Accordingly, a coder / decoder (also known as a "codec") in video coding is improved over the current codec. In practice, an improved video coding process provides a better user experience to the user when the video is transmitted, received, and / or viewed.

[0006] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that enabling BDOF has setting the BDOF flag to a first value, and disabling BDOF has setting the BDOF flag to a second value.

[0007] Optionally, in any of the foregoing aspects, another implementation of that aspect provides generating a reference picture for a current picture based on a reference picture list according to a bidirectional inter prediction mode.

[0008] Optionally, in any of the foregoing aspects, another implementation of that aspect provides selectively enabling and disabling BDOF for blocks within a plurality of pictures depending on whether the resolution of each picture is different from or the same as the resolution of a reference picture associated with the picture.

[0009] Optionally, in any of the foregoing aspects, another implementation of that aspect provides enabling reference picture resampling (RPR) for an entire coded video sequence (CVS) including the current picture when BDOF is disabled.

[0010] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that the resolution of the current picture is arranged within a parameter set of the coded video bitstream and the current block is obtained from a slice of the current picture.

[0011] Optionally, in any of the foregoing aspects, another implementation of that aspect provides displaying an image generated using the current block on a display of an electronics device.

[0012] A second aspect relates to a method of encoding a video bitstream implemented by a video encoder. The method includes determining by the video encoder whether a resolution of a current picture being encoded is the same as a resolution of a reference picture specified in a reference picture list associated with the current picture; enabling bidirectional optical flow (BDOF) for a current block of the current picture by the video encoder when it is determined that the resolution of the current picture is the same as each resolution of the reference pictures; disabling BDOF for the current block of the current picture by the video encoder when it is determined that the resolution of the current picture is different from any resolution of the reference pictures; and refining a motion vector corresponding to the current block using BDOF by the video encoder when BDOF is enabled for the current block.

[0013] The method provides a technique that enables BDOF to be selectively disabled when a spatial resolution of a current picture is different from a spatial resolution of a reference picture, instead of disabling BDOF for the entire CVS when reference picture resampling (RPR) is enabled. By having the ability to selectively disable BDOF in this way, coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced at both the encoder and the decoder. Accordingly, a coder / decoder (also known as a “codec”) in video coding is improved over current codecs. In practice, an improved video coding process provides a better user experience to a user when video is transmitted, received, and / or viewed.

[0014] Optionally, in any of the foregoing aspects, another implementation of that aspect provides for determining motion vectors for a current picture based on reference pictures by a video encoder, encoding the current picture based on the motion vectors by the video encoder, and decoding the current picture using a hypothetical reference decoder by the video encoder.

[0015] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that enabling BDOF has the BDOF flag set to a first value, and disabling BDOF has the BDOF flag set to a second value.

[0016] Optionally, in any of the foregoing aspects, another implementation of that aspect provides for generating reference pictures for a current picture based on reference picture lists according to a bi - directional inter - prediction mode.

[0017] Optionally, in any of the foregoing aspects, another implementation of that aspect provides for selectively enabling and disabling BDOF for blocks within a plurality of pictures according to whether the resolution of each picture is different from or the same as the resolution of the reference pictures associated with the picture.

[0018] Optionally, in any of the foregoing aspects, another implementation of that aspect provides for enabling reference picture resampling (RPR) for an entire coded video sequence (CVS) including the current picture even when BDOF is disabled.

[0019] Optionally, in any of the foregoing aspects, another implementation of that aspect provides for transmitting a video bitstream including the current block to a video decoder.

[0020] A third aspect relates to a decoding apparatus. The decoding apparatus includes a receiver configured to receive a coded video bitstream, a memory coupled to the receiver and storing instructions, and a processor coupled to the memory and executing the instructions to cause the decoding apparatus to determine whether a resolution of a current picture being decoded is the same as a resolution of a reference picture specified by a reference picture list associated with a current block, enable bidirectional optical flow (BDOF) for a current block of the current picture when it is determined that the resolution of the current picture is the same as each resolution of the reference pictures, disable BDOF for the current block of the current picture when it is determined that the resolution of the current picture is different from any resolution of the reference pictures, and refine a motion vector corresponding to the current block when BDOF is enabled for the current block.

[0021] The decoding apparatus provides a technique that enables BDOF to be selectively disabled when a spatial resolution of a current picture is different from a spatial resolution of a reference picture, instead of disabling BDOF for the entire coded video sequence (CVS) when reference picture resampling (RPR) is enabled. By having the ability to selectively disable BDOF in this way, coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Accordingly, a coder / decoder (also known as a “codec”) in video coding is improved over the current codec. In practice, the improved video coding process provides a better user experience for the user when the video is transmitted, received, and / or viewed.

[0022] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that when BDOF is disabled, reference picture resampling (RPR) is enabled for the entire coded video sequence (CVS) including the current picture.

[0023] Optionally, in any of the foregoing aspects, another implementation of that aspect provides a display configured to display an image generated based on the current block.

[0024] A fourth aspect relates to an encoding apparatus. The encoding apparatus includes a memory storing instructions and a processor coupled to the memory. The instructions are implemented to cause the encoding apparatus to determine whether the resolution of the current picture being encoded is the same as the resolution of a reference picture specified by a reference picture list associated with the current picture. When it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures, bidirectional optical flow (BDOF) is enabled for the current block of the current picture. When it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures, BDOF is disabled for the current block of the current picture. When BDOF is enabled for the current block, a processor configured to refine a motion vector corresponding to the current block, and a transmitter coupled to the processor, the transmitter being configured to transmit a video bitstream including the current block towards a video decoder.

[0025] The encoding apparatus provides a technique that enables BDOF to be selectively disabled when the spatial resolution of the current picture is different from the spatial resolution of the reference picture, instead of having to disable BDOF for the entire CVS when reference picture resampling (RPR) is enabled. By having the ability to selectively disable BDOF in this way, the coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Accordingly, a coder / decoder (also known as a “codec”) in video coding is improved over the current codec. Practically, the improved video coding process provides a better user experience for the user when the video is transmitted, received, and / or viewed.

[0026] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that reference picture resampling (RPR) is enabled for the entire coded video sequence (CVS) including the current picture even when BDOF is disabled.

[0027] Optionally, in any of the foregoing aspects, another implementation of that aspect provides that the memory stores the video bitstream prior to the transmitter transmitting the bitstream towards the video decoder.

[0028] A fifth aspect relates to a coding apparatus. The coding apparatus includes a receiver configured to receive a picture to be coded or to receive a bitstream to be decoded, a transmitter coupled to the receiver, the transmitter being configured to transmit the bitstream to a decoder or to transmit the decoded picture to a display, a memory coupled to at least one of the receiver or the transmitter, the memory being configured to store instructions, and a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory to perform any of the methods disclosed herein.

[0029] Instead of having to disable BDOF for the entire CVS when reference picture sampling (RPR) is enabled, a technique is provided that enables BDOF to be selectively disabled when the spatial resolution of the current picture is different from the spatial resolution of the reference picture. By having the ability to selectively disable BDOF in this way, coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Accordingly, a coder / decoder (also known as a "codec") in video coding is improved over the current codec. As a practical matter, an improved video coding process provides a better user experience to the user when the video is transmitted, received, and / or viewed.

[0030] A sixth aspect relates to a system. The system includes an encoder and a decoder that communicates with the encoder, and the encoder or the decoder includes a decoding device, an encoding device, or a coding device disclosed herein.

[0031] Instead of having to disable BDOF for the entire CVS when reference picture sampling (RPR) is enabled, a technique is provided that enables BDOF to be selectively disabled when the spatial resolution of the current picture is different from the spatial resolution of the reference picture. By having the ability to selectively disable BDOF in this way, coding efficiency can be improved. Accordingly, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Accordingly, a coder / decoder (also known as a "codec") in video coding is improved over the current codec. As a practical matter, an improved video coding process provides a better user experience to the user when the video is transmitted, received, and / or viewed.

Brief Description of the Drawings

[0032] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in connection with the accompanying drawings and detailed description, in which like reference numerals represent like parts.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

DETAILED DESCRIPTION OF THE INVENTION

[0033] It should be understood initially that, while exemplary implementations of one or more embodiments are presented below, the disclosed system and / or method may be implemented using a number of techniques, whether currently known or yet to be developed. This disclosure should in no way be limited to the exemplary implementations, figures, and techniques illustrated below, including the exemplary designs and implementations depicted and described herein, but may be modified within the scope of the appended claims and their full equivalents.

[0034] As used herein, resolution describes the number of pixels within a video file. That is, resolution is the width and height of the projected image measured in pixels. For example, a video may have a resolution of 1280 (horizontal pixels) × 720 (vertical pixels). This is typically simply written as 1280×720, or abbreviated as 720p. BDOF is a process, algorithm, or coding tool used to refine motion or motion vectors for a predicted block. BDOF enables finding motion vectors for sub-coding units based on the gradient of the difference between two reference pictures. The RPR function is the ability to change the spatial resolution of a coded picture in the middle of a bitstream without the need for intra-coding of the picture at the location where the resolution changes.

[0035] FIG. 1 is a block diagram showing an example of a coding system 10 that can utilize the video coding techniques described herein. As shown in FIG. 1, the coding system 10 includes a source device 12 that provides encoded video data to be decoded by a destination device 14 at a later time. In particular, the source device 12 can provide the video data to the destination device 14 via a computer-readable medium 16. The source device 12 and the destination device 14 can each have any of a wide range of devices, including a desktop computer, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a telephone such as a so-called “smart” phone, a so-called “smart” pad, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, or the like. In some examples, the source device 12 and the destination device 14 can be provided for wireless communication.

[0036] Destination device 14 can receive encoded video data to be decoded via a computer-readable medium 16. The computer-readable medium 16 can have any type of medium or device capable of moving the encoded video data from source device 12 to destination device 14. In one example, the computer-readable medium 16 can have a communication medium that enables source device 12 to directly transmit the encoded video data to destination device 14 in real time. The encoded video data can be modulated according to a communication standard such as, for example, a wireless communication protocol and transmitted to destination device 14. The communication medium can have any wireless or wired communication medium such as, for example, a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as, for example, a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other device useful for facilitating communication from source device 12 to destination device 14.

[0037] In some examples, the encoded data can be output from the output interface 22 to a storage device. Similarly, the encoded data can be accessed from the storage device by an input interface. The storage device can include any of a variety of distributed or locally accessible data storage media, such as, for example, a hard drive, a Blu-ray disc, a digital video disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other digital storage media suitable for storing encoded video data. In a further example, the storage device may correspond to a file server or other intermediate storage device that can store the encoded video generated by the source device 12. The destination device 14 can access the stored video data from the storage device via streaming or download. The file server can be any type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Examples of file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network-attached storage (NAS) devices, or local disk drives. The destination device 14 can access the encoded video data via any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device can be a streaming transmission, a download transmission, or a combination thereof.

[0038] The technology disclosed herein is not necessarily limited to wireless applications or settings. The technology can be applied to video coding to support any of a variety of multimedia applications, such as, for example, over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the coding system 10 can be configured to support one-way or two-way video transmission to support applications such as, for example, video streaming, video playback, video broadcast, and / or video telephony.

[0039] In the example of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to this disclosure, the video encoder 20 of the source device 12 and / or the video decoder 30 of the destination device 14 can be configured to apply the technology for video coding. In other examples, the source device and the destination device may include other components or configurations. For example, the source device 12 may receive video data from an external video source such as an external camera. Similarly, the destination device 14 may interface with an external display device rather than including an integrated display device.

[0040] The illustrated coding system 10 of FIG. 1 is merely an example. Techniques for video coding can be performed by any digital video encoding and / or decoding device. The techniques of this disclosure are generally performed by a video coding device, but the techniques may also typically be performed by a video encoder / decoder, commonly referred to as a “CODEC”. Additionally, the techniques of this disclosure may also be performed by a video preprocessor. The video encoder and / or decoder can be a graphics processing unit (GPU) or a similar device.

[0041] Source device 12 and destination device 14 are merely examples of such coding devices where source device 12 generates video data coded for transmission to destination device 14. In some examples, source device 12 and destination device 14 can operate in a substantially symmetric manner such that each of source device and destination devices 12, 14 includes video encoding and decoding components. Accordingly, coding system 10 can support one-way or two-way video transmission between video devices 12, 14 for, e.g., video streaming, video playback, video broadcast, or video telephony.

[0042] The video source 18 of source device 12 can include, for example, a video capture device such as a video camera, a video archive storing previously captured video, and / or a video feed interface that receives video from a video content provider. As a further alternative, video source 18 can generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video.

[0043] In some cases, when the video source 18 is a video camera, the source device 12 and the destination device 14 may form a so-called camera phone or video phone. However, as described above, the technology described in this disclosure is generally applicable to video coding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. And the encoded video information may be output to the computer-readable medium 16 by the output interface 22.

[0044] The computer-readable medium 16 may include a transient medium such as a wireless broadcast or a wired network transmission, or a storage medium (i.e., a non-transient storage medium) such as a hard disk, a flash drive, a compact disk, a digital video disk, a Blu-ray disk, or other computer-readable media. In some examples, a network server (not shown) may receive the encoded video data from the source device 12 and provide the encoded video data to the destination device 14 via, for example, a network transmission. Similarly, a computing device of a media production facility such as a disk stamping facility may receive the encoded video data from the source device 12 and produce a disk containing the encoded video data. Therefore, the computer-readable medium 16 can be understood to include one or more computer-readable media in various forms in various examples.

[0045] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 can include syntax information defined by the video encoder 20, and this syntax information is also used by the video decoder 30 and includes syntax elements that describe the characteristics and / or processing of blocks and / or other coding units such as, for example, a group of pictures (GOP). The display device 32 displays the decoded video data to the user and can include any of a variety of display devices such as, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0046] Video encoder 20 and video decoder 30 may operate in accordance with video coding standards such as the currently under - development High Efficiency Video Coding (HEVC) standard, and may conform to the HEVC Test Model (HM). Alternatively, video encoder 20 and video decoder 30 may operate in accordance with other proprietary or industry standards, such as the International Telecommunication Union Telecommunication Standardization Sector (ITU - T) H.264 standard, also known as Moving Picture Experts Group (MPEG) - 4 Part 10, Advanced Video Coding (AVC), H.265 / HEVC, or extensions of such standards. However, the techniques of this disclosure are not limited to any particular coding standard. Other examples of video coding standards include MPEG - 2 and ITU - T H.263. Although not shown in FIG. 1, in some aspects, video encoder 20 and video decoder 30 may each be integrated with an audio encoder and decoder, and may include a multiplexer - demultiplexer (MUX - DEMUX) unit suitable for handling the encoding of both audio and video in a common data stream or separate data streams, or other hardware and software. When applicable, the MUX - DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as, for example, the User Datagram Protocol (UDP).

[0047] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder circuits, such as, for example, one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If the technology is implemented partially in software, the apparatus may store the software instructions on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, and they may both be integrated as part of a combined encoder / decoder (CODEC) within their respective devices. An apparatus including video encoder 20 and / or video decoder 30 may have an integrated circuit, a microprocessor, and / or a wireless communication device such as, for example, a cellular phone.

[0048] FIG. 2 is a block diagram illustrating an example of video encoder 20 that may implement video coding techniques. Video encoder 20 may perform intra coding and inter coding of video blocks within a video slice. Intra coding relies on spatial prediction to reduce or remove spatial redundancy in video within a given video frame or picture. Inter coding relies on temporal prediction to reduce or remove temporal redundancy in video within adjacent frames or pictures of a video sequence. Intra mode (I mode) may refer to any of several spatial-based coding modes. Inter modes, such as, for example, uni-directional (also known as uni prediction) prediction (P mode) or bi-prediction (also known as bi prediction) (B mode), may refer to any of several time-based coding modes.

[0049] As shown in FIG. 2, video encoder 20 receives a current video block within a video frame to be encoded. In the example of FIG. 2, video encoder 20 includes a mode selection unit 40, a reference frame memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Instead, mode selection unit 40 includes a motion compensation unit 44, a motion estimation unit 42, an intra prediction (also known as intra-prediction) unit 46, and a partitioning unit 48. For video block reconstruction, video encoder 20 also includes an inverse quantization unit 58, an inverse transform unit 60, and an adder 62. A deblocking filter (not shown in FIG. 2) that filters block boundaries to remove blocky artifacts from the reconstructed video may also be included. Optionally, the deblocking filter typically filters the output of adder 62. Also, in addition to the deblocking filter, further filters (in-loop or post-loop) may be used. Such filters are not shown for simplicity, but optionally may filter the output of adder 50 (as an in-loop filter).

[0050] In the encoding process, video encoder 20 receives a video frame or slice to be coded. The frame or slice may be divided into a plurality of video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-prediction coding of the received video block with respect to one or more blocks in one or more reference frames to provide temporal prediction. Instead, intra prediction unit 46 may perform intra-prediction coding of the received video block with respect to one or more adjacent blocks in the same frame or slice as the block to be coded to provide spatial prediction. Video encoder 20 may execute a plurality of coding paths to select an appropriate coding mode for, for example, each block of video data.

[0051] Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on an evaluation of a previous partitioning method in a previous coding path. For example, the partitioning unit 48 may first partition a frame or a slice into maximum coding units (LCUs), and then partition each of the LCUs into sub-coding units based on rate distortion analysis (e.g., rate distortion optimization). The mode selection unit 40 may further create a quadtree data structure indicating the partitioning of the LCUs into sub-CUs. A leaf node CU of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs).

[0052] This disclosure uses the term “block” to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure in the context of other standards (e.g., macroblocks and their sub-blocks in H.264 / AVC). A CU includes a coding node, a PU, and a TU associated with the coding node. The size of a CU corresponds to the size of the coding node and is square-shaped. The size of a CU may range from 8×8 pixels up to the size of a tree block with a maximum of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. Syntax data associated with a CU may describe, for example, the partitioning of the CU into one or more PUs. The partitioning mode may vary depending on whether the CU is coded in skip or direct mode, intra prediction mode, or inter prediction (also known as inter-prediction) mode. A PU may be partitioned into a non-square shape. Syntax data associated with a CU may also describe, for example, the partitioning of the CU into one or more TUs according to a quadtree. A TU may be square or non-square (e.g., rectangular) in shape.

[0053] The mode selection unit 40 can select one of the intra or inter coding modes, for example, based on the error result, and provide the obtained intra or inter coded block to the adder 50 that generates residual block data and the adder 62 that reconstructs the coded block for use as a reference frame. The mode selection unit 40 also provides syntax elements such as, for example, motion vectors, intra mode indicators, partition information, and other such syntax information to the entropy coding unit 56.

[0054] The motion estimation unit 42 and the motion compensation unit 44 can be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion regarding a video block. The motion vector can indicate, for example, the displacement of the PU of the video block in the current video frame or picture with respect to the predicted block (or other coded unit) in the reference frame with respect to the currently coded current block (or other unit to be coded) in the current frame. The predicted block is a block that is found to match well with the block to be coded regarding the pixel difference that can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some examples, the video encoder 20 can calculate the values of the sub - integer pixel positions of the reference picture stored in the reference frame memory 64. For example, the video encoder 20 can interpolate the values of the 1 / 4 - pixel position, 1 / 8 - pixel position, or other fractional pixel positions of the reference picture. Therefore, the motion estimation unit 42 can perform a motion search for full - pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0055] The motion estimation unit 42 calculates the motion vector for the PU of the video block within the inter-coded slice by comparing the position of the PU with the position of the prediction block in the reference picture. The reference picture can be selected from the first reference picture list (List0) or the second reference picture list (List1) that each identify one or more reference pictures stored in the reference frame memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.

[0056] The motion compensation executed by the motion compensation unit 44 may include fetching or generating a prediction block based on the motion vector determined by the motion estimation unit 42. Again, in some examples, the motion estimation unit 42 and the motion compensation unit 44 may be functionally integrated. Receiving the motion vector for the current video block's PU, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector within one of the reference picture lists. The adder 50 forms a residual video block by subtracting the pixel value of the prediction block from the pixel value of the currently coded video block, as will be described later, to form the value of the pixel difference. Generally, the motion estimation unit 42 performs motion estimation on the luma component, and the motion compensation unit 44 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The mode selection unit 40 may also generate syntax elements related to the video block and the video slice for use by the video decoder 30 when decoding the video blocks of the video slice.

[0057] As described above, the intra prediction unit 46 can perform intra prediction on the current block as an alternative to the inter prediction executed by the motion estimation unit 42 and the motion compensation unit 44. In particular, the intra prediction unit 46 can determine an intra prediction mode to be used for encoding the current block. In some examples, the intra prediction unit 46 can encode the current block using various intra prediction modes, for example, between different encoding paths, and the intra prediction unit 46 (or, in some examples, the mode selection unit 40) can select an appropriate intra prediction mode for use from the tested modes.

[0058] For example, the intra prediction unit 46 can calculate rate distortion values using rate distortion analysis for various tested intra prediction modes and select an intra prediction mode having the best rate distortion characteristics among those tested modes. Rate distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block from which the encoded block was generated, and the bit rate (i.e., the number of bits) used to generate the encoded block. The intra prediction unit 46 calculates a ratio from the distortion and rate for various encoded blocks and determines which intra prediction mode exhibits the best rate distortion value for that block.

[0059] In addition, the intra prediction unit 46 can be configured to code depth blocks of a depth map using a depth modeling mode (DMM). The mode selection unit 40 can determine, for example, using rate distortion optimization (RDO), whether the available DMM mode produces better coding results than the intra prediction mode and other DMM modes. The data of the texture image corresponding to the depth map can be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 can also be configured to perform inter prediction on the depth blocks of the depth map.

[0060] After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include configuration data in the bitstream to be transmitted, and the configuration data may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of coding contexts for various blocks, and the most likely intra prediction mode, intra prediction mode index table, and modified intra prediction mode index table used for each of those contexts.

[0061] The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the original video block being coded. An adder 50 represents one or more components that perform this subtraction operation.

[0062] The transform processing unit 52 applies a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform, to the residual block to generate a video block having residual transform coefficient values. The transform processing unit 52 may perform other transforms conceptually similar to the DCT. A wavelet transform, integer transform, subband transform, or other type of transform may also be used.

[0063] The conversion processing unit 52 applies the conversion to the residual blocks to generate a block of residual conversion coefficients. The conversion can convert the residual information from the pixel value domain to a conversion domain such as, for example, the frequency domain. The conversion processing unit 52 can send the obtained conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be changed by adjusting the quantization parameter. In some examples, the quantization unit 54 can then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0064] Following quantization, the entropy coding unit 56 entropy-codes the quantized conversion coefficients. For example, the entropy coding unit 56 can perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques). In the case of context-based entropy coding, the context can be based on adjacent blocks. Following the entropy coding by the entropy coding unit 56, the encoded bit stream can be sent to another device (e.g., the decoder 30 in the case of video), or archived for later transmission or retrieval.

[0065] The inverse quantization unit 58 and the inverse transform unit 60 each apply inverse quantization and inverse transform to reconstruct a residual block in the pixel domain, for example, for later use as a reference block. The motion compensation unit 44 may calculate a reference block by adding the residual block to a predicted block of one of the frames in the reference frame memory 64. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values used for motion estimation. The adder 62 adds the reconstructed residual block to the motion-compensated predicted block generated by the motion compensation unit 44 to generate a reconstructed video block stored in the reference frame memory 64. The reconstructed video block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-coding blocks in subsequent video frames.

[0066] FIG. 3 is a block diagram illustrating an example of a video decoder 30 that may implement video coding techniques. In the example of FIG. 3, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. The video decoder 30 may, in some examples, perform a decoding path generally inverse to the encoding path described with respect to the video encoder 20 (FIG. 2). The motion compensation unit 72 can generate prediction data based on the motion vectors received from the entropy decoding unit 70, and the intra prediction unit 74 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 70.

[0067] In the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements from video encoder 20. Entropy decoding unit 70 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Entropy decoding unit 70 transfers the motion vectors and other syntax elements to motion compensation unit 72. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level.

[0068] When the video slice is coded as an intra-coded (I) slice, intra prediction unit 74 may generate prediction data for the video blocks of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current frame or picture. When the video frame is coded as an inter-coded (e.g., B, P, or GPB) slice, motion compensation unit 72 generates a prediction block for the video blocks of the current video slice based on the motion vectors and other syntax elements received from entropy decoding unit 70. The prediction block may be generated from one of the reference pictures within one of the reference picture lists. Video decoder 30 may construct reference frame lists List0 and List1 using default construction techniques based on the reference pictures stored in reference frame memory 82.

[0069] The motion compensation unit 72 determines prediction information about the video blocks of the current video slice by syntax-analyzing the motion vectors and other syntax elements, and generates a prediction block for the currently decoded video block using the prediction information. For example, the motion compensation unit 72 uses a part of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to code the video blocks of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), one or more construction information of the reference picture list related to the slice, the motion vectors for each inter-coded video block of the slice, the inter prediction status for each inter-coded video block of the slice, and other information for decoding the video blocks within the current video slice.

[0070] The motion compensation unit 72 may also perform interpolation based on an interpolation filter. The motion compensation unit 72 may calculate an interpolation value for sub-integer pixels of a reference block using the interpolation filter used by the video encoder 20 in the encoding of the video block. In this case, the motion compensation unit 72 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and generate a prediction block using the interpolation filter.

[0071] Data about the texture image corresponding to the depth map may be stored in the reference frame memory 82. The motion compensation unit 72 may also be configured to inter-predict the depth blocks of the depth map.

[0072] In one embodiment, video decoder 30 includes a user interface (UI) 84. The user interface 84 is configured to receive input from a user (e.g., a network administrator) of the video decoder 30. Through the user interface 84, the user can manage or change settings for the video decoder 30. For example, the user can input or otherwise provide values for parameters (e.g., flags) in order to control the configuration and / or operation of the video decoder 30 according to user preferences. The user interface 84 can be a graphical user interface (GUI) that enables the user to interact with the video decoder 30, for example, via graphical icons, drop-down menus, and check boxes. In some cases, the user interface 84 can receive information from the user via a keyboard, mouse, or other peripheral device. In one embodiment, the user can access the user interface 84 via a smartphone, tablet device, and personal computer, etc., which are remotely located from the video decoder 30. As used herein, the user interface 84 may be referred to as an external input or external means.

[0073] With the above in mind, video compression techniques perform spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in a video sequence. In block-based video coding, a video slice (i.e., a video picture, or a part of a video picture) can be partitioned into a plurality of video blocks that may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks within an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks within an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in adjacent blocks within the same picture, or temporal prediction with respect to reference samples in other reference pictures. A picture may sometimes be referred to as a frame, and a reference picture may sometimes be referred to as a reference frame.

[0074] Spatial or temporal prediction results in a prediction block for the block to be coded. Residual data represents the pixel difference between the original block to be coded and the prediction block. An inter-coded block is encoded according to a motion vector that points to a block of reference samples forming the prediction block, with the residual data indicating the difference between the block being coded and the prediction block. An intra-coded block is encoded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain, resulting in residual transform coefficients that may then be quantized. The quantized transform coefficients, which were initially arranged in a two-dimensional array, are scanned to generate a one-dimensional vector of transform coefficients, and entropy coding may be applied to achieve further compression.

[0075] Image and video compression has undergone rapid growth and led to various coding standards. Such video coding standards include ITU-T H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) MPEG-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC) also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC) and 3D HEVC (3D-HEVC).

[0076] There is also a new video coding standard named Versatile Video Coding (VVC) under development by the Joint Video Expert Team (JVET) of ITU-T and ISO / IEC. The VVC standard has several working drafts, and in particular, one working draft (WD) of VVC, namely “Versatile Video Coding (Draft 5)” by B. Bross, J. Chen, and S. Liu, JVET-N1001-v3, 13th JVET Meeting, March 27, 2019 (VVC Draft 5), is hereby incorporated by reference in its entirety.

[0077] The description of the technology disclosed herein is based on the Versatile Video Coding (VVC), a video coding standard under development by the Joint Video Expert Team (JVET) of ITU-T and ISO / IEC. However, the technology is also applicable to other video codec specifications.

[0078] Figure 4 shows a relationship 400 between an intra random access point (IRAP) picture 402 and leading pictures 404 and trailing pictures 406 in decoding order 408 and presentation order 410. In one embodiment, the IRAP picture 402 is referred to as a clean random access (CRA) picture, or alternatively as an instantaneous decoder refresh (IDR) picture having a random access decodable (RADL) picture. In HEVC, all IDR pictures, CRA pictures, and Broken Link Access (BLA) pictures are considered to be IRAP pictures 402. In VVC, it was agreed at the 12th JVET meeting in October 2018 that both IDR pictures and CRA pictures are to be IRAP pictures. In one embodiment, Broken Link Access (BLA) pictures and Gradual Decoder Refresh (GDR) pictures may also be considered to be IRAP pictures. The decoding process of a coded video sequence always starts at an IRAP.

[0079] As shown in Figure 4, the leading pictures 404 (e.g., pictures 2 and 3) are behind the IRAP picture 402 in decoding order 408, but are ahead of the IRAP picture 402 in presentation order 410. The trailing picture 406 is behind the IRAP picture 402 in both decoding order 408 and presentation order 410. Although Figure 4 shows two leading pictures 404 and one trailing picture 406, those skilled in the art will understand that in actual applications, there may be more or fewer leading pictures 404 and / or trailing pictures 406 in decoding order 408 and presentation order 410.

[0080] The leading picture 404 in FIG. 4 is divided into two types, namely random access skipped leading (RASL) and RADL. When decoding starts with an IRAP picture 402 (e.g., picture 1), the RADL picture (e.g., picture 3) can be decoded properly, but the RASL picture (e.g., picture 2) cannot be decoded properly. Therefore, the RASL picture is discarded. In view of the distinction between the RADL picture and the RASL picture, for efficient and proper coding, the type of the leading picture 404 related to the IRAP picture 402 should be specified as either RADL or RASL. In HEVC, when there are RASL pictures and RADL pictures, there is a constraint that for the RASL pictures and RADL pictures related to the same IRAP picture 402, the RASL picture shall precede the RADL picture in the presentation order 410.

[0081] The IRAP picture 402 provides the following two important functions / benefits. First, the presence of the IRAP picture 402 indicates that the decoding process can start from that picture. This function enables the random accessibility that the decoding process starts not necessarily at the beginning of the bitstream but at that position within the bitstream as long as the IRAP picture 402 exists at that position. Second, the presence of the IRAP picture 402 refreshes the decoding process so that the pictures coded starting from the IRAP picture 402, except for the RASL pictures, are coded without any reference to the previous pictures. The presence of the IRAP picture 402 in the bitstream thus stops the errors that may occur during the decoding of the pictures coded prior to the IRAP picture 402 from propagating to the IRAP picture 402 and the pictures after the IRAP picture 402 in the decoding order 408.

[0082] While the IRAP picture 402 provides important functions, it comes with a disadvantage in terms of compression efficiency. The presence of the IRAP picture 402 causes a bitrate surge. This disadvantage in compression efficiency is due to two reasons. First, since the IRAP picture 402 is an intra-predicted picture, the picture itself requires relatively more bits to represent compared to other pictures that are inter-predicted pictures (e.g., the leading picture 404, the trailing picture 406). Second, the presence of the IRAP picture 402 causes the temporal prediction to be interrupted (this is because the decoder will refresh the decoding process, and one of the operations of the decoding process for this purpose is to remove the reference pictures in the decoded picture buffer (DPB)), so the IRAP picture 402 makes the coding of the pictures after the IRAP picture 402 in the decoding order 408 less efficient (i.e., requires even more bits to represent). This is because the reference pictures for their inter-prediction coding that those pictures have become fewer.

[0083] Among the picture types regarded as IRAP pictures 402, the IDR pictures in HEVC have different signaling and derivation compared to other picture types. Some of the differences are as follows.

[0084] In the signaling and derivation of the picture order count (POC) value of the IDR picture, the most significant bit (MSB) part of the POC is not derived from the preceding key picture and is simply set equal to 0.

[0085] Regarding signaling the information necessary for reference picture management, the slice header of an IDR picture does not contain the information that needs to be signaled to assist reference picture management. In the case of other picture types (i.e., CRA, trailing, temporal sub-layer access (TSA), etc.), for the reference picture marking process (i.e., the process of determining the status of reference pictures in the decoded picture buffer (DPB) as either used or not used for reference), information such as, for example, the reference picture set (RPS) described later or other forms of similar information (e.g., reference picture list) is required. However, in the case of an IDR picture, such information does not need to be signaled. This is because the presence of an IDR indicates that the decoding process simply marks all reference pictures in the DPB as not used for reference.

[0086] In HEVC and VVC, an IRAP picture 402 and a leading picture 404 may each be included within a single network abstraction layer (NAL) unit. A set of NAL units may be referred to as an access unit. The IRAP picture 402 and the leading picture 404 are given different NAL unit types so that they can be easily identified by a system-level application. For example, a video splicer needs to understand the coded picture type without having to understand too many details of the syntax elements in the coded bitstream, especially to identify the IRAP picture 402 from non-IRAP pictures and to identify the leading picture 404 from trailing pictures 406, including determining RASL pictures and RADL pictures. A trailing picture 406 is a picture that is associated with the IRAP picture 402 and is after the IRAP picture 402 in the presentation order 410. A certain picture may be after a specific IRAP picture 402 in the decoding order 408 and may precede any other IRAP picture 402 in the decoding order 408. In contrast, giving the IRAP picture 402 and the leading picture 404 their own NAL unit types helps with such applications.

[0087] In HEVC, the NAL unit types for IRAP pictures include the following: BLA with leading pictures (BLA_W_LP): The NAL unit of a broken link access (BLA) picture that may be followed by one or more leading pictures in the decoding order; BLA with RADL (BLA_W_RADL): The NAL unit of a BLA picture that may be followed by one or more RADL pictures in the decoding order but not by RASL pictures; BLA without leading pictures (BLA_N_LP): The NAL unit of a BLA picture that is not followed by a leading picture in the decoding order; IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture that can be followed by one or more RADL pictures in decoding order but not by a RASL picture; IDR without a leading picture (IDR_N_LP): A NAL unit of an IDR picture that is not followed by a leading picture in decoding order; CRA: A NAL unit of a clean random access (CRA) picture that can be followed by a leading picture (i.e., either a RASL picture or a RADL picture, or both); RADL: A NAL unit of a RADL picture; RASL: A NAL unit of a RASL picture.

[0088] In VVC, the NAL unit types for IRAP picture 402 and leading picture 404 are as follows: IDR with RADL (IDR_W_RADL): A NAL unit of an IDR picture that can be followed by one or more RADL pictures in decoding order but not by a RASL picture; IDR without a leading picture (IDR_N_LP): A NAL unit of an IDR picture that is not followed by a leading picture in decoding order; CRA: A NAL unit of a clean random access (CRA) picture that can be followed by a leading picture (i.e., either a RASL picture or a RADL picture, or both); RADL: A NAL unit of a RADL picture; RASL: A NAL unit of a RASL picture.

[0089] The reference picture sampling (RPR) function is the ability to change the spatial resolution of the picture being coded in the middle of the bitstream without the need for intra coding of the picture at the resolution change position. To enable this, the picture needs to be able to reference one or more reference pictures with a spatial resolution different from that of the current picture for inter prediction purposes. Therefore, resampling of such a reference picture or a part thereof is required for the encoding and decoding of the current picture. Hence the name RPR. This function may also be called adaptive resolution change (ARC) or by other names. There are use cases or application scenarios that benefit from the RPR function, including the following.

[0090] Rate adaptation in video telephony and conferencing. This is for adapting the coded video to changing network conditions. When the network condition deteriorates and the available bandwidth becomes small, the encoder can adapt by encoding pictures with a lower resolution.

[0091] Active speaker change in multiparty video conferencing. In a multiparty video conferencing, it is common that the video size for the active speaker is larger or wider than that for the remaining conference participants. When the active speaker changes, the picture resolution for each participant may also need to be adjusted. The need for the ARC function becomes even more important when the active speaker changes frequently.

[0092] Fast start in streaming. In a streaming application, it is common for the application to buffer until it reaches a certain number of decoded pictures before starting to display the pictures. Starting the bitstream at a lower resolution enables the application to have enough pictures in the buffer to start the display faster.

[0093] Adaptive stream switching in streaming. The Dynamic Adaptive Streaming over HTTP (DASH) specification includes a feature named @mediaStreamStructureId. This feature enables switching between different representations at open GOP random access points with non-decodable leading pictures, such as CRA pictures with associated RASL pictures in HEVC for example. When two different representations of the same video have different bitrates but the same spatial resolution and they have the same value of @mediaStreamStructureId, switching between the two representations can be performed at a CRA picture with an associated RASL picture, and the RASL picture associated with the CRA picture at the switching position can be decoded with an acceptable quality, thus enabling seamless switching. By using ARC, the @mediaStreamStructureId feature can also be used for switching between DASH representations with different spatial resolutions.

[0094] Various methods, such as signaling like a list of picture resolutions, some constraints on resampling of reference pictures in the DPB, etc., assist the basic techniques for supporting RPR / ARC. Also, at the 14th JVET meeting in Geneva, there were several contributions proposing the constraints to be applied to VVC for supporting RPR. The proposed constraints include the following.

[0095] When the current picture references a reference picture with a different resolution from the current picture, some tools shall be disabled for the coding of blocks in the current picture. Those tools include the following.

[0096] Temporal motion vector prediction (TMVP) and advanced TMVP (ATMVP). This is proposed by JVET-N0118.

[0097] Decoder Side Motion Vector Refinement (DMVR). This was proposed by JVET-N0279.

[0098] Bi-directional Optical Flow (BIO). This was proposed by JVET-N0279.

[0099] Bi-prediction of blocks from reference pictures with a different resolution than the current picture is not allowed. This was proposed by JVET-N0118.

[0100] Regarding motion compensation, sample filtering is to be applied only once, i.e., when resampling and interpolation are required to reach a finer per-pel resolution (e.g., 1 / 4 pel resolution), these two filters need to be combined and applied only once. This was proposed by JVET-N0118.

[0101] Scalability in video coding is typically supported by using multi-layer coding techniques. A multi-layer bitstream includes a base layer (BL) and one or more enhancement layers (EL). Examples of scalability include spatial scalability, quality / signal-to-noise (SNR) scalability, multi-view scalability, etc. When multi-layer coding techniques are used, a picture or a part thereof can be coded by (1) without using a reference picture, i.e., using intra prediction, (2) referring to a reference picture within the same layer, i.e., using inter prediction, or (3) referring to a reference picture within (one or more) other layers, i.e., using inter-layer prediction. The reference picture used for inter-layer prediction of the current picture is called the inter-layer reference picture (ILRP).

[0102] FIG. 5 shows an example of multi-layer coding for spatial scalability 500. The picture 502 in layer N has a different resolution (e.g., a lower resolution) than the picture 504 in layer N+1. In one embodiment, layer N is considered the base layer and layer N+1 is considered the enhancement layer as described above. The picture 502 in layer N and the picture 504 in layer N+1 can be coded using inter prediction (as indicated by the solid arrows). The picture 502 may also be coded using inter-layer prediction (as indicated by the dashed arrows).

[0103] In the context of RPR, the reference picture can be resampled either by selecting the reference picture from a lower layer or by generating the reference picture for a higher layer based on the reference picture of a lower layer using inter-layer prediction.

[0104] Previous H.26x video coding families have provided support for scalability in (one or more) profiles different from the (one or more) profiles for single-layer coding. Scalable Video Coding (SVC) is a scalable extension of AVC / H.264 that provides support for spatial, temporal, and quality scalability. In SVC, within each macroblock (MB) in an EL picture, a flag indicating whether the EL MB is predicted using a collocated block from a lower layer is signaled. Prediction from a collocated block can include texture, motion vectors, and / or coding mode. Implementations of SVC cannot directly reuse unmodified H.264 / AVC implementations in their designs. The SVC EL macroblock syntax and decoding process are different from the H.264 / AVC syntax and decoding process.

[0105] Scalable HEVC (SHVC) is an extension of the HEVC / H.265 standard that provides support for spatial and quality scalability. Multi-View HEVC (MV-HEVC) is an extension of HEVC / H.265 that provides support for multi-view scalability. 3D HEVC (3D-HEVC) is an extension of HEVC / H.264 that provides support for more advanced and efficient three-dimensional (3D) video coding than MV-HEVC. Note that temporal scalability is included as an integral part of the single-layer HEVC codec. The design of the multi-layer extension of HEVC adopts the idea that the decoded pictures used for inter-layer prediction are treated as long-term reference pictures (LTRP) derived only from the same access unit (AU) and are assigned reference indices in the (one or more) reference picture lists together with other temporal reference pictures within the current layer. Inter-layer prediction (ILP) is achieved at the prediction unit (PU) level by setting the values of the reference indices for referring to the (one or more) inter-layer reference pictures within the (one or more) reference picture lists.

[0106] In particular, both the reference picture resampling function and the spatial scalability function require resampling of the reference picture or a part thereof. The reference picture resampling can be realized either at the picture level or at the coding block level. However, when referring to RPR as a coding function, it is a function for single-layer coding. Even so, it is possible to use the same resampling filter for both the RPR function of single-layer coding and the spatial scalability function for multi-layer coding, and it is even preferable from the perspective of codec design.

[0107] JVET-N0279 proposed to disable BIO for RPR. More precisely, what it proposed was to disable the use of BIO (also known as BDOF) for the entire coded video sequence (CVS) when RPR is enabled. It should be noted that even when the RPR function is enabled, the current picture often does not refer to reference pictures with different resolutions. Therefore, disabling BIO for the entire CVS can unnecessarily limit and compromise coding efficiency.

[0108] Disclosed herein is a technique that enables selective disabling of BDOF when the spatial resolution of the current picture is different from that of the reference picture, instead of disabling BDOF for the entire CVS when RPR is enabled. By having the ability to selectively disable BDOF in this way, coding efficiency can be improved. Therefore, the use of processor, memory, and / or network resources can be reduced in both the encoder and the decoder. Consequently, the coder / decoder (also known as "codec") in video coding is improved over the current codec. In practice, the improved video coding process provides a better user experience for the user when the video is transmitted, received, and / or viewed.

[0109] FIG. 6 is a schematic diagram showing an example of one-way inter prediction 600. One-way inter prediction 600 can be used to determine motion vectors for the coded and / or decoded blocks created when partitioning a picture.

[0110] The slice direction inter-prediction 600 uses a reference frame 630 having a reference block 631 to predict a current block 611 within a current frame 610. The reference frame 630 may be positioned temporally after the current frame 610 (e.g., as a subsequent reference frame) as shown, but in some examples, it may also be positioned temporally before the current frame 610 (e.g., as a preceding reference frame). The current frame 610 is an example of a frame / picture that is encoded / decoded at a particular point in time. The current frame 610 includes an object within the current block 611 that matches an object within the reference block 631 of the reference frame 630. The reference frame 630 is a frame used as a reference for encoding the current frame 610, and the reference block 631 is a block within the reference frame 630 that includes an object also included in the current block 611 of the current frame 610.

[0111] Current block 611 is any coding unit that is encoded / decoded at a specific point in the coding process. Current block 611 may be the entire partitioned block or, alternatively, a sub-block when using the affine inter prediction mode. Current frame 610 is separated from reference frame 630 by some temporal distance (TD) 633. TD 633 indicates the amount of time between current frame 610 and reference frame 630 in the video sequence and can be measured in frame units. Prediction information for current block 611 may refer to reference frame 630 and / or reference block 631 by a reference index that indicates the direction and temporal distance between these frames. Throughout the period represented by TD 633, an object within current block 611 moves from its position in current frame 610 to another position in reference frame 630 (e.g., the position of reference block 631). For example, the object may move along motion trajectory 613, which is the direction of movement of the object over time. Motion vector 635 describes the direction and magnitude of the movement of the object along motion trajectory 613 throughout TD 633. Thus, the encoded motion vector 635, reference block 631, and a residual including the difference between current block 611 and reference block 631 provide sufficient information to reconstruct current block 611 and position current block 611 within current frame 610.

[0112] FIG. 7 is a schematic diagram showing an example of bidirectional inter prediction 700. Bidirectional inter prediction 700 can be used to determine motion vectors for encoded and / or decoded blocks created when partitioning a picture.

[0113] The bidirectional inter prediction 700 is similar to the unidirectional inter prediction 600, but uses a pair of reference frames to predict the current block 711 within the current frame 710. Thus, the current frame 710 and the current block 711 are substantially the same as the current frame 610 and the current block 611, respectively. The current frame 710 is temporally located between a previous reference frame 720 that appears before the current frame 710 in the video sequence and a subsequent reference frame 730 that appears after the current frame 710 in the video sequence. The previous reference frame 720 and the subsequent reference frame 730 are substantially the same as the reference frame 630 in other respects.

[0114] The current block 711 matches a previous reference block 721 within the previous reference frame 720 and a subsequent reference block 731 within the subsequent reference frame 730. Such a match indicates that, in the course of the video sequence, an object moves from the position in the previous reference block 721, along the motion trajectory 713, through the current block 711, to the position in the subsequent reference block 731. The current frame 710 is separated from the previous reference frame 720 by some previous time distance (TD0) 723 and is separated from the subsequent reference frame 730 by some subsequent time distance (TD1) 733. TD0 723 indicates the amount of time in frame units between the previous reference frame 720 and the current frame 710 in the video sequence. TD1 733 indicates the amount of time in frame units between the current frame 710 and the subsequent reference frame 730 in the video sequence. Thus, the object moves from the previous reference block 721 to the current block 711 along the motion trajectory 713 over the period indicated by TD0 723. The object also moves from the current block 711 to the subsequent reference block 731 along the motion trajectory 713 over the period indicated by TD1 733. The prediction information for the current block 711 can be referenced by the previous reference frame 720 and / or the previous reference block 721 and the subsequent reference frame 730 and / or the subsequent reference block 731 by a pair of reference indices that indicate the direction and time distance between these frames.

[0115] The previous motion vector (MV0) 725 describes the direction and magnitude of the movement of an object along the motion trajectory 713 through the TD0 723 (e.g., between the previous reference frame 720 and the current frame 710). The subsequent motion vector (MV1) 735 describes the direction and magnitude of the movement of the object along the motion trajectory 713 through the TD1 733 (e.g., between the current frame 710 and the subsequent reference frame 730). Thus, in the bidirectional inter prediction 700, the current block 711 can be coded and reconstructed by using the previous reference block 721 and / or the subsequent reference block 731, the MV0 725, and the MV1 735.

[0116] In one embodiment, the inter prediction and / or the bidirectional inter prediction may be performed on a per-sample (e.g., per-pixel) basis rather than on a per-block basis. That is, for each sample in the current block 711, a motion vector pointing to each sample in the previous reference block 721 and / or the subsequent reference block 731 can be determined. In such an embodiment, the motion vectors 725 and 735 shown in FIG. 7 represent a plurality of motion vectors corresponding to a plurality of samples in the current block 711, the previous reference block 721, and the subsequent reference block 731.

[0117] In both the merge mode and the advanced motion vector prediction (AMVP) mode, a candidate list is generated by adding candidate motion vectors to the candidate list in an order defined by a candidate list determination pattern. Such candidate motion vectors may include motion vectors according to unidirectional inter prediction 600, bidirectional inter prediction 700, or a combination thereof. Specifically, for adjacent blocks, motion vectors are generated when those blocks are encoded. Such motion vectors are added to the candidate list for the current block, and from the candidate list, a motion vector for the current block is selected. And the motion vector can be signaled as the index of the selected motion vector in the candidate list. The decoder can construct the candidate list using the same process as the encoder and can determine the motion vector selected from the candidate list based on the signaled index. Therefore, the candidate motion vectors include motion vectors generated according to unidirectional inter prediction 600 and / or bidirectional inter prediction 700 depending on which approach was used when such adjacent blocks were encoded.

[0118] Figure 8 shows a video bitstream 800. When used herein, the video bitstream 800 may also be referred to as an encoded video bitstream, a bitstream, or a variation thereof. As shown in Figure 8, the bitstream 800 includes a sequence parameter set (SPS) 802, a picture parameter set (PPS) 804, a slice header 806, and picture data 808.

[0119] The SPS 802 contains data common to all pictures within a sequence of pictures (SOP). In contrast, the PPS 804 contains data common to the entire picture. The slice header 806 contains information about the current slice, such as, for example, the slice type and which of the reference pictures will be used. The SPS 802 and PPS 804 are sometimes generally referred to as parameter sets. The SPS 802, PPS 804, and slice header 806 are types of network abstraction layer (NAL) units. An NAL unit is a syntax structure that contains an indication of the type of data it follows (e.g., coded video data). NAL units are classified into video coding layer (VCL) and non-VCL NAL units. A VCL NAL unit contains data representing the values of samples within a video picture, and a non-VCL NAL unit contains some related additional information such as, for example, parameter sets (important header data applicable to a number of VCL NAL units) and supplementary enhancement information (timing information and other supplementary data that are not necessary to decode the values of samples within a video picture but can enhance the usefulness of the decoded video signal). As those skilled in the art will understand, in actual applications, the bitstream 800 may contain other parameters and information.

[0120] The picture data 808 of FIG. 8 has data related to an encoded or decoded image or video. The picture data 808 may simply be referred to as the payload or data carried within the bitstream 800. In one embodiment, the picture data 808 has a coded video sequence (CVS) 814 (or CLVS) that includes a plurality of pictures 810. The CVS 814 is the coded video sequence for all coded layer video sequences (CLVS) within the video bitstream 800. In particular, when the video bitstream 800 contains a single layer, the CVS and CLVS are the same. The CVS and CLVS differ only when the video bitstream 800 contains multiple layers.

[0121] As shown in FIG. 8, the slices of each picture 810 can be included within its own VCL NAL unit 812. A set of VCL NAL units 812 within the CVS 814 may be referred to as an access unit.

[0122] FIG. 9 shows a partitioning technique 900 for a picture 910. The picture 910 may be similar to any of the pictures 810 in FIG. 8. As shown, the picture 910 can be partitioned into a plurality of slices 912. A slice is a spatially distinct region of a frame (e.g., a picture) that is encoded separately from any other region within the same frame. Three slices 912 are shown in FIG. 9, but in actual applications, a greater or lesser number of slices may be used. Each slice 912 can be partitioned into a plurality of blocks 914. The blocks 914 in FIG. 9 may be similar to the current block 711, the previous reference block 721, and the subsequent reference block 731 in FIG. 7. The blocks 914 may represent CUs. Four blocks 914 are shown in FIG. 9, but in actual applications, a greater or lesser number of blocks may be used.

[0123] Each block 914 can be partitioned into a plurality of samples 916 (e.g., pixels). In one embodiment, the size of each block 914 is measured in luma samples. Sixteen samples 916 are shown in FIG. 9, but in actual applications, a greater or lesser number of samples may be used.

[0124] FIG. 10 is an embodiment of a method 1000 for decoding a coded video bitstream implemented by a video decoder (e.g., video decoder 30). Method 1000 can be executed after the bitstream to be decoded is received directly or indirectly from a video encoder (e.g., video encoder 20). Method 1000 improves the decoding process by enabling selective disabling of BDOF when the spatial resolution of the current picture is different from the spatial resolution of the reference picture, instead of having to disable BDOF for the entire CVS when RPR is enabled. Thus, by having the ability to selectively disable BDOF, the coding efficiency can be improved. Therefore, in practice, the codec performance is improved, which leads to a better user experience.

[0125] In block 1002, the video decoder determines whether the resolution of the current picture being decoded is the same as the resolution of the reference picture specified by the reference picture list. In one embodiment, the video decoder receives a coded video bitstream (e.g., bitstream 800). The coded video bitstream includes a reference picture list, indicates the resolution of the current picture, and indicates the bidirectional inter prediction mode. In one embodiment, the reference picture list structure includes the reference picture list. In one embodiment, the reference picture list is used for bidirectional inter prediction. In one embodiment, the resolution of the current picture is arranged within the parameter set of the coded video bitstream. In one embodiment, the resolution of the reference picture is derived based on the current picture, estimated based on the resolution of the current picture, syntax analyzed from the bitstream, or obtained by other means. In one embodiment, the reference picture for the current picture is generated based on the reference picture list according to the bidirectional inter prediction mode.

[0126] When it is determined in block 1004 that the resolution of the current picture is the same as each resolution of the reference pictures, the video decoder enables BDOF for the current block of the current picture. In one embodiment, the video decoder enables BDOF by setting the BDOF flag to a first value (e.g., true, 1, etc.). In one embodiment, even when BDOF is enabled, BDOF is an optional process. That is, even when BDOF is enabled, it is not necessarily required to execute BDOF.

[0127] When it is determined in block 1006 that the resolution of the current picture is different from any of the resolutions of the reference pictures, the video decoder disables BDOF for the current block of the current picture. In one embodiment, the video decoder disables BDOF by setting the BDOF flag to a second value (e.g., false, 0).

[0128] When the BDOF flag is set to the first value in block 1008, the video decoder refines the motion vector corresponding to the current block. In one embodiment, method 1000 further selectively enables and disables BDOF for other blocks within the current picture according to whether the resolution of the current picture is different from or the same as the resolution of the reference pictures.

[0129] In one embodiment, the method further enables reference picture resampling (RPR) for the entire coded video sequence (CVS) including the current picture even when BDOF is disabled.

[0130] In one embodiment, the current block is obtained from a slice of the current picture. In one embodiment, the current picture has a plurality of slices, and the current block is obtained from a slice among the plurality of slices.

[0131] In one embodiment, an image generated based on a current picture is displayed to a user of an electronics device (e.g., a smartphone, a tablet, a laptop, a personal computer, etc.).

[0132] FIG. 11 is an embodiment of a method 1100 for encoding a video bitstream implemented by a video encoder (e.g., video encoder 20). Method 900 can encode a picture (e.g., from a video) into a video bitstream and be executed when transmitted to a video decoder (e.g., video decoder 30). Method 1100 improves the encoding process by enabling selective disabling of BDOF when the spatial resolution of the current picture is different from the spatial resolution of the reference picture, instead of having to disable BDOF for the entire CVS when RPR is enabled. Thus, by having the ability to selectively disable BDOF, the coding efficiency can be improved. Therefore, in practice, the codec performance is improved, which leads to a better user experience.

[0133] In block 1102, the video encoder determines whether the resolution of the current picture being encoded is the same as the resolution of the reference picture specified by the reference picture list. In one embodiment, the reference picture list structure includes a reference picture list. In one embodiment, the reference picture list is used for bidirectional inter prediction. In one embodiment, the resolution of the current picture is encoded within the parameter set of the video bitstream. In one embodiment, the reference picture for the current picture is generated based on the reference picture list according to the bidirectional inter prediction mode.

[0134] In block 1104, when it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures, the video encoder enables BDOF for the current block of the current picture. In one embodiment, the video encoder enables BDOF by setting a BDOF flag to a first value (e.g., true, 1, etc.). In one embodiment, even when BDOF is enabled, BDOF is an optional process. That is, even when BDOF is enabled, it is not necessarily required to execute BDOF.

[0135] In one embodiment, the method includes determining a motion vector for a current picture based on a reference picture, encoding the current picture based on the motion vector, and decoding the current picture using a hypothetical reference decoder (HRD).

[0136] In block 1106, when the resolution of the current picture is different from the resolution of any of the reference pictures, the video encoder disables BDOF for the current block of the current picture. In one embodiment, the video encoder disables BDOF by setting the BDOF flag to a second value (e.g., false, 0).

[0137] In block 1108, when the BDOF flag is set to the first value, the video encoder refines the motion vector corresponding to the current block. In one embodiment, method 1100 further includes selectively enabling and disabling BDOF for other blocks in the current picture according to whether the resolution of the current picture is different from or the same as the resolution of the reference picture.

[0138] In one embodiment, the method further includes enabling reference picture resampling (RPR) for the entire coded video sequence (CVS) including the current picture even when BDOF is disabled.

[0139] In one embodiment, the current block is obtained from a slice of the current picture. In one embodiment, the current picture has a plurality of slices, and the current block is obtained from a slice among the plurality of slices.

[0140] In one embodiment, a video encoder generates a video bitstream including the current block and transmits the video bitstream towards a video decoder. In one embodiment, the video encoder stores the video bitstream for transmission towards the video decoder.

[0141] In one embodiment, a method for decoding a video bitstream is disclosed. The video bitstream has at least one picture. Each picture has a plurality of slices. Each slice of the plurality of slices has a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists has a plurality of reference pictures that can be used for inter prediction of coding blocks within the slice.

[0142] The method includes parsing a parameter set to obtain resolution information of the current picture, obtaining two reference picture lists of the current slice within the current picture, determining a reference picture for decoding a current coding block within the current slice, determining the resolution of the reference picture, determining whether bidirectional optical flow (BIO) is used or enabled for decoding the current coding block based on the resolutions of the current picture and the reference picture, and decoding the current coding block.

[0143] In one embodiment, the method includes that when the resolutions of the current picture and the reference picture are different, BIO is not used or disabled for decoding the current coding block.

[0144] In one embodiment, a method for decoding a video bitstream is disclosed. The video bitstream has at least one picture. Each picture has a plurality of slices. A header including a plurality of syntax elements is associated with each slice of the plurality of slices. Each slice of the plurality of slices has a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists has a plurality of reference pictures that can be used for inter prediction of coding blocks within the current slice.

[0145] The method includes parsing a parameter set to obtain a flag that defines whether a bidirectional optical flow (BIO) coding tool / technique can be used for decoding pictures in the current coded video sequence, obtaining the current slice within the current picture, and when the value of the flag that defines whether a bidirectional optical flow (BIO) coding tool / technique can be used for decoding pictures in the current coded video sequence defines that BIO can be used, parsing the slice header associated with the current slice to obtain a flag that defines whether the BIO coding tool can be used for decoding coding blocks within the current slice.

[0146] In one embodiment, when the value of the flag that defines whether the BIO coding tool can be used for decoding the current coding block within the current slice defines that the coding tool cannot be used for decoding the current slice, the BIO coding tool is not used or is disabled for decoding the current coding block.

[0147] In one embodiment, when it does not exist, the value of the flag that defines whether the BIO coding tool can be used for decoding the current coding block within the current slice is presumed to be the same as the value of the flag that defines whether a bidirectional optical flow (BIO) coding tool / technique can be used for decoding pictures in the current coded video sequence.

[0148] In one embodiment, a method for encoding a video bitstream is disclosed. The video bitstream has at least one picture. Each picture has a plurality of slices. A header including a plurality of syntax elements is associated with each slice of the plurality of slices. Each slice of the plurality of slices has a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists has a plurality of reference pictures that can be used for inter prediction of coding blocks within the current slice.

[0149] The method includes determining whether a bi-directional optical flow (BIO) coding tool / technique can be used for encoding pictures in the current coded video sequence, parsing a parameter set to obtain resolution information of each picture bitstream, obtaining two reference picture lists of the current slice in the current picture, parsing the reference picture list of the current slice to obtain active reference pictures that can be used for decoding coding blocks of the current slice, and restricting that the BIO coding tool cannot be used for encoding coding blocks within the current slice when at least one of the following conditions is satisfied: the condition that the BIO coding tool cannot be used for encoding pictures in the current coded video sequence, and the condition that the resolution of the current picture is different from the resolution of at least one of the reference pictures.

[0150] In one embodiment, a method for decoding a video bitstream is disclosed. The bitstream has at least one picture. Each picture has a plurality of slices. A header containing a plurality of syntax elements is associated with each slice of the plurality of slices. Each slice of the plurality of slices has a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists has a plurality of reference pictures that can be used for inter prediction of coding blocks within the current slice.

[0151] The method includes parsing a parameter set to obtain a flag that defines whether a bi-directional optical flow (BIO) coding tool / technique can be used for decoding pictures within the current coded video sequence, and parsing the parameter set to obtain a flag that defines whether a bi-directional optical flow (BIO) coding tool / technique can be used for decoding a picture that references a parameter set that is a picture parameter set (PPS).

[0152] In one embodiment, when the value of the flag that defines whether the BIO coding tool / technique can be used for decoding a picture that references the PPS defines that the coding tool cannot be used, the BIO coding tool is not used or is disabled for decoding the current coding block.

[0153] In one embodiment, a method for encoding a video bitstream is disclosed. The video bitstream has at least one picture. Each picture has a plurality of slices. A header containing a plurality of syntax elements is associated with each slice of the plurality of slices. Each slice of the plurality of slices has a plurality of coding blocks and a plurality of reference picture lists. Each reference picture list of the plurality of reference picture lists has a plurality of reference pictures that can be used for inter prediction of coding blocks within the current slice.

[0154] In one embodiment, the method includes determining whether a bidirectional optical flow (BIO) coding tool / technique can be used for coding a picture in the currently coded video sequence, determining whether a bidirectional optical flow (BIO) coding tool / technique can be used for coding a picture that references the current PPS, and restricting that the BIO coding tool cannot be used for coding a picture that references the current PPS when the BIO coding tool cannot be used for coding a picture in the currently coded sequence.

[0155] To implement the embodiments disclosed herein, the following syntax and semantics may be used. The following description is in comparison with the base text of the latest VVC draft specification. In other words, only the differences are described, and the text in the base text not mentioned below is applied as it is. The text added to the base text is shown in bold, and the text to be deleted is shown in italics.

[0156] Update the reference picture list construction process as follows.

[0157] The reference picture lists RefPicList[0] and RefPicList[1] are constructed as follows: (Outer 1) TIFF2025087712000002.tif195170

[0158] Derivation of a flag for determining whether BIO is used.

[0159] predSamplesL0 L 、predSamplesL1 L 、and predSamplesIntra L are assumed to be (cbWidth)×(cbHeight) arrays of predicted luma sample values, and predSamplesL0 Cb 、predSamplesL1 Cb 、predSamplesL0Cr and predSamplesL1 Cr predSamplesIntra Cb and predSamplesIntra Cr are assumed to be (cbWidth / 2)×(cbHeight / 2) arrays of predicted chroma sample values. (Outside 2) TIFF2025087712000003.tif91170

[0160] Sequence parameter set syntax and semantics. (Outside 3) TIFF2025087712000004.tif47170

[0161] The sps_bdof_enabled_flag equal to 0 specifies that the bidirectional optical flow inter prediction is disabled. The sps_bdof_enabled_flag equal to 1 specifies that the bidirectional optical flow inter prediction is enabled.

[0162] Slice header syntax and semantics. (Outside 4) TIFF2025087712000005.tif55170

[0163] (Outside 5) TIFF2025087712000006.tif24170

[0164] Derivation of a flag that determines whether BIO is used.

[0165] Assume that predSamplesL0L, predSamplesL1L, and predSamplesIntraL are arrays of (cbWidth)×(cbHeight) for predicted luma sample values, and predSamplesL0Cb, predSamplesL1Cb, predSamplesL0Cr, and predSamplesL1Cr, predSamplesIntraCb, and predSamplesIntraCr are arrays of (cbWidth / 2)×(cbHeight / 2) for predicted chroma sample values. (Outer 6) TIFF2025087712000007.tif85170

[0166] Sequence parameter set syntax and semantics. (Outer 7) TIFF2025087712000008.tif45170

[0167] The sps_bdof_enabled_flag equal to 0 specifies that the bidirectional optical flow inter prediction is disabled. The sps_bdof_enabled_flag equal to 1 specifies that the bidirectional optical flow inter prediction is enabled.

[0168] Picture parameter set syntax and semantics. (Outer 8) TIFF2025087712000009.tif46170

[0169] (Outer 9) TIFF2025087712000010.tif26170

[0170] (Outer 10) TIFF2025087712000011.tif13170

[0171] Derivation of the flag for determining whether BIO is used.

[0172] predSamplesL0L 、predSamplesL1 L 、and predSamplesIntra L are assumed to be (cbWidth)×(cbHeight) arrays of predicted luma sample values, and predSamplesL0 Cb 、predSamplesL1 Cb 、predSamplesL0 Cr 、and predSamplesL1 Cr 、predSamplesIntra Cb 、and predSamplesIntra Cr are assumed to be (cbWidth / 2)×(cbHeight / 2) arrays of predicted chroma sample values. (External 11) TIFF2025087712000012.tif85170

[0173] FIG. 12 is a schematic diagram of a video coding apparatus 1200 (e.g., video encoder 20 or video decoder 30) according to an embodiment of the disclosure. The video coding apparatus 1200 is suitable for implementing the disclosure embodiments described herein. The video coding apparatus 1200 includes an input port 1210 and a receiver unit (Rx) 1220 for receiving data, a processor, logic unit, or central processing unit (CPU) 1230 for processing data, a transmitter unit (Tx) 1240 and an output port 1250 for transmitting data, and a memory 1260 for storing data. The video coding apparatus 1200 may also have optical - electrical (OE) components and electrical - optical (EO) components coupled to the input port 1210, the receiver unit 1220, the transmitter unit 1240, and the output port 1250 for an exit or entry of optical or electrical signals.

[0174] Processor 1230 is implemented by hardware and software. Processor 1230 can be implemented as one or more of a CPU chip, cores (e.g., as a multi-core processor), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and a digital signal processor (DSP). Processor 1230 communicates with an input port 1210, a receiver unit 1220, a transmitter unit 1240, an output port 1250, and a memory 1260. Processor 1230 has a coding module 1270. The coding module 1270 implements the disclosed embodiments described above. For example, the coding module 1270 implements, processes, prepares, or provides various codec functions. Including the coding module 1270 thus provides a substantial improvement to the functionality of the video coding apparatus 1200 and enables the conversion of the video coding apparatus 1200 to different states. Alternatively, the coding module 1270 is implemented as instructions stored in the memory 1260 and executed by the processor 1230.

[0175] The video coding apparatus 1200 may also include an input and / or output (I / O) device 1280 for communicating with the user and for data. The I / O device 1280 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 1280 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0176] Memory 1260 has one or more disks, tape drives, and solid state drives, and is used as an overflow data storage device to store such programs when a program is selected for execution, and can store instructions and data read during program execution. Memory 1260 can be volatile and / or non-volatile, and can be a read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).

[0177] Figure 13 is a schematic diagram of an embodiment of means 1300 for coding. In one embodiment, means 1300 for coding is implemented in a video coding device 1302 (e.g., video encoder 20 or video decoder 30). The video coding device 1302 includes receiving means 1301. The receiving means 1301 is configured to receive a picture to be encoded or a bitstream to be decoded. The video coding device 1302 includes transmitting means 1307 coupled to the receiving means 1301. The transmitting means 1307 is configured to transmit a bitstream to a decoder or to transmit a decoded image to a display means (e.g., one of the I / O devices 1280).

[0178] The video coding device 1302 includes storage means 1303. The storage means 1303 is coupled to at least one of the receiving means 1301 or the transmitting means 1307. The storage means 1303 is configured to store instructions. The video coding device 1302 also includes processing means 1305. The processing means 1305 is coupled to the storage means 1303. The processing means 1305 is configured to execute instructions stored in the storage means 1303 to execute the methods disclosed herein.

[0179] It should also be understood that the steps of the exemplary methods described herein need not necessarily be executed in the order described, and that the order of such method steps should be understood to be merely exemplary. Similarly, additional steps may be included in such methods, and in methods consistent with various embodiments of the present disclosure, certain steps may be omitted or combined.

[0180] Although several embodiments have been presented in the present disclosure, it should be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The examples herein are not limiting but should be regarded as exemplary, and the intention is that the disclosure should not be limited to the details given herein. For example, these various elements or components may be combined or integrated in other systems, or certain mechanisms may be omitted or not implemented.

[0181] Also, the technologies, systems, subsystems, and methods described and illustrated as separate or discrete in various embodiments may be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of the present disclosure. Other items shown or described as being coupled, directly or indirectly, or communicating with each other may be indirectly coupled or communicate through some interface, device, or intermediate component, regardless of whether they are electrical, mechanical, or otherwise. Other examples of variations, substitutions, and modifications can be elucidated by those skilled in the art and made without departing from the spirit and scope disclosed herein.

Claims

1. 1. A method for encoding a current picture into a video bitstream, comprising: determining whether a resolution of the current picture is the same as a resolution of a reference picture identified in a reference picture list associated with the current picture; when it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures, enabling bidirectional optical flow (BDOF) for the current block of the current picture; disabling the BDOF for the current block of the current picture when it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; performing inter prediction based on a reference block of the reference picture to obtain a predicted block for the current block; obtaining a residual block based on the predicted block and the current block; applying a transform and quantization to the residual block to obtain quantized residual transform coefficients; encoding the quantized residual transform coefficients into the video bitstream; and The method according to claim 1,

2. encoding sps_bdof_enabled_flag into the video bitstream, where sps_bdof_enabled_flag equal to 0 specifies that the BDOF inter prediction is disabled and sps_bdof_enabled_flag equal to 1 specifies that the BDOF inter prediction is enabled; The method of claim 1 further comprising:

3. The method of claim 1 , wherein the resolution of the current picture is represented by PicWidthInSamplesY and PicHeightInSamplesY of the current picture.

4. The method of claim 1 , wherein the current picture comprises a plurality of slices, and the current block is included in one of the slices.

5. 1. A method of decoding implemented by a video decoder, comprising: performing an entropy decoding process, an inverse quantization process, and an inverse transform process on the received encoded bitstream to obtain a reconstructed residual block; determining whether a resolution of a current picture being decoded is the same as a resolution of a reference picture identified by a reference picture list associated with the current picture; when the resolution of the current picture is determined to be the same as the resolution of each of the reference pictures, enabling bidirectional optical flow (BDOF) for the current block of the current picture; disabling the BDOF for the current block of the current picture when it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; Obtaining a reconstructed block based on the reconstructed residual block and a prediction block, the prediction block being obtained based on a reference block of the reference picture and the current block; The method according to claim 1,

6. The method of claim 5 , wherein the bitstream further includes sps_bdof_enabled_flag, where sps_bdof_enabled_flag equal to 0 specifies that the BDOF inter prediction is disabled and sps_bdof_enabled_flag equal to 1 specifies that the BDOF inter prediction is enabled.

7. 7. The method of claim 5, wherein the resolution of the current picture is represented by PicWidthInSamplesY and PicHeightInSamplesY of the current picture.

8. The method according to claim 5 , wherein the current picture comprises a number of slices, and the current block is included in one of the slices.

9. 1. An encoding device, comprising: A memory storing instructions; a processor coupled to the memory and configured to implement the instructions to cause the encoding device to: determining whether a resolution of a current picture is the same as a resolution of a reference picture identified in a reference picture list associated with the current picture; when it is determined that the resolution of the current picture is the same as the resolution of each of the reference pictures, enabling bidirectional optical flow (BDOF) for the current block of the current picture; when it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures, invalidating the BDOF for the current block of the current picture; performing inter prediction based on a reference block of the reference picture to obtain a predicted block for the current block; obtaining a residual block based on the predicted block and the current block; applying a transform and a quantization to the residual block to obtain quantized residual transform coefficients; encoding the quantized residual transform coefficients into the video bitstream. a processor configured to a transmitter coupled to the processor, the transmitter configured to transmit a video bitstream including the current block to a video decoder; An encoding device having the following construction.

10. The encoding apparatus of claim 9 , wherein when the BDOF is disabled, reference picture resampling (RPR) is enabled for an entire coded video sequence (CVS) including the current picture.

11. The encoding device according to claim 9 or 10, wherein the memory is configured to store the video bitstream before the transmitter transmits the video bitstream towards the video decoder.

12. A decoding device, comprising: a receiver configured to receive a coded video bitstream; a memory coupled to the receiver, the memory storing instructions; a processor coupled to the memory for executing the instructions to cause the decoding device to: performing an entropy decoding process, an inverse quantization process, and an inverse transform process on the received encoded bitstream to obtain a reconstructed residual block; determining whether a resolution of a current picture being decoded is the same as a resolution of a reference picture identified by a reference picture list associated with the current picture; when the resolution of the current picture is determined to be the same as the resolution of each of the reference pictures, enabling bidirectional optical flow (BDOF) for the current block of the current picture; when it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures, invalidating the BDOF for the current block of the current picture; Obtaining a reconstructed block based on the reconstructed residual block and a prediction block, the prediction block being obtained based on a reference block of the reference picture and the current block; a processor configured to A decoding device having the above configuration.

13. The decoding apparatus of claim 12 , wherein when the BDOF is disabled, reference picture resampling (RPR) is enabled for an entire coded video sequence (CVS) including the current picture.

14. 14. A decoding device according to claim 12 or 13, further comprising a display arranged to display an image generated based on the current block.

15. 9. A video coding device, the device having a processor and a memory storing instructions which, when executed by the processor, configure the device to perform a method according to any one of claims 1 to 4 or to perform a method according to any one of claims 5 to 8.

16. An encoder; a decoder in communication with the encoder; wherein the encoder comprises an encoding device according to any one of claims 9 to 11 and the decoder comprises a decoding device according to any one of claims 12 to 14. system.

17. 1. A processor-implemented method for storing a video bitstream, comprising: receiving an encoded bitstream; receiving the bitstream comprising data and syntax elements representing a current block of a current picture, the syntax elements comprising a parameter set, a motion vector, an intra mode indicator, and partition information, the parameter set comprising resolution information of the current picture, the resolution information of the current picture indicating a resolution of the current picture, wherein bidirectional optical flow (BDOF) is enabled for the current block when the resolution of the current picture is the same as a resolution of each of the reference pictures identified in a reference picture list associated with the current picture, and wherein the BDOF is disabled for the current block when it is determined that the resolution of the current picture differs from the resolution of any of the reference pictures; storing the encoded bitstream on a storage medium; The method according to claim 1,

18. 1. A system for storing an encoded bitstream of video data, comprising: a receiver configured to receive the encoded bitstream, a receiver, wherein the encoded bitstream comprises data and syntax elements representing a current block of a current picture, the syntax elements comprising a parameter set, a motion vector, an intra mode indicator, and partition information, the parameter set comprising resolution information of the current picture, the resolution information of the current picture indicating a resolution of the current picture, wherein bidirectional optical flow (BDOF) is enabled for the current block when the resolution of the current picture is the same as a resolution of each of the reference pictures identified in a reference picture list associated with the current picture, and wherein the BDOF is disabled for the current block when it is determined that the resolution of the current picture is different from the resolution of any of the reference pictures; a storage medium configured to store the encoded bitstream; and A system having

19. 1. A method for transmitting an encoded bitstream of video data, comprising the steps of: Obtaining a bitstream from a storage medium; transmitting the bitstream; transmitting the bitstream comprising data and syntax elements representing a current block of a current picture, the syntax elements comprising a parameter set, a motion vector, an intra mode indicator, and partition information, the parameter set comprising resolution information of the current picture, the resolution information of the current picture indicating a resolution of the current picture, wherein bidirectional optical flow (BDOF) is enabled for the current block when the resolution of the current picture is the same as a resolution of each of the reference pictures identified in a reference picture list associated with the current picture, and wherein the BDOF is disabled for the current block when it is determined that the resolution of the current picture differs from the resolution of any of the reference pictures; The method according to claim 1,

20. 1. A system for transmitting an encoded bitstream of video data, comprising: a receiver configured to obtain the bitstream from the storage medium; a transmission unit configured to transmit the bitstream, a transmission unit, the bitstream comprising data and syntax elements representing a current block of a current picture, the syntax elements comprising a parameter set, a motion vector, an intra mode indicator, and partition information, the parameter set comprising resolution information of the current picture, the resolution information of the current picture indicating a resolution of the current picture, wherein bidirectional optical flow (BDOF) is enabled for the current block when the resolution of the current picture is the same as a resolution of each of the reference pictures identified in a reference picture list associated with the current picture, and wherein the BDOF is disabled for the current block when it is determined that the resolution of the current picture differs from the resolution of any of the reference pictures; A system having

21. A computer readable storage medium having stored thereon a computer program executable by a processor, the computer program causing the processor to perform the method according to any one of claims 1 to 8 when executed by the processor.

22. A program having a program code for performing the method according to any one of claims 1 to 8 when the program is run on a computer or processor.

Citation Information

Patent Citations

  • Improved BI-directional optical flow for video coding

    WO2017058899A1

  • Signaling of adaptive picture size in video bitstream

    WO2020185814A1

  • Selective use of coding tools in video processing

    WO2020228660A1

  • Interplay between reference picture resampling and video coding tools

    WO2021073488A1