Video coding method, code stream transmission or storage method, equipment and storage medium

By introducing a joint incremental motion vector signaling mechanism, the signaling method of inter prediction mode is optimized, and the problem of low signaling efficiency of inter prediction mode and motion vector difference in existing video encoding technology is solved, video compression efficiency is improved, bandwidth and storage requirements are reduced, and encoding efficiency of high-resolution and high-frame rate videos is improved.

CN120455661APending Publication Date: 2025-08-08TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510888913.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-03-22
Filing Date
2022-04-15
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the existing video encoding technology, the inter-frame prediction mode and motion vector difference signaling efficiency are low, resulting in insufficient video compression efficiency, especially in the transmission and storage of high-resolution and high-frame rate video data, with high bandwidth and storage requirements.

Method used

The joint incremental motion vector (joint_delta_mv) signaling mechanism is adopted, and the inter prediction mode and incremental motion vector of the current block are extracted, and the decoding process of the current block is derived using the joint incremental motion vector to optimize the signaling method of the inter prediction mode.

Benefits of technology

It improves the compression efficiency of video encoding, reduces bandwidth and storage requirements, especially in the transmission and storage of high-resolution and high-frame rate video data, and improves the encoding efficiency and decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455661A_ABST
    Figure CN120455661A_ABST
Patent Text Reader

Abstract

The invention relates to a video encoding and decoding method and apparatus, and a storage medium. The method comprises the following steps: receiving an encoded video code stream; extracting an inter-frame prediction mode and a joint incremental motion vector (MV) of a current block in a current frame from the coded video code stream; extracting a first flag from the encoded video stream, the first flag indicating whether a first increment MV of a first reference frame in a reference list 0 and a second increment MV of a second reference frame in a reference list 1 are jointly signaled; in response to the first flag indicating that the first delta MV and the second delta MV are jointly signaled, deriving the first delta MV and the second delta MV based on the joint delta MV; and decoding the current block in the current frame based on the first increment MV and the second increment MV.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by reference

[0001] This application is based upon and claims the benefit of U.S. Provisional Application No. 63 / 245,655, filed on September 17, 2021, which is hereby incorporated by reference in its entirety. This application is also based upon and claims the benefit of U.S. Non-Provisional Application No. 17 / 700,745, filed on March 22, 2022, which is hereby incorporated by reference in its entirety. Technical Field

[0002] The present disclosure relates to video encoding and / or decoding techniques, and in particular to an improved design and signaling of joint motion vector differences for encoding and / or decoding. Background Art

[0003] The background description provided herein is intended to present the background of the present application as a whole. To the extent that the work of the presently named inventors is described in the background section and in various aspects of this specification, it is not intended that it be prior art at the time of filing this application, and it is neither expressly nor impliedly admitted that it is prior art to the present application.

[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. An uncompressed digital video may comprise a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated fully sampled or subsampled chrominance samples. The series of pictures has a fixed or variable picture rate (or frame rate), for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and chrominance subsampling of 4:2:0, at 8 bits per pixel per color channel requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.

[0005] One goal of video encoding and decoding is to reduce redundant information in an uncompressed input video signal through compression. Video compression can help reduce the aforementioned bandwidth and / or storage requirements, in some cases by two or more orders of magnitude. Both lossless and lossy compression, as well as combinations of the two, can be employed. Lossless compression refers to techniques that reconstruct an exact replica of the original signal from a compressed original signal through the decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not completely preserved during encoding and cannot be fully recovered during decoding. When using lossy compression, the reconstructed signal may differ from the original, but the distortion between the original and the reconstructed signal is small enough to make the reconstructed signal usable for the intended application, despite some information loss. For video, lossy compression is widely used in many applications. The amount of tolerable distortion depends on the application. For example, users of some consumer video streaming applications may tolerate higher distortion than users of film or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows encoding algorithms that produce higher loss and higher compression ratios.

[0006] Video encoders and decoders may utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transforms, quantization, and entropy coding.

[0007] Video codec techniques may include known intra-frame coding techniques. In intra-frame coding, sample values are represented without reference to samples or other data of a previously reconstructed reference picture. In some video codecs, a picture is spatially subdivided into blocks of samples. When all sample blocks are encoded in intra-frame mode, the picture may be referred to as an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and can therefore be used as the first picture in a coded video bitstream and video session, or as a still image. The samples of the intra-frame predicted block can then be transformed into the frequency domain, and the transform coefficients thus generated can be quantized before entropy coding. Intra-frame prediction represents a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the DC value after transformation and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding at a given quantization step size.

[0008] Conventional intra-frame coding, as known from coding techniques such as MPEG-2, does not use intra-frame prediction. However, some newer video compression techniques include attempts to encode / decode blocks based on, for example, surrounding sample data and / or metadata obtained during spatially adjacent encoding and / or decoding and preceding the data block being intra-coded or decoded in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and not reference data from other reference pictures.

[0009] There can be many different forms of intra-frame prediction. When more than one such technique is available in a given video coding technology, the technique used may be referred to as an intra-frame prediction mode. One or more intra-frame prediction modes may be provided in a particular codec. In some cases, a mode may have sub-modes and / or may be associated with various parameters, and the mode / sub-mode information and intra-frame coding parameters for a video block may be contained in a mode codeword, which may be encoded separately or together. For a given mode, sub-mode and / or parameter combination, which codeword is used can have an impact on the coding efficiency gain through intra-frame prediction, and the same is true for the entropy coding technique used to convert the codeword into the codestream.

[0010] Some modes of intra prediction were introduced with H.264, revised in H.265, and further revised in newer coding techniques such as Joint Detection Mode (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Typically, for intra prediction, the values of neighboring samples that have become available can be used to form a predictor block. For example, the available values of a specific set of neighboring samples along a specific direction and / or row can be copied into the predictor block. A reference to the direction used can be encoded in the bitstream or can itself be predicted.

[0011] refer to Figure 1A , depicted at the bottom right is a subset of the 9 predictor directions specified in H.265's 33 possible intra predictor directions (corresponding to the 33 angular modes of the 35 intra modes specified in H.265). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the directions according to which the sample at 101 is predicted using neighboring samples. For example, arrow (102) indicates that sample (101) is predicted based on one or more neighboring samples to the upper right that are at an angle of 45 degrees to the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more neighboring samples to the lower left of sample (101) that are at an angle of 22.5 degrees to the horizontal.

[0012] Still refer to Figure 1A, a square block (104) consisting of 4×4 samples is shown in the upper left (indicated by the thick dashed line). The square block (104) consists of 16 samples, each of which is labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is a 4×4 size sample, S44 is located in the lower right corner. Example reference samples following a similar numbering scheme are also shown. Reference samples are labeled with "R" and their Y position (e.g., row index) and X position (e.g., column index) relative to the block (104). In H.264 and H.265, neighboring prediction samples that are adjacent to the block being reconstructed are used.

[0013] The intra-picture prediction for block 104 can begin by copying reference sample values from neighboring samples according to a signaled prediction direction. For example, assume that the encoded video stream includes signaling indicating the prediction direction of arrow (102) for this block 104 - that is, predicting samples based on one or more prediction samples to the upper right at a 45-degree angle to the horizontal. In such a case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Sample S44 is then predicted based on reference sample R08.

[0014] In some cases, such as by interpolation, the values of multiple reference samples may be combined in order to calculate the reference sample, particularly when the direction is not divisible by 45 degrees.

[0015] As video coding technology continues to develop, the number of possible directions increases. For example, in H.264 (2003), 9 different directions can be used for intra prediction. This has increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of this disclosure. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and certain techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, accepting some bit cost for the direction. In addition, the direction itself can sometimes be predicted based on the adjacent directions used for intra prediction of adjacent blocks that have already been decoded.

[0016] Figure 1B A diagram (180) depicting 65 intra prediction directions according to the JEM is shown to illustrate the increase in the number of prediction directions in various coding techniques over time.

[0017] The manner in which the bits representing the intra-frame prediction direction are mapped to the prediction direction in the coded video stream can vary between different video coding techniques and can range from simple direct mappings of prediction directions to intra-frame prediction modes, to codewords, to complex adaptive schemes involving most probable modes, and similar techniques. However, in all cases, there may be certain directions for intra-frame prediction that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, those less probable directions will be represented by a larger number of bits than the more probable directions.

[0018] Inter-picture prediction or inter-frame prediction can be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or portion thereof (reference picture) is spatially shifted in the direction indicated by a motion vector (hereinafter MV) and can be used to predict a newly reconstructed picture or picture portion (e.g., block). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, where the third dimension is an indication of the reference picture in use (approximately the temporal dimension).

[0019] In some video compression techniques, the current MV applicable to a region of sample data can be predicted from other MVs, such as those related to other regions of sample data spatially adjacent to the region being reconstructed and preceding the current MV in decoding order. This can significantly reduce the total amount of data required to encode the MV by removing redundancy in related MVs, thereby increasing compression efficiency. MV prediction can be performed efficiently, for example, because when encoding an input video signal derived from a camera (referred to as natural video), there is a statistical probability that regions larger than the region for which a single MV applies will move in similar directions within the video sequence. Therefore, in some cases, similar motion vectors derived from MVs in neighboring regions can be used for prediction. This results in the actual MV for a given region being similar or identical to the MV predicted from surrounding MVs. After entropy coding, such MVs can be represented using fewer bits than would be used if the MV were encoded directly rather than predicted from one or more neighboring MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, MV prediction can itself be lossy, for example due to rounding errors when calculating predicted values from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding," December 2016). Among the various MV prediction mechanisms specified in H.265, this application describes a technique referred to below as "spatial merging."

[0021] Please refer to Figure 2 , the current block (201) includes samples that have been discovered by the encoder during the motion search process, and the samples can be predicted based on the previous block of the same size that has been spatially offset. In addition, the MV can be derived from metadata associated with one or more reference pictures instead of encoding the MV directly. For example, the MV associated with any of the five surrounding samples A0, A1 and B0, B1, B2 (corresponding to 202 to 206) is used to derive the MV from the metadata of the nearest reference picture (in decoding order). In H.265, MV prediction can use the prediction value of the same reference picture also used by the neighboring blocks. Summary of the Invention

[0022] This disclosure describes various embodiments of methods, apparatus, and storage media for video encoding and decoding.

[0023] According to one aspect, an embodiment of the present disclosure provides a method for video decoding. The method includes receiving an encoded video stream. The method also includes extracting an inter-frame prediction mode and a joint incremental motion vector MV of a current block in a current frame from the encoded video stream; extracting a first flag from the encoded video stream, the first flag indicating whether a first incremental MV of a first reference frame in reference list 0 and a second incremental MV of a second reference frame in reference list 1 are jointly signaled; in response to the first flag indicating that the first incremental MV and the second incremental MV are jointly signaled, deriving the first incremental MV and the second incremental MV based on the joint incremental MV; and decoding the current block in the current frame based on the first incremental MV and the second incremental MV. The present application also discloses a method for video encoding, comprising: receiving an encoded video stream; signaling, in the encoded video stream, an inter-frame prediction mode and a joint incremental motion vector MV (joint_delta_mv) of a current block in a current frame; signaling, in the encoded video stream, a first flag (joint_mvd_flag), wherein the first flag (joint_mvd_flag) indicates whether a first incremental MV of a first reference frame in reference list 0 and a second incremental MV of a second reference frame in reference list 1 are jointly signaled; wherein, when the first flag indicates that the first incremental MV and the second incremental MV are jointly signaled, the joint incremental MV is used to derive the first incremental MV and the second incremental MV, and the first incremental MV and the second incremental MV are used to decode the current block in the current frame.

[0024] According to another aspect, embodiments of the present disclosure provide a video encoding and / or decoding apparatus. The apparatus includes a memory storing instructions; and a processor in communication with the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above-described method for video decoding and / or encoding.

[0025] Aspects of the present disclosure also provide a video encoding or decoding device or apparatus, comprising a circuit configured to perform any of the above-mentioned method implementations.

[0026] In another aspect, an embodiment of the present disclosure provides a non-volatile computer-readable medium storing instructions that, when executed by a computer for video encoding and decoding, cause the computer to perform the above video encoding and decoding method. A video encoding method is also provided, comprising: determining whether to apply an inter-frame prediction mode and an incremental motion vector for a current block in a current frame of video data; generating a first flag joint_mvd_flag, the first flag joint_mvd_flag indicating whether to jointly signal a first incremental MV of a first reference frame in reference list 0 and a second incremental MV of a second reference frame in reference list 1 for the current block; and encoding the first flag joint_mvd_flag in a bitstream of the video data. A method for storing a video bitstream is also provided, executing a video encoding method to generate a video bitstream, and storing the video bitstream. A method for transmitting a video bitstream is also provided, executing a video encoding method to generate a video bitstream, and transmitting the video bitstream. A method for processing a video block is also provided, comprising converting the video block into a video code stream, wherein the video code stream includes: a first indication for determining whether to apply an inter-frame prediction mode and an incremental motion vector to the video block; a first flag joint_mvd_flag, wherein the first flag joint_mvd_flag indicates whether to jointly signal a first incremental MV of a first reference frame in reference list 0 and a second incremental MV of a second reference frame in reference list 1 for the current block; wherein, when the first flag joint_mvd_flag indicates that the first incremental MV and the second incremental MV are jointly signaled, the video code stream also includes a joint incremental MV for the first incremental MV and the second incremental MV.

[0027] The above aspects and other aspects and embodiments thereof are described in more detail in the drawings, the description and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.

[0029] Figure 1A A schematic diagram showing an exemplary subset of intra prediction direction modes.

[0030] Figure 1B A diagram showing exemplary intra prediction directions is shown.

[0031] Figure 2 A schematic diagram showing spatial merging candidates of a current block and its surroundings for motion vector prediction in one example is shown.

[0032] Figure 3 A schematic diagram illustrating a simplified block diagram of a communication system according to an example embodiment.

[0033] Figure 4 A schematic diagram illustrating a simplified block diagram of a communication system according to an example embodiment.

[0034] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment.

[0035] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment.

[0036] Figure 7 A block diagram of a video encoder according to another example embodiment is shown.

[0037] Figure 8 A block diagram of a video decoder according to another example embodiment is shown.

[0038] Figure 9 A scheme of coding block partitioning according to an exemplary embodiment of the present disclosure is shown.

[0039] Figure 10 Another scheme of coding block partitioning according to an example embodiment of the present disclosure is shown.

[0040] Figure 11 Another scheme of coding block partitioning according to an example embodiment of the present disclosure is shown.

[0041] Figure 12 An example of dividing a basic block into coding blocks according to an example partitioning scheme is shown.

[0042] Figure 13 An example ternary partitioning scheme is shown.

[0043] Figure 14 An example quadtree binary tree coding block partitioning scheme is shown.

[0044] Figure 15 A scheme for partitioning a coding block into multiple transform blocks and an encoding order of the transform blocks according to an example embodiment of the present disclosure is shown.

[0045] Figure 16 Another scheme for partitioning a coding block into multiple transform blocks and an encoding order of the transform blocks according to an example embodiment of the present disclosure is shown.

[0046] Figure 17 Another scheme for partitioning a coding block into multiple transform blocks according to an example embodiment of the present disclosure is shown.

[0047] Figure 18 A flow chart of a method according to an example embodiment of the present disclosure is shown.

[0048] Figure 19 A schematic diagram of a computer system according to an example embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0049] The present invention will now be described in detail hereinafter with reference to the accompanying drawings, which form a part hereof and which illustrate, by way of illustration, specific examples of embodiments. However, it is noted that the present invention may be embodied in a variety of different forms, and thus, covered or claimed subject matter is intended to be construed as not being limited to any of the embodiments set forth below. It is also noted that the present invention may be embodied as a method, apparatus, component, or system. Thus, embodiments of the present invention may, for example, take the form of hardware, software, firmware, or any combination thereof.

[0050] Throughout the specification and claims, terms may have subtle meanings that are suggested or implied by the context beyond their explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Likewise, the phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include all or part of any combination of exemplary embodiments / embodiments.

[0051] In general, terms can be understood at least in part based on their use in the context. For example, terms such as "and", "or", or "and / or" as used herein can include multiple meanings, which can depend at least in part on the context in which such terms are used. Typically, "or", if used in an association list such as A, B, or C, is intended to mean A, B, and C, which are used herein in an inclusive sense, and A, B, or C, which are used herein in an exclusive sense. In addition, the terms "one or more" or "at least one" as used herein, at least in part depending on the context, can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as "a", "an," or "the" can also be understood to express singular usage or plural usage, which depends at least in part on the context. In addition, the term "based on" or "determined by..." can be understood to not necessarily be intended to express an exclusive set of factors, but can allow the presence of additional factors that are not necessarily explicitly described again, which depends at least in part on the context.

[0052] Figure 33 is a simplified block diagram of a communication system (300) according to an embodiment disclosed herein. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first terminal device (310) and a second terminal device (320) interconnected via the network (350). Figure 3 In an embodiment, the first terminal device (310) and the second terminal device (320) perform unidirectional data transmission. For example, the first terminal device (310) may encode video data (e.g., a video picture stream captured by the first terminal device (310)) for transmission to the second terminal device (320) via the network (350). The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to restore the video data, and display the video picture based on the restored video data. Unidirectional data transmission is more common in applications such as media services.

[0053] In another embodiment, a communication system (300) includes a third terminal device (330) and a fourth terminal device (340) that perform bidirectional transmission of encoded video data, which can be implemented, for example, during a video conference. For bidirectional data transmission, each of the third terminal device (330) and the fourth terminal device (340) can encode video data (e.g., a video picture stream collected by the terminal device) for transmission to the other of the third terminal device (330) and the fourth terminal device (340) via a network (350). Each of the third terminal device (330) and the fourth terminal device (340) can also receive the encoded video data transmitted by the other of the third terminal device (330) and the fourth terminal device (340), and can decode the encoded video data to restore the video data, and can display the video picture on an accessible display device based on the restored video data.

[0054] exist Figure 3In the embodiment of the present invention, the first terminal device (310), the second terminal device (320), the third terminal device (330) and the fourth terminal device (340) may be servers, personal computers and smart phones, but the scope of application of the basic principles disclosed in this application is not limited thereto. The embodiments disclosed in this application are applicable to desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. The network (350) represents any number or type of network that transmits encoded video data between the first terminal device (310), the second terminal device (320), the third terminal device (330) and the fourth terminal device (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit switching, packet switching and / or other types of channels. The network may include a telecommunications network, a local area network, a wide area network and / or the Internet. For the purposes of this application, unless otherwise explicitly explained herein, the architecture and topology of the network (350) may be irrelevant to the operations disclosed in this application.

[0055] As an example, Figure 4 The video encoder and video decoder are shown in a video streaming environment. The subject matter disclosed in this application is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, compressed video storage on digital media including CDs, DVDs, memory sticks, etc.

[0056] The video streaming system may include an acquisition subsystem (413) that may include a video source (401), such as a digital camera, to create, for example, an uncompressed video picture or image stream (402). In an embodiment, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. The uncompressed video picture stream (402) is depicted as a thick line to emphasize the high data volume of the video picture stream compared to the encoded video data (404) (or encoded video bitstream), and the video picture stream (402) may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or embody various aspects of the disclosed subject matter as described in more detail below. Compared to the uncompressed video picture stream (402), the encoded video data (404) (or the encoded video code stream (404)) is depicted as a thin line to emphasize the lower amount of data of the encoded video data (404) (or the encoded video code stream (404)), which can be stored on the streaming server (405) for future use or directly used by downstream video devices (not shown). One or more streaming client subsystems, such as Figure 4The client subsystem (406) and the client subsystem (408) in the streaming server (405) can access the streaming server (405) to retrieve the copy (407) and the copy (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in the electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an uncompressed output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or another presentation device (not depicted). The video decoder 410 can be configured to perform some or all of the various functions described in the present disclosure. In some streaming systems, the encoded video data (404), the video data (407), and the video data (409) (e.g., a video bitstream) can be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T H.265. In an embodiment, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application can be used in the context of the VVC standard and other video coding standards.

[0057] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0058] Figure 5 1 is a block diagram of a video decoder (510) according to an embodiment disclosed below in the present application. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 A video decoder (410) of an embodiment.

[0059] A receiver (531) may receive one or more encoded video sequences to be decoded by a video decoder (510); in the same or another embodiment, one encoded video sequence is decoded at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. Each video sequence may be associated with multiple video frames or images. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective processing circuits (not shown). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be configured between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be provided external to the video decoder (510) and separate from the video decoder (510) (not shown). In other applications, a buffer memory (not shown) may be provided external to the video decoder (510), for example, to prevent network jitter, and another buffer memory (515) may be provided internal to the video decoder (510), for example, to handle broadcast timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be required, or the buffer memory may be made smaller. Of course, for use on a traffic packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and the buffer memory may be relatively large. Such a buffer memory may have an adaptive size and may be at least partially implemented in an operating system or similar component (not shown) external to the video decoder (510).

[0060] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the encoded video sequence. The types of symbols include information for managing the operation of the video decoder (510) and potentially information for controlling a display device such as a display device (512) (e.g., a display screen), which may or may not be part of the electronic device (530) but may be coupled to the electronic device (530), such as Figure 5As shown in . The control information for the display device may be a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). The entropy coding of the coded video sequence may be performed according to a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (520) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (eg, transform coefficients), quantizer parameter values, motion vectors, and so on.

[0061] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).

[0062] Depending on the type of coded video picture or portion of a coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (521) may involve multiple different processing or functional units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For the sake of brevity, the flow of such subgroup control information between the parser (520) and the multiple processing or functional units described below is not described.

[0063] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these functional units interact closely with each other and may be integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the conceptual subdivision of functional units is adopted in the following disclosure.

[0064] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantized transform coefficients as symbols (521) from the parser (520) along with control information including information indicating which type of inverse transform to use, block size, quantization factors / parameters, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output a block including sample values, which may be input to an aggregator (555).

[0065] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; for example, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses reconstructed surrounding block information stored in the current picture buffer (558) to generate surrounding blocks of the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers partially reconstructed current pictures and / or fully reconstructed current pictures. In some implementations, the aggregator (555) adds the prediction information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0066] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to inter-frame coded and potentially motion compensated blocks. In this case, the motion compensated prediction unit (553) may access the reference picture memory (557) to extract samples for inter-frame picture prediction. After the extracted samples are motion compensated according to the symbols (521), these samples may be added to the output of the scaler / inverse transform unit (551) by the aggregator (555) (the output of the unit 551 is called residual samples or residual signal), thereby generating output sample information. The retrieval of the prediction samples by the motion compensated prediction unit (553) from the address in the reference picture memory (557) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (553) in the form of the symbols (521), for example, including X and Y components (displacement) and a reference picture component (time). Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, motion compensation may also be associated with a motion vector prediction mechanism, etc.

[0067] The output samples of the aggregator (555) may be employed by various loop filtering techniques in a loop filter unit (554). The video compression techniques may include in-loop filter techniques that are controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and that are available to the loop filter unit (556) as symbols (521) from the parser (520). However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, as well as to previously reconstructed and loop filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in further detail below.

[0068] The output of the loop filter unit (556) may be a sample stream that may be output to a display device (512) and stored in a reference picture memory (557) for subsequent inter-picture prediction.

[0069] Once fully reconstructed, certain coded pictures can be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557) and a new current picture buffer can be reallocated before starting to reconstruct a subsequent coded picture.

[0070] The video decoder (510) may perform decoding operations according to a predetermined video compression technique employed, for example, in the ITU-T H.265 standard. A coded video sequence may conform to the syntax specified by the video compression technique or standard used in the sense that the coded video sequence follows the syntax of the video compression technique or standard and a profile documented in the video compression technique or standard. Specifically, the profile may select certain tools from among all the tools available in the video compression technique or standard as the only tools available for use under the profile. For conformance, the complexity of the coded video sequence is within a range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by the Hypothetical Reference Decoder (HRD) specification and metadata about the HRD buffer management signaled in the coded video sequence.

[0071] In one embodiment, a receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0072] Figure 6 1 is a block diagram of a video encoder (603) according to an embodiment disclosed in the present application. The video encoder (603) is provided in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace Figure 4 A video encoder (403) in an embodiment.

[0073] The video encoder (603) can be used to generate a video from a video source (601) (not Figure 6 In another embodiment, the video source (601) can be implemented as part of the electronic device (620) to receive video samples, and the video source can capture video images to be encoded by the video encoder (603). In another embodiment, the video source (601) can be implemented as part of the electronic device (620).

[0074] The video source (601) may provide a source video sequence in the form of a stream of digital video samples to be encoded by the video encoder (603), wherein the stream of digital video samples may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, XYZ, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that are imparted with motion when viewed sequentially. The pictures themselves may be constructed as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. being used. The relationship between pixels and samples may be readily understood by those skilled in the art. The following description focuses on samples.

[0075] According to an embodiment, the video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, the coupling is not shown in the figure. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate-distortion optimization technology, etc.), picture size, group of pictures (GOP) layout, maximum allowed motion vector search range, etc. The controller (650) may be used to have other suitable functions that relate to the video encoder (503) optimized for a certain system design.

[0076] In some embodiments, the video encoder (603) operates in a coding loop. As a simplified description, in embodiments, the coding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data, even though the embedded decoder 633 processes the encoded video stream through the source encoder 630 without entropy coding (because any compression between the symbols and the encoded video stream is lossless in the video compression techniques considered in this application). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because the decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same sample values that the decoder will "see" when using prediction during decoding. This fundamental principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained, eg due to channel errors) is used to improve coding quality.

[0077] The operation of the "local" decoder (633) can be combined with the operation of Figure 5 The "remote" decoder described in detail for the video decoder (510) is identical. However, additional brief reference is made to Figure 5 , when symbols are available and the entropy encoder (645) and parser (520) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the encoder's local decoder (633).

[0078] At this point, it can be observed that any decoder technology other than parsing / entropy decoding present in the decoder must also be present in a substantially identical functional form in the corresponding encoder. For this reason, this application sometimes focuses on the decoder operation, which is related to the decoding portion of the encoder. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described in detail. The encoder is described in more detail below only in certain areas or aspects.

[0079] During operation, in some embodiments, the source encoder (630) may perform motion-compensated predictive coding. Motion-compensated predictive coding predictively encodes an input picture with reference to one or more previously encoded pictures in a video sequence designated as "reference pictures." In this manner, the encoding engine (632) encodes the differences (residuals) between blocks of pixels in color channels of the input picture and blocks of pixels in a reference picture that may be selected as a prediction reference for the input picture. The term "residual" and its adjective form "residual" may be used interchangeably.

[0080] The local video decoder (633) may decode the coded video data of a picture that may be designated as a reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the coded video data is available at the video decoder ( Figure 6 When decoded at a local (not shown) location, the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that the video decoder may perform on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the remote video decoder.

[0081] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (635) may operate on a pixel-by-pixel-block basis based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (635), it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).

[0082] The controller (650) can manage encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0083] The outputs of all the above functional units may be entropy coded in an entropy encoder (645). The entropy encoder (645) losslessly compresses the symbols generated by the various functional units using techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0084] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission over a communication channel (660), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0085] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types:

[0086] An intra picture (I picture) can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and features.

[0087] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values for each block.

[0088] Bidirectionally predictive pictures (B pictures) can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict sample values for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0089] A source picture is typically spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or they may be predictively coded (spatial or intra-predicted) with reference to already coded blocks of the same picture. Pixel blocks of a P picture may be predictively coded using spatial prediction or temporal prediction with reference to a previously coded reference picture. Blocks of a B picture may be predictively coded using spatial prediction or temporal prediction with reference to one or two previously coded reference pictures. Source pictures or intermediately processed pictures may be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same pattern, as described in further detail below.

[0090] The video encoder (603) may perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. In operation, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video coding technique or standard used.

[0091] In an embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, redundant pictures and slices, and other forms of redundant data, SEI messages, VUI parameter set fragments, and the like.

[0092] The captured video may be presented as a temporal sequence of multiple source pictures (video pictures). Intra-picture prediction (often shortened to intra prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In an embodiment, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a previously encoded and buffered reference picture in the video, the block in the current picture can be encoded using a vector called a motion vector. The motion vector points to the reference block in a reference picture, and when multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0093] In some embodiments, bidirectional prediction techniques can be used for inter-picture prediction. According to bidirectional prediction techniques, two reference pictures are used, for example, a first reference picture and a second reference picture, both of which precede the current picture in the video in decoding order (but may be in the past or future, respectively, in display order). A block in the current picture can be encoded using a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be jointly predicted using a combination of the first reference block and the second reference block.

[0094] In addition, merge mode technology can be used in inter-picture prediction to improve coding efficiency.

[0095] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, a picture in a video picture sequence is divided into coding tree units (CTUs) for compression, and the CTUs in the picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally speaking, a CTU includes three parallel coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Furthermore, each CTU can be split into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or four 32×32 pixel CUs. Each of the one or more 32×32 blocks can be further divided into four 16×16 pixel CUs. In one embodiment, each CU can be analyzed during encoding to determine the prediction type for the CU among various prediction types, such as inter-frame prediction type or intra-frame prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Splitting the CU into PUs (or PBs of different color channels) can be performed in various spatial modes. For example, a luma or chroma PB may include a matrix of values of samples (e.g., luma values), such as 8x 8 pixels, 16x 16 pixels, 8x 16 pixels, and 16x 8 samples.

[0096] Figure 7is a diagram of a video encoder (703) according to another exemplary embodiment disclosed herein. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and to encode the processed block into an encoded picture that is part of an encoded video sequence. In this embodiment, the video encoder (703) is configured to replace Figure 4 A video encoder (303) in an embodiment.

[0097] In one embodiment, the video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) then uses, for example, rate-distortion optimization (RDO) to determine whether to use intra mode, inter mode, or bi-prediction mode to encode the processing block. When it is determined that the processing block is to be encoded in intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when it is determined that the processing block is to be encoded in inter mode or bi-prediction mode, the video encoder (703) may use inter prediction or bi-prediction techniques, respectively, to encode the processing block into an encoded picture. In some exemplary embodiments, merge mode may be a sub-mode of inter-picture prediction, in which a motion vector is derived from one or more motion vector predictors without the aid of an encoded motion vector component external to the predictor. In certain other exemplary embodiments, there may be a motion vector component applicable to the subject block. Thus, the video encoder (703) includes a method for encoding the processing block in the following manner: Figure 7 Components explicitly shown in the , such as a mode decision module (not shown) for determining a prediction mode for a processing block.

[0098] exist Figure 7 In an embodiment of the present invention, the video encoder (703) includes Figure 7 An inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721) and an entropy encoder (725) coupled together are shown in an exemplary arrangement.

[0099] The inter-frame encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in display order), generate inter-frame prediction information (e.g., redundant information description, motion vectors, merge mode information according to an inter-frame coding technique), and calculate an inter-frame prediction result (e.g., a prediction block) based on the inter-frame prediction information using any suitable technique. In some embodiments, the reference picture is encoded using an embedded image. Figure 6 The decoding unit 633 in the example encoder 620 (such as Figure 7 , as shown in the residual decoder 728, which will be described in detail below), which decodes the decoded reference pictures based on the encoded video information.

[0100] The intra-frame encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with previously encoded blocks in the same picture in some cases, generate quantization coefficients after transformation, and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques). The intra-frame encoder (722) calculates an intra-frame prediction result (e.g., a prediction block) based on the intra-frame prediction information and a reference block in the same picture.

[0101] The general controller (721) is used to determine general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines a prediction mode for a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is intra mode, the general controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information and add the intra prediction information to the bitstream; and when the prediction mode for the block is inter mode, the general controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select inter prediction information and add the inter prediction information to the bitstream.

[0102] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result of the block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is configured to convert the residual data from the time domain to the frequency domain to generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra prediction information. The decoded blocks are processed appropriately to generate decoded pictures, and the decoded pictures may be buffered in memory circuitry (not shown) and used as reference pictures.

[0103] The entropy encoder (725) is used to format the codestream to produce encoded blocks and perform entropy encoding. The entropy encoder (725) is configured to include various information in the codestream. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other appropriate information in the codestream. It should be noted that when encoding a block in inter-frame mode or the merge submode of bidirectional prediction mode, there is no residual information.

[0104] Figure 8 FIG is a diagram of a video decoder (810) according to another embodiment disclosed herein. The video decoder (810) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (810) is configured to replace Figure 4 A video decoder (410) of an embodiment.

[0105] exist Figure 8 In one embodiment, the video decoder (810) includes Figure 8 The schematic arrangement shown in FIG. 8 is an entropy decoder ( 871 ), an inter-frame decoder ( 880 ), a residual decoder ( 873 ), a reconstruction module ( 874 ), and an intra-frame decoder ( 872 ) coupled together.

[0106] The entropy decoder (871) can be used to reconstruct certain symbols from the encoded picture, which represent syntax elements that constitute the encoded picture. Such symbols may include, for example, the mode used to encode the block (e.g., intra mode, inter mode, bidirectional prediction mode, merge submode, or another submode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata for use by the intra decoder (872) or inter decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, and the like. In an embodiment, when the prediction mode is inter or bidirectional prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be inverse quantized and provided to the residual decoder (873).

[0107] The inter-frame decoder (880) is configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0108] The intra-frame decoder (872) is configured to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information.

[0109] The residual decoder (873) is used to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also use certain control information (to obtain the quantizer parameter QP), which may be provided by the entropy decoder (871) (the data path is not shown because this is only low-data-volume control information).

[0110] The reconstruction module (874) is used to combine the residual output by the residual decoder (873) with the prediction result (which can be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block. The reconstructed block forms part of the reconstructed picture, and the reconstructed picture can be used as part of the reconstructed video. It should be noted that other suitable operations such as deblocking operations can be performed to improve visual quality.

[0111] It should be noted that the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810) may be implemented using any suitable technology. In one embodiment, the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810) may be implemented using one or more processors executing software instructions.

[0112] Turning to block partitioning for encoding and decoding, partitioning can generally start from basic blocks and can follow a predefined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. Partitioning can be hierarchical and recursive. After dividing or partitioning the basic blocks according to the example partitioning process or any one or combination of the other processes described below, a final set of partitions or codec blocks can be obtained. Each of these partitions can be at one of the various partition levels in the partition hierarchy and can have various shapes. Each of the partitions can be referred to as a codec block (CB). For the various example partitioning implementations described further below, each resulting CB can have any of the allowed sizes and partition levels. Such partitions are referred to as codec blocks because they can form units for which some basic encoding / decoding decisions can be made and encoding / decoding parameters can be optimized, determined, and signaled in the coded video stream. The highest or deepest level in the final partition represents the depth of the codec block partition structure of the tree. The codec block can be a luminance codec block or a chrominance codec block. The CB tree structure for each color can be referred to as a codec block tree (CBT).

[0113] The codec blocks of all color channels may be collectively referred to as a codec unit (CU). The hierarchical structure of all color channels may be collectively referred to as a coding tree unit (CTU). The partition patterns or structures of the various color channels in a CTU may be the same or different.

[0114] In some embodiments, the partition tree schemes or structures for luma and chroma channels may not have to be the same. In other words, the luma and chroma channels may have separate coding tree structures or modes. Further, whether the luma and chroma channels use the same or different codec partition tree structures and the actual codec partition tree structures to be used may depend on whether the slice being encoded is a P, B, or I slice. For example, for an I slice, the chroma channel and the luma channel may have their own codec partition tree structures or codec partition tree structure modes, while for a P or B slice, the luma and chroma channels may share the same codec partition tree scheme. When separate codec partition tree structures or modes are applied, the luma channel may be partitioned into CBs by one codec partition tree structure, and the chroma channels may be partitioned into chroma CBs by another codec partition tree structure.

[0115] In some example embodiments, a basic block may apply a predetermined partitioning scheme. Figure 9As shown, an example 4-way partition tree may be employed. The 4-way partition tree may start from a first predetermined level (e.g., 64×64 block level or other size, such as basic block size), and basic blocks may be partitioned hierarchically down to a predefined lowest level (e.g., 4×4 level). For example, basic blocks may be restricted to four predetermined partitioning options or modes indicated by 902, 904, 906, and 908, where the partition designated as R allows for recursive partitioning. Figure 9 The same partitioning options shown can be repeated at lower scales, down to the lowest level (e.g., a 4x4 level). In some embodiments, additional restrictions may be applied to Figure 9 The partition scheme of Figure 9 In an embodiment of the invention, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed, but they may not be recursive, while square partitions may be allowed to be recursive. If necessary, follow Figure 9 The recursive partitioning of generates the final set of coding blocks. The coding tree depth can be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 64x64 block) can be set to 0, and the coding tree depth of the root block can be set to 0. Figure 9 After being further split once, the coding tree depth increases by 1. For the above scheme, the maximum or deepest level from 64x 64 basic block to 4x 4 minimum partition is 4 (starting from level 0). Such a partitioning scheme can be applied to one or more color channels. Figure 9 The scheme partitions each color channel independently (e.g., a partitioning pattern or option in a predefined pattern may be determined independently for each color channel at each level of the hierarchy). Alternatively, two or more color channels may share Figure 9 The same hierarchical pattern tree (for example, the same partitioning pattern or option in the predefined pattern can be selected for two or more color channels at each hierarchical level).

[0116] Figure 10 Another example predefined partitioning scheme that allows recursive partitioning to form a partition tree is shown. Figure 10 As shown, an example 10-way partition structure or pattern may be predefined. A root block may start at a predefined level (eg, from a basic block at a 128x128 level, or a 64x64 level). Figure 10 Example partition structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Figure 10 The second row 1002, 1004, 1006 and 1008 indicate a partition type with three sub-partitions, which may be referred to as a "T-type" partition. The "T-type" partitions 1002, 1004, 1006 and 1008 may be referred to as a left T-type, a top T-type, a right T-type and a bottom T-type. In some example embodiments, Figure 10The rectangular partition of does not allow further subdivision. The coding tree depth can be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128×128 block) can be set to 0, and the coding tree depth of the root block can be set to 0. Figure 10 After the pattern is further split once, the coding tree depth is increased by 1. In some embodiments, all square partitions in 1010 may only be allowed to be divided according to Figure 10 In other words, for the square partitions in the T-patterns 1002, 1004, 1006, and 1008, recursive partitioning may not be allowed. If necessary, follow Figure 10 The recursive partitioning process generates the final set of coding blocks. This scheme can be applied to one or more color channels. In some embodiments, more flexibility can be added when using partitions below the 8x8 level. For example, in some cases, 2x2 chroma inter-frame prediction can be used.

[0117] In some other example embodiments for codec block partitioning, a quadtree structure can be used to partition a basic block or intermediate block into quadtree partitions. This type of quadtree partitioning can be applied hierarchically and recursively to any square partition. Whether a basic block or intermediate block or partition is further quadtree partitioned can be adapted to various local characteristics of the basic block or intermediate block / partition. The quadtree partitioning at picture boundaries can be further adjusted. For example, implicit quadtree partitioning can be performed at picture boundaries so that the block will remain quadtree partitioned until it is sized to fit within the picture boundary.

[0118] In some other example embodiments, hierarchical binary partitioning can be used starting from a basic block. For such a scheme, a basic block or intermediate block can be partitioned into two partitions. The binary partitioning can be horizontal or vertical. For example, horizontal binary partitioning can split a basic block or intermediate block into equal right and left partitions. Similarly, vertical binary partitioning can split a basic block or intermediate block into equal upper and lower partitions. Such binary partitioning can be hierarchical and recursive. A decision can be made at each of the basic blocks or intermediate blocks whether the binary partitioning scheme should continue, and if the scheme does continue further, a decision is made whether horizontal or vertical binary partitioning should be used. In some embodiments, further partitioning is stopped at a predefined minimum partition size (in one or two dimensions). Alternatively, further partitioning can be stopped once a predefined partition level or depth from the basic block is reached. In some embodiments, the aspect ratio of the partitions can be limited. For example, the aspect ratio of the partitions can be no less than 1:4 (or greater than 4:1). As such, a vertical strip partition having a vertical to horizontal aspect ratio of 4:1 can only be further vertically binary partitioned into an upper partition and a lower partition each having a vertical to horizontal aspect ratio of 2:1.

[0119] In yet other examples, a ternary partitioning scheme may be used to partition a basic block or any intermediate block, such as Figure 13 The ternary graph can be Figure 13 1302 is implemented vertically, or as shown in Figure 13 The 1304 level is implemented. Figure 13 In the example partition ratio of vertical or horizontal is shown as 1:2:1, but other ratios can also be predefined. In some embodiments, two or more different ratios can be predefined. Such ternary partitioning schemes can be used to supplement quadtree or binary partitioning structures because such ternary tree partitioning can capture objects located at the center of a block in one continuous partition, while quadtree and binary trees always partition along the center of the block and thus divide the object into separate partitions. In some embodiments, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transformations.

[0120] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree and binary partitioning schemes can be combined to partition a basic block into a quadtree-binary tree (QTBT) structure. In such a scheme, if specified, subject to a set of predefined conditions, a basic block or intermediate block / partition can be either quadtree partitioned or binary partitioned. Figure 14 A specific example is shown in . Figure 14 In the example of FIG, a basic block is first quadtree partitioned into four partitions, as shown in 1402, 1404, 1406, and 1408. Thereafter, each of the resulting partitions is either quadtree partitioned into four further partitions (such as 1408), or binary partitioned at the next level into two further partitions (horizontally or vertically, such as 1402 or 1406, both of which are symmetrical), or not partitioned (such as 1404). For square partitions, binary or quadtree partitioning can be recursively allowed, as shown in the overall example partitioning diagram of 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree partitioning and dashed lines represent binary partitioning. A flag can be used for each binary partition node (non-leaf binary partition) to indicate whether the binary partition is horizontal or vertical. For example, as shown in 1420, consistent with the partitioning structure of 1410, a flag of "0" can represent a horizontal binary partition, and a flag of "1" can represent a vertical binary partition. For quadtree partitioning, there is no need to indicate the partition type, because quadtree partitioning always splits the block or partition horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some embodiments, the flag "1" can indicate horizontal binary partitioning, and the flag "0" can indicate vertical binary partitioning.

[0121] In some example implementations of QTBT, the quadtree and binary segmentation rule set may be represented by the following predefined parameters and corresponding functions associated therewith: —CTU size: the root node size of the quadtree (the size of the basic block) —MinQTSize: Minimum allowed quadtree leaf node size —MaxBTSize: Maximum allowed binary tree root node size —MaxBTDepth: Maximum allowed binary tree depth —MinBTSize: minimum allowed binary tree leaf node size In some example implementations of the QTBT partitioning structure, the CTU size can be set to 128x128 luma samples with two corresponding 64x64 chroma sample blocks (when example chroma subsampling is considered and used), MinQTSize can be set to 16x16, MaxBTSize can be set to 64x64, MinBTSize (for both width and height) can be set to 4x4, and MaxBTDepth can be set to 4. Quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes from their minimum allowed size of 16x16 (i.e., MinQTSize) to 128x128 (i.e., the CTU size). If the node is 128x128, it will not be split by the binary tree first because the size exceeds the MaxBTSize (i.e., 64x64). Otherwise, the nodes that do not exceed the MaxBTSize can be partitioned by the binary tree. In Figure 14 In the example of , the basic block is 128×128. According to a predefined set of rules, the basic block can only be split by a quadtree. The partition depth of the basic block is 0. Each of the four resulting partitions is 64×64, does not exceed MaxBTSize, and can be further split by a quadtree or a binary tree at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further splitting can be considered. When the width of the binary tree node is equal to MinBTSize (i.e., 4), no further horizontal splitting can be considered. Similarly, when the height of the binary tree node is equal to MinBTSize, no further vertical splitting can be considered.

[0122] In some example embodiments, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure or separate QTBT structures for luma and chroma. For example, for P and B slices, the luma and chroma CTBs in one CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CBs by the QTBT structure, and the chroma CTB can be partitioned into chroma CBs by another QTBT structure. This means that a CU can be used to represent different color channels in an I slice, for example, an I slice can consist of a codec block for the luma component or a codec block for two chroma components, and a CU in a P or B slice can consist of codec blocks for all three color components.

[0123] In some other embodiments, the QTBT scheme can be supplemented with the above-mentioned ternary scheme. Such an embodiment can be called a multi-type tree (MTT) structure. For example, in addition to the binary split of the node, you can also choose Figure 13 In some embodiments, only square nodes can be triadically partitioned. An additional flag may be used to indicate whether the triad partition is horizontal or vertical.

[0124] Two-level or multi-level tree designs, such as the QTBT implementation and the QTBT implementation supplemented by ternary partitioning, can be designed primarily to reduce complexity. In theory, the complexity of traversing the tree is T D , where T represents the number of split types and D is the depth of the tree. A trade-off can be made by using multiple types (T) while reducing the depth (D).

[0125] In some embodiments, CB can be further partitioned. For example, for the purpose of intra-frame or inter-frame prediction during encoding and decoding, CB can be further partitioned into multiple prediction blocks (PBs). In other words, CB can be further divided into different sub-partitions, in which separate prediction decisions / configurations can be made. In parallel, for the purpose of depicting the level of transformation or inverse transformation of video data, CB can be further partitioned into multiple transform blocks (TBs). The partitioning schemes of CB to PB and TB can be the same or different. For example, each partitioning scheme can be performed using its own process based on various characteristics of video data, for example. In some example embodiments, PB and TB partitioning schemes can be independent. In some other example embodiments, PB and TB partitioning schemes and boundaries can be related. In some embodiments, for example, TB can be partitioned after PB partitioning, in particular, each PB is further partitioned into one or more TBs after being determined according to the partitioning of the coding block. For example, in some embodiments, PB can be divided into one, two, four or other number of TBs.

[0126] In some embodiments, in order to partition the basic blocks into coding blocks and further into prediction blocks and / or transform blocks, the luma channel and the chroma channels can be treated differently. For example, in some embodiments, the coding blocks of the luma channel can be allowed to be partitioned into prediction blocks and / or transform blocks, while the coding blocks of the chroma channels can not be partitioned into prediction blocks and / or transform blocks. In such embodiments, the transformation and / or prediction of the luma block can therefore be performed only at the coding block level. For another example, the minimum transform block size of the luma channel and the chroma channels can be different, for example, the coding blocks of the luma channel can be allowed to be partitioned into transform blocks and / or prediction blocks that are smaller than those of the chroma channels. For yet another example, the maximum depth of partitioning the coding blocks into transform blocks and / or prediction blocks can be different between the luma channel and the chroma channels, for example, the coding blocks of the luma channel can be allowed to be partitioned into transform blocks and / or prediction blocks that are deeper than those of the chroma channels. For a specific example, a luma coding block may be partitioned into transform blocks of multiple sizes, which may be represented by recursive partitioning down to up to 2 levels, and may allow transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transform block sizes from 4×4 to 64×64. However, for chroma blocks, only the largest possible transform block may be specified for the luma block.

[0127] In some example implementations for partitioning a coding block into PBs, the depth, shape, and / or other characteristics of partitioning the PB may depend on whether the PB is intra-coded or inter-coded.

[0128] In various example schemes, partitioning of a coding block (or prediction block) into transform blocks may be implemented, including but not limited to recursive or non-recursive quadtree partitioning and predefined pattern partitioning, with additional consideration of transform blocks at the boundaries of the coding block or prediction block. In general, the resulting transform blocks may be at different partitioning levels, may not have the same size, and may not need to be square in shape (e.g., they may be rectangular with some allowed size and aspect ratio). Figure 15 、 16 and 17 describe further examples in more detail.

[0129] However, in some other embodiments, the CB obtained via any of the above partitioning schemes can be used as a basic or minimum encoding block for prediction and / or transformation. In other words, no further partitioning is performed for the purpose of performing inter-frame prediction / intra-frame prediction and / or for the purpose of transformation. For example, the CB obtained from the above QTBT scheme can be used directly as a unit for performing prediction. Specifically, such a QTBT structure removes the concept of multiple partition types, that is, it removes the separation of CU, PU and TU, and supports more flexibility in the CU / CB partition shape as described above. In such a QTBT block structure, the CU / CB can have a square or rectangular shape. The leaf nodes of such a QTBT are used as units of prediction and transformation processing without any further partitioning. This means that the CU, PU and TU have the same block size in such an example QTBT encoding and decoding block structure.

[0130] The above various CB partitioning schemes and further partitioning of CB into PB and / or TB (including no PB / TB partitioning) can be combined in any manner.The following specific embodiments are provided as non-limiting examples.

[0131] Specific example implementations of coding block and transform block partitioning are described below. In such example implementations, the recursive quadtree partitioning described above or predefined partitioning patterns (e.g. Figure 9 and 10 The basic block is partitioned into coding blocks (those modes in ). At each level, whether further quadtree partitioning of a particular partition should continue can be determined by local video data characteristics. The resulting CBs can be at various quadtree partitioning levels of various sizes. A decision can be made at the CB level (or CU, for all three color channel channels) whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode a picture area. Each CB can be further partitioned into one, two, four or other number of PBs according to a predefined PB partitioning type. Within a PB, the same prediction process can be applied, and relevant information is sent to the decoder based on the PB. After obtaining the residual block by applying the prediction process based on the PB partitioning type, the CB can be partitioned into TBs according to another quadtree structure similar to the coding tree of the CB. In this particular embodiment, the CB or TB can be, but is not necessarily limited to, a square. In addition, in this particular example, the PB can be a square or rectangular shape for inter-frame prediction and can be only a square for intra-frame prediction. The coding block can be partitioned into, for example, four square-shaped TBs. Each TB can be further recursively split (using quadtree partitioning) into smaller TBs, called residual quadtrees (RQTs).

[0132] Another example embodiment for partitioning a basic block into CBs, PBs, and / or TBs is described below. For example, a quadtree of nested multi-type trees may be used that has a segmentation structure using binary and ternary partitioning (e.g., a QTBT as described above or a QTBT with ternary partitioning) rather than using a quadtree such as Figure 9 Or multiple partition unit types shown in 10. The separation of CB, PB and TB concepts (i.e., partitioning CB into PB and / or TB, and partitioning PB into TB) can be abandoned except when a CB with a size that is too large for the maximum transform length is required, where such CB may need to be further split. This example partitioning scheme can be designed to support more flexibility in the shape of the CB partitions, so that both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, the CU can have a square or rectangular shape. Specifically, the coding tree block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by nesting multiple types of tree structures. Figure 11 An example of a nested multi-type tree structure using binary or ternary partitioning is shown in FIG. Figure 11 The example multi-type tree structure includes four split types, which are called vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical ternary split (SPLIT_TT_VER) (1106) and horizontal ternary split (SPLIT_TT_HOR) (1108). The CB then corresponds to the leaves of the multi-type tree. In this example embodiment, unless the CB is too large for the maximum transform length, the segment is used for prediction and transform processing without any further partitioning. This means that in most cases, CB, PB and TB have the same block size in the quadtree with nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CB. In some embodiments, in addition to binary or ternary splits, such as Figure 11 The nested patterns that can be performed include quadtree partitioning.

[0133] Figure 12 An example of a quadtree with nested multi-type tree coding block structure for block partitioning (including quadtree, binary and ternary partitioning options) for one basic block is shown in FIG. In more detail, Figure 12 The basic block 1200 is shown to be partitioned by a quadtree into four square partitions 1202, 1204, 1206 and 1208. Figure 11 The multi-type tree structure and quadtree further split each of the quadtree partitions. Figure 12In the example, partition 1204 is not further split. Partitions 1202 and 1208 each use another quadtree split. For partition 1202, the upper left, upper right, lower left and lower right partitions of the second level quadtree split use the third level of quadtree split, Figure 11 Horizontal binary segmentation 1104, no segmentation, and Figure 11 The horizontal ternary partition 1108 is divided into two parts. Partition 1208 adopts another quadtree partition, and the upper left, upper right, lower left and lower right partitions of the second level quadtree partition adopt Figure 11 The third level of vertical ternary segmentation 1106 is segmentation, no segmentation, no segmentation, and Figure 11 The horizontal binary partition 1104. The two sub-partitions of the third level upper left partition 1208 are based on Figure 11 The horizontal binary segmentation 1104 and the horizontal ternary segmentation 1108 are further divided respectively. Figure 11 The second level segmentation pattern of the vertical binary segmentation 1102 is divided into two partitions, which are divided into two partitions according to Figure 11 The horizontal ternary segmentation 1108 and vertical binary segmentation 1102 are further segmented by a third level. Figure 11 1104, a fourth level segmentation is further applied to one of them.

[0134] For the specific example above, the maximum luma transform size may be 64x64, and the maximum supported chroma transform size may be different for luma at, for example, 32x32. Figure 12 The example CB in is typically not further divided into smaller PBs and / or TBs. When the width or height of a luminance coding block or a chrominance coding block is larger than the maximum transform width or height, the luminance coding block or the chrominance coding block may be automatically split in the horizontal and / or vertical directions to satisfy the transform size constraints along that direction.

[0135] In a specific example for partitioning basic blocks into the above CBs, as described above, the coding tree scheme can support the ability for luma and chroma to have separate block tree structures. For example, for P and B slices, the luma and chroma CTBs in one CTU can share the same coding tree structure. For example, for I slices, luma and chroma can have separate coding block tree structures. When a separate block tree structure is applied, the luma CTB is partitioned into CBs through one coding tree structure, and the chroma CTB is partitioned into chroma CBs through another coding tree structure. This means that a CU in an I slice can consist of coding blocks for the luma component or coding blocks for two chroma components, and a CU in a P or B slice always consists of coding blocks for all three color components unless the video is monochrome.

[0136] When a coding block is further partitioned into multiple transform blocks, the transform blocks are ordered in the bitstream, following various orders or scanning methods. Example implementations for partitioning a coding block or prediction block into transform blocks and the encoding order of the transform blocks are described in further detail below. In some example implementations, as described above, transform partitioning can support a variety of shapes, such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1 transform blocks, with transform block sizes ranging from, for example, 4×4 to 64×64. In some implementations, if the coding block is less than or equal to 64×64, transform block partitioning can be applied only to the luma component, so that for chroma blocks, the transform block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, then the luma and chroma coding blocks can be implicitly split into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.

[0137] In some example implementations of transform block partitioning, for both intra-coded blocks and inter-coded blocks, the coding block can be further partitioned into multiple transform blocks with a partition depth of up to a predetermined number of levels (e.g., 2 levels). The transform block partition depth and size can be related. For some example implementations, an example mapping from the transform size of the current depth to the transform size of the next depth is shown in Table 1 below. Table 1: Change partition size settings

[0138] Based on the example mapping in Table 1, for a 1:1 square block, the next level of transform partitioning can create four 1:1 square sub-transform blocks. Transform partitioning can, for example, stop at 4×4. In this way, the transform size of 4×4 at the current depth corresponds to the same size 4×4 at the next depth. In the example in Table 1, for a 1:2 / 2:1 non-square block, the next level of transform partitioning will create two 1:1 square sub-transform blocks, while for a 1:4 / 4:1 non-square block, the next level of transform partitioning will create two 1:2 / 2:1 sub-transform blocks.

[0139] In some example embodiments, for the luma component of an intra-coded block, additional restrictions may be applied to the transform block partitioning. For example, for each level of transform partitioning, all sub-transform blocks may be restricted to have the same size. For example, for a 32×16 coding block, a level 1 transform partitioning creates two 16×16 sub-transform blocks and a level 2 transform partitioning creates eight 8×8 sub-transform blocks. In other words, the second level partitioning must be applied to all first level sub-blocks to keep the transform unit sizes equal. Figure 151 shows an example of transform block partitioning for intra-coding square blocks as shown in Table 1, and the coding order is illustrated by arrows. Specifically, 1502 shows a square coding block. 1504 shows a first-level partitioning into four equal-sized transform blocks according to Table 1, where the coding order is indicated by arrows. 1506 shows a second-level partitioning of all first-level equal-sized blocks into 16 equal-sized transform blocks according to Table 1, where the coding order is indicated by arrows.

[0140] In some example embodiments, the above restrictions on intra-frame coding may not apply to the luma component of an inter-frame coded block. For example, after the first level of transform partitioning, any of the sub-transform blocks may be further independently partitioned one level further. Thus, the resulting transform blocks may or may not have the same size. Figure 16 An example of dividing an inter-frame coding block into transform blocks is shown in FIG, and the transform blocks have a coding order. Figure 16 In the example of , according to Table 1, the inter-coded block 1602 is split into transform blocks at two levels. At the first level, the inter-coded block is split into four transform blocks of equal size. Then, as shown in 1604, only one of the four transform blocks (not all) is further split into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes. The example coding order of these 7 transform blocks is represented by Figure 16 The arrow in 1604 is shown.

[0141] In some example embodiments, for one or more chroma components, some additional restrictions on transform blocks may be applied. For example, for one or more chroma components, the transform block size may be as large as the coding block size, but not smaller than a predefined size, such as 8×8.

[0142] In some other example embodiments, for coding blocks with a width (W) or height (H) greater than 64, both luma and chroma coding blocks may be implicitly split into multiples of min(W, 64)×min(H, 64) and min(W, 32)×min(H, 32) transform units, respectively. Here, in the present disclosure, "min(a, b)" may return the smaller value between a and b.

[0143] Figure 17 Another optional example scheme for partitioning a coding block or prediction block into transform blocks is further shown. Figure 17 As shown, instead of using recursive transform partitioning, a predefined set of partition types can be applied to the coding block according to its transform type. Figure 17 In the specific example shown in , one of six example partition types can be applied to split a coding block into various numbers of transform blocks. Such a scheme for generating transform block partitions can be applied to coding blocks or prediction blocks.

[0144] In more detail, Figure 17 The partitioning scheme of

[0055] provides up to 6 example partition types for any given transform type (transform type refers to, for example, the type of preliminary transformation, such as ADST and others). In this scheme, a transform partition type can be assigned to each coding block or prediction block based on, for example, rate-distortion cost. In an example, the transform partition type assigned to a coding block or prediction block can be determined based on the transform type of the coding block or prediction block. A specific transform partition type can correspond to a transform block partition size and mode, such as Figure 17 As shown in the 6 transform partition types illustrated in FIG. The correspondence between various transform types and various transform partition types can be predefined. The corresponding examples are shown below, where the capitalized mark indicates the transform partition type that can be assigned to the coding block or prediction block based on the rate-distortion cost:

[0145] PARTITION_NONE: Allocate a transform size equal to the block size.

[0146] PARTITION_SPLIT: The allocated transform size is 1 / 2 the width of the block size and 1 / 2 the height of the block size.

[0147] PARTITION_HORZ: Allocate a transform size with the same width as the block size and 1 / 2 the height of the block size.

[0148] PARTITION_VERT: Allocate a transform size with a width of 1 / 2 the block size and the same height as the block size.

[0149] PARTITION_HORZ4: Allocates a transform size with the same width as the block size and 1 / 4 of the height of the block size.

[0150] PARTITION_VERT4: Allocates a transform size with a width 1 / 4 of the block size and the same height as the block size.

[0151] In the above example, if Figure 17 All transform partition types shown include a uniform transform size for the partition transform blocks. This is merely an example and not a limitation. In some other embodiments, mixed transform block sizes may be used for partition transform blocks of a particular partition type (or mode).

[0152] The PB (or CB, also referred to as a PB when not further partitioned into prediction blocks) obtained from any of the above partitioning schemes can then become the individual blocks used for encoding and decoding via intra-frame or inter-frame prediction. For inter-frame prediction of the current PB, the residual between the current block and the predicted block can be generated, encoded, and included in the encoded bitstream.

[0153] Inter prediction can be implemented, for example, in a single reference mode or a composite reference mode. In some embodiments, a skip flag may first be included in the codestream for the current block (or at a higher level) to indicate whether the current block is inter-coded and will not be skipped. If the current block is inter-coded, another flag may be further included in the codestream to signal whether a single reference mode or a composite reference mode is used for the current block. For a single reference mode, one reference block may be used to generate a prediction block for the current block. For a composite reference mode, two or more reference blocks may be used to generate a prediction block, for example, by weighted averaging. A composite reference mode may be referred to as more than one reference mode, two reference modes, or multiple reference modes. One or more reference blocks may be identified using one or more reference frame indices and additionally using corresponding one or more motion vectors that indicate one or more shifts in position (e.g., in horizontal and vertical pixels) between the one or more reference blocks and the current block. For example, an inter-frame prediction block for a current block can be generated from a single reference block identified by a motion vector in a reference frame as a prediction block in a single reference mode, while for a composite reference mode, a prediction block can be generated by taking a weighted average of two reference blocks in two reference frames indicated by two motion vectors. One or more motion vectors can be encoded and included in a codestream in various ways.

[0154] In some embodiments, the encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be kept in the DPB waiting to be displayed (in the decoding system), and some images / pictures in the DPB may be used as reference frames to implement inter-frame prediction. In some embodiments, the reference frames in the DPB may be marked as short-term references or long-term references for the current image being encoded or decoded. For example, a short-term reference frame may include a frame used for inter-frame prediction of a block in the current frame, or a frame for inter-frame prediction of a block in a predefined number (e.g., 2) of subsequent video frames that are closest to the current frame in decoding order. Long-term reference frames may include frames in the DPB that can be used to predict image blocks in frames that are more than a predetermined number of frames from the current frame in decoding order. Information about such tags for short-term and long-term reference frames may be referred to as a reference picture set (RPS) and may be added to the header of each frame in the encoded bitstream. Each frame in the coded video stream can be identified by a picture order counter (POC), and each frame is numbered in an absolute manner according to the playback sequence, or numbered relative to a group of pictures starting from, for example, an I frame.

[0155] In some example embodiments, one or more reference picture lists may be formed based on information in the RPS, the lists containing identifiers of short-term and long-term reference frames for inter-frame prediction. For example, a single picture reference list, denoted as L0 reference (or reference list 0), may be formed for unidirectional inter-frame prediction, while two picture reference lists, denoted as L0 (or reference list 0) and L1 (or reference list 1), may be formed for bidirectional inter-frame prediction, for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. When multiple references used to generate a prediction block by weighted averaging in composite prediction mode are on the same side of the block to be predicted, unidirectional inter-frame prediction may be in single reference mode or composite reference mode. Bidirectional inter-frame prediction can only be in composite mode because bidirectional inter-frame prediction involves at least two reference blocks.

[0156] In some embodiments, a merge mode (MM) for inter-frame prediction may be implemented. Generally, with merge mode, one or more motion vectors in a single reference prediction or a composite reference prediction for the current PB can be derived from one or more other motion vectors, rather than being independently calculated and signaled. For example, in the encoding system, the one or more current motion vectors for the current PB can be simplified to one or more differences between the one or more current motion vectors and one or more other encoded motion vectors (referred to as reference motion vectors). The one or more differences of the one or more motion vectors (rather than the entire one or more current motion vectors) can be encoded and included in the bitstream and can be linked to one or more reference motion vectors. Correspondingly, in the decoding system, one or more motion vectors corresponding to the current PB can be derived based on the decoded motion vector differences and the one or more decoded reference motion vectors linked to them. As a specific form of general merge mode (MM) inter-frame prediction, this type of inter-frame prediction based on motion vector differences can be referred to as merge mode with motion vector difference (MMVD). Therefore, general MM, or specific MMVD, can be implemented to exploit the correlation between motion vectors associated with different PBs to improve encoding and decoding efficiency. For example, adjacent PBs may have similar motion vectors.For another example, for similarly positioned / placed blocks in space, the motion vectors may be correlated in time (between frames).

[0157] In some example embodiments, an MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in merge mode. Additionally, or alternatively, an MMVD flag may be included in the bitstream during the encoding process and signaled in the bitstream to indicate whether the current PB is in MMVD mode. The MM and / or MMVD flags or indicators may be provided at the PB level, CB level, CU level, CTB level, CTU level, slice level, picture level, etc. For a specific example, for the current CU, both the MM flag and the MMVD flag may be included, and the MMVD flag may be signaled immediately after the skip flag and the MM flag to specify whether the MMVD mode is used for the current CU.

[0158] In some example embodiments of MMVD, a merge candidate list for motion vector prediction can be formed for the block being predicted. The merge candidate list can contain a predetermined number (e.g., 2) of MV predictor candidate blocks whose motion vectors can be used to predict the current motion vector. MVD candidate blocks can include blocks selected from neighboring blocks and / or temporal blocks in the same frame (e.g., blocks co-located in a previous or subsequent frame relative to the current frame). These options represent blocks that may have similar or identical motion vectors as the current block at a spatial or temporal location relative to the current block. The size of the MV predictor candidate list can be predetermined. For example, the list can contain two candidates. To be on the merge candidate list, a candidate block needs to, for example, have the same reference frame (or frames) as the current block, must exist (e.g., when the current block is near the edge of a frame, a boundary check needs to be performed), and must have been encoded during the encoding process and / or decoded during the decoding process. In some embodiments, if spatial neighboring blocks are available and the above conditions are met, the merge candidate list can be first filled with spatial neighboring blocks (scanned in a specific predefined order), and then filled with temporal blocks if space is still available in the list. For example, neighboring candidate blocks can be selected from the left block and the top block of the current block. The merge MV predictor candidate list can be signaled in the bitstream.

[0159] In some embodiments, the actual merge candidate being used as the reference motion vector for predicting the motion vector of the current block can be signaled. In the case where the merge candidate list contains two candidates, a one-bit flag called the merge candidate flag can be used to indicate the selection of the reference merge candidate. For the current block being predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor can be associated with a reference motion vector from the merge candidate list.

[0160] In some example embodiments of MMVD, after a merge candidate is selected and used as a base motion vector predictor for a motion vector to be predicted, a motion vector difference (MVD or delta MV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) may be calculated in the encoding system. Such MVD may include information representing the magnitude of the MV difference and the direction of the MV difference, which may be signaled in the bitstream. The magnitude of the motion difference and the direction of the motion difference may be signaled in various ways.

[0161] In some example implementations of the MMVD, a distance index can be used to specify the magnitude information of the motion vector difference and indicate one of a set of predefined offsets representing the predefined motion vector difference from the starting point (reference motion vector). The MV offset according to the signaled index can then be added to the horizontal or vertical component of the starting (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset is determined by the example direction information of the MVD. An example predefined relationship between the distance index and the predefined offset is specified in Table 2. Table 2 - Example relationship between distance index and predefined MV offset

[0162] In some example embodiments of the MMVD, a direction index may be further signaled and used to indicate the direction of the MVD relative to a reference motion vector. In some embodiments, the direction may be limited to one of a horizontal direction and a vertical direction. An exemplary 2-bit direction index is shown in Table 3. In the example of Table 3, the interpretation of the MVD may vary depending on the information of the start / reference MV. For example, when the start / reference MV corresponds to a unidirectional prediction block or corresponds to a bidirectional prediction block (wherein two reference frame lists point to the same side of the current picture) (i.e., the POCs of both reference pictures are greater than the POC of the current picture, or are less than the POC of the current picture), the symbols in Table 3 may specify the symbol (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a bidirectionally predicted block (in which two reference pictures are on different sides of the current picture) (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the symbols in Table 3 may specify the sign of the MV offset added to the reference MV, the reference MV corresponds to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 may have an opposite value (opposite sign of the offset). Conversely, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the symbols in Table 3 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset of the reference MV associated with picture reference list 0 may have an opposite value. Table 3 - Example implementation of the sign of the MV offset specified by the direction index Direction IDX 00 01 10 11 x-axis (horizontal) + - not applicable not applicable y-axis (vertical) not applicable not applicable + -

[0163] In some example embodiments, the MVD can be scaled based on the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling is required. Otherwise, if the difference in POC in reference list 0 is greater than the difference in POC in reference list 1, the MVD of reference list 1 is scaled. If the POC difference in reference list 1 is greater than that in list 0, the MVD of list 0 can be scaled in the same manner. If the starting MV is unidirectionally predicted, the MVD is added to the available or reference MV.

[0164] In some example implementations of MVD encoding and signaling for bidirectional composite prediction, in addition to or as an alternative to separately encoding and signaling both MVDs, symmetric MVD encoding can be implemented such that only one MVD needs to be signaled, and the other MVD can be derived from the signaled MVD. In such implementations, motion information is signaled, including reference picture indices for both List-0 and List-1. However, only the MVD associated with, for example, Reference List-0 is signaled, while the MVD associated with Reference List-1 is not signaled but derived. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" can be included in the codestream to indicate whether Reference List-1 is not signaled in the codestream. If this flag is 1, indicating that Reference List-1 is equal to zero (and therefore not signaled), the bidirectional prediction flag (called "BiDirPredFlag") can be set to 0, indicating that bidirectional prediction is not present. Otherwise, if mvd_l1_zero_flag is zero, if the closest reference picture in list-0 and the closest reference picture in list-1 form a pair of preceding and following reference pictures or a pair of preceding and following reference pictures, BiDirPredFlag can be set to 1, and both list-0 and list-1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. BiDirPredFlag is 1 to indicate that the symmetric mode flag is additionally signaled in the code stream. When BiDirPredFlag is 1, the decoder can extract the symmetric mode flag from the code stream. For example, the symmetric mode flag can be signaled at the CU level (if necessary), and it indicates whether the symmetric MVD encoding and decoding mode is being used for the corresponding CU. When the symmetric mode flag is 1, it indicates that the symmetric MVD codec mode is used, and the reference picture indices of both list-0 and list-1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") are signaled with the MVD associated with list-0 (referred to as "MVD0"), and another motion vector difference "MVD1" will be derived instead of being signaled. For example, MVD1 can be derived as -MVD0. As such, only one MVD is signaled in the example symmetric MVD mode. In some other example implementations of MV prediction, for both single reference and composite reference mode MV prediction, a coordination scheme can be used to implement general merge mode, MMVD, and some other types of MV prediction. Various syntax elements can be used to signal the manner in which the MV of the current block is predicted.

[0165] For example, for single reference mode, the following MV prediction modes may be signaled:

[0166] NEARMV—directly uses one of the motion vector predictors (MVPs) in the list without using any MVD. This MVP is indicated by the DRL (Dynamic Reference List) index.

[0167] NEWMV—Uses one of the motion vector predictors (MVPs) in the list as a reference and applies the delta to the MVP (e.g., using MVD), which is signaled by the DRL index.

[0168] GLOBALMV—Use motion vectors based on frame-level global motion parameters.

[0169] Likewise, for the composite reference inter prediction mode using two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes may be signaled:

[0170] NEAR_NEARMV—For each of the two MVs to be predicted, one of the motion vector predictors (MVPs) in the list is used instead of MVD, and this MVP is signaled by the DRL index.

[0171] NEAR_NEWMV—To predict the first of the two motion vectors, one of the motion vector predictors (MVPs) in the list is used as the reference MV without MVD, and this MVP is signaled by the DRL index; to predict the second of the two motion vectors, one of the motion vector predictors (MVPs) in the list is used as the reference MV, in combination with an additionally signaled delta MV (MVD), and this MVP is signaled by the DRL index.

[0172] NEW_NEARMV—To predict the second of the two motion vectors, one of the motion vector predictors (MVPs) in the list is used as the reference MV without MVD, and this MVP is signaled by the DRL index; to predict the first of the two motion vectors, one of the motion vector predictors (MVPs) in the list is used as the reference MV in combination with an additionally signaled delta MV (MVD), and this MVP is signaled by the DRL index.

[0173] NEW_NEWMV—Uses one of the motion vector predictors (MVPs) in the list as a reference MV and uses it in conjunction with an additionally signaled delta MV to predict each of the two MVs, where this MVP is signaled via a DRL index.

[0174] GLOBAL_GLOBALMV—Use MV from each reference based on frame-level global motion parameters.

[0175] Thus, the term "NEAR" above refers to MV prediction using a reference MV (without MVD) as a general merge mode, while the term "NEW" refers to MV prediction involving the use of a reference MV and offsetting it with a signaled MVD as in MMVD mode. For composite inter prediction, both the reference base motion vector and the motion vector delta above can typically be different or independent between the two references (even though they can be correlated, and such correlation can be exploited to reduce the amount of information required to signal the two motion vector deltas). In such cases, joint signaling of the two MVDs can be implemented and indicated in the codestream.

[0176] The above dynamic reference list (DRL) can be used to store a set of indexed motion vectors that are dynamically maintained and considered as candidate motion vector predictors.

[0177] In some example embodiments, for composite prediction, an optical flow-based approach can be used to refine motion vectors (MVs) on a per-subblock basis. Specifically, the optical flow equation can be applied to formulate a least-squares problem, whereby fine motion can be derived from the gradients of composite inter-prediction samples. Utilizing these fine motions, the MV of each sub-block can be refined within the prediction block, which can enhance inter-prediction quality. Some codec features can be an extension of the concept of bidirectional optical flow (BDOF), as it supports MV refinement when two reference blocks have arbitrary temporal distances to the current block.

[0178] In some embodiments, four additional composite inter modes listed below may be added: NEAR_NEARMV_OPTFLOW, NEAR_NEWMV_OPTFLOW, NEW_NEARMV_OPTFLOW, and / or NEW_NEWMV_OPTFLOW.

[0179] These modes may be referred to as optical flow modes, and the reference MV type may be defined as a regular composite mode (eg, NEAR_NEWMV_OPTFLOW has the same reference MV type as in NEAR_NEWMV). Composite prediction may be performed based on per-subblock modified MVs rather than the original MVs.

[0180] The various embodiments and / or implementations described in this disclosure may be used alone or in any combination. Further, a portion, all, or any portion or all of these embodiments and / or implementations may be embodied as a part of an encoder and / or decoder and may be implemented with hardware and / or software. For example, they may be hard-coded in a dedicated processing circuit (e.g., one or more integrated circuits). In another example, they may be implemented by executing a program stored in a non-volatile computer-readable medium by one or more processors.

[0181] There may be some concerns / problems associated with some implementations of signaling methods for motion vector differences, for example, how to signal one or more delta MVs in NEW_NEARMV mode, NEAR_NEWMV mode, or NEW_NEWMV mode. One concern / problem is that the correlation of motion vector differences in the two reference lists is not exploited, thus reducing encoding / decoding efficiency and performance.

[0182] The present disclosure describes various embodiments for signaling motion vector differences (MVD or delta MV) for inter-prediction mode encoding and / or decoding, addressing at least one of the aforementioned concerns / problems and enabling efficient software / hardware implementations for improved inter-prediction mode encoding / decoding.

[0183] In various embodiments, reference Figure 18 , a method 1800 for video decoding is provided, which is executed in a device, the device including a memory storing instructions and a processor in communication with the memory. The method 1800 may include some or all of the following steps: step 1810, receiving an encoded video stream; step 1820, extracting an inter-frame prediction mode and a joint incremental motion vector MV of a current block in a current frame from the encoded video stream; step 1830, extracting a first flag indicating whether a first incremental MV of a first reference frame and a second incremental MV of a second reference frame are jointly signaled from the encoded video stream; step 1840, in response to the first flag indicating that the first incremental MV and the second incremental MV are jointly signaled, deriving a first incremental MV and a second incremental MV based on the joint incremental MV; and / or step 1850, decoding the current block in the current frame based on the first incremental MV and the second incremental MV.

[0184] In some embodiments, a joint delta MV may be denoted as joint_delta_mv, which may be an element indicating a joint delta MV. In some embodiments, a first flag indicating whether a first delta MV of a first reference frame and a second delta MV of a second reference frame are jointly signaled may be denoted as joint_mvd_flag. In some embodiments, the first reference frame may be a frame in a reference list (reference list 0); and / or the second reference frame may be a frame in another reference list (reference list 1).

[0185] In some embodiments, step 1830 may include extracting a first flag (joint_mvd_flag) from the encoded video stream, the first flag (joint_mvd_flag) indicating whether the first incremental MV of the first reference frame in reference list 0 and the second incremental MV of the second reference frame in reference list 1 are jointly signaled.

[0186] In various embodiments of the present disclosure, the size of a block (such as but not limited to a coding block, a prediction block, or a transform block) may refer to the width or height of the block. The width or height of the block may be an integer in pixels. In various embodiments of the present disclosure, the size of a block may refer to the area size of the block. The area size of the block may be an integer calculated in pixels by multiplying the width of the block by the height of the block. In some different embodiments of the present disclosure, the size of a block may refer to the maximum value of the width or height of the block, the minimum value of the width and height of the block, or the aspect ratio of the block. The aspect ratio of the block may be calculated as the width of the block divided by the height of the block, or may be calculated as the height divided by the width of the block.

[0187] Here, in some embodiments of the present disclosure, the "first" reference frame may refer not only to "one" reference frame, but also to the "first" reference frame among multiple reference frames (e.g., having the smallest index, or appearing earliest in the sequence), and the "second" reference frame may refer not only to "another" reference frame, but also to the "second" reference frame among multiple reference frames (e.g., having the second smallest index, or appearing second earliest in the sequence).

[0188] Here, in various embodiments of the present disclosure, “signaling XYZ” may refer to encoding XYZ into an encoded codestream during an encoding process; and / or, after transmitting the encoded codestream from one device to another device, “signaling XYZ” may refer to decoding / extracting XYZ from the encoded codestream during a decoding process.

[0189] Here, in various embodiments of the present disclosure, the direction of the reference frame can be determined by whether the reference frame is before or after the current frame in the display order. In some embodiments of the composite reference mode, when the picture order count (POC) of both reference frames of a motion vector pair is greater than or less than the POC of the current frame, the directions of the two reference frames are the same. Otherwise, when the POC of one reference frame is greater than the POC of the current frame and the POC of the other reference frame is less than the POC of the current frame, the directions of the two reference frames are different.

[0190] Here, in various embodiments of the present disclosure, a 'block' may refer to a prediction block, a coding block, a transform block, or a coding unit (CU).

[0191] Referring to step 1810, the device may be Figure 5 electronic device (530) or Figure 8 In some embodiments, the device may be a video decoder (810) in Figure 6 In other embodiments, the device may be Figure 5a portion of an electronic device (530) in Figure 8 a portion of a video decoder (810) in, or Figure 6 The coded video stream can be a part of the decoder (633) in the encoder (620). Figure 8 coded video sequence in , or Figure 6 or Figure 7 The intermediate encoded data in .

[0192] Referring to step 1820, the inter prediction mode and joint incremental motion vector (MV) for the current block in the current frame are extracted from the encoded video stream. The current block may be in composite reference mode. The inter prediction mode may include one of the following: NEAR_NEAR mode, NEW_NEARMV mode, NEAR_NEWMV mode, or NEW_NEWMV mode. The joint incremental MV may be referred to as an MV difference (MVD).

[0193] Referring to step 1830 , a first flag indicating whether a first incremental MV of a first reference frame and a second incremental MV of a second reference frame are jointly signaled is extracted from the encoded video stream.

[0194] In some implementations of one or more inter-frame prediction modes, the first flag is encoded in the encoded video stream and can be extracted from the encoded video stream.

[0195] In some embodiments of one or more inter-frame prediction modes, the first flag is not encoded in the coded video bitstream and can be derived from one or more inter-frame prediction models based on a default value. For example, the inter-frame prediction mode of the current block is NEAR_NEARMV; and step 1830 can include determining the first flag as a default value. In some embodiments, the default value is 0, indicating that the first delta MV of the first reference frame and the second delta MV of the second reference frame are not jointly signaled.

[0196] In some embodiments of the composite reference mode, a first flag (which may be named joint_mvd_flag) may be sent to the device to indicate whether the delta MVs of the first reference list (reference list 0) and the second reference list (reference list 1) are jointly signaled.

[0197] In some other embodiments, in response to the value of the first flag (joint_mvd_flag) indicating that the delta MVs of reference list 0 and reference list 1 are jointly signaled, only one joint delta MV (which may be named joint_delta_mv) is signaled and sent to the decoder, and the delta MV of reference list 0 and the delta MV of reference list 1 may be derived from the joint delta MV (joint_delta_mv). In response to the value of the first flag (joint_mvd_flag) indicating that the delta MV of reference list 0 and the delta MV of reference list 1 are not jointly signaled, zero, one, or two delta MVs of reference list 0 and / or reference list 1 may be signaled separately based on the inter prediction mode.

[0198] In some embodiments, the value of the first flag is 0, which may indicate that the incremental MVs of reference list 0 and reference list 1 are jointly signaled, and only one joint incremental MV is signaled and sent; and the value of the first flag is 1, which may indicate that the incremental MVs of reference list 0 and reference list 1 are not jointly signaled, and zero (or one or two) joint incremental MVs may be signaled and sent. Vice versa, in some other embodiments, the value of the first flag is 1, which may indicate that the incremental MVs of reference list 0 and reference list 1 are jointly signaled, and only one joint incremental MV is signaled and sent; and the value of the first flag is 0, which may indicate that the incremental MVs of reference list 0 and reference list 1 are not jointly signaled, and zero (or one or two) joint incremental MVs may be signaled and sent.

[0199] Referring to step 1840 , in response to the first flag indicating that the first delta MV and the second delta MV are jointly signaled, the first delta MV and the second delta MV may be derived based on the joint delta MV.

[0200] In some embodiments, when the inter-frame prediction mode of the current block is NEW_NEWMV and the first flag (e.g., joint_mvd_flag) indicates that the delta MVs of reference list 0 and reference list 1 are jointly signaled, the delta MVs of reference list 0 and / or reference list 1 can be derived from joint_delta_mv based on the POC distances of the first reference frame and the second reference frame to the current frame and the directions of the two reference frames.

[0201] In some other embodiments, the inter prediction mode of the current block is NEW_NEWMV; step 1840 may include determining the first delta MV as a joint delta MV, and determining the second delta MV by scaling the joint delta MV according to at least one of: a first picture order count (POC) distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or a directional relationship between the first reference frame and the second reference frame relative to the current frame. Here, the directional relationship is, for example, the same direction or opposite direction.

[0202] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; step 1840 may include determining the second incremental MV as a joint incremental MV, and determining the first incremental MV by scaling the joint incremental MV according to at least one of the following: a first picture order count (POC) distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or a directional relationship between the first reference frame and the second reference frame relative to the current frame.

[0203] In some embodiments, the delta MV in reference list 0 (or list 1) can always be set equal to joint_delta_mv, and the delta MV in reference list 1 (or list 0) can be scaled from joint_delta_mv according to the POC distance from the reference frame to the current frame and / or the direction of the two reference frames.

[0204] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; step 1840 may include, in response to a first absolute POC distance between the first reference frame and the current frame being greater than a second absolute POC distance between the second reference frame and the current frame: determining the first incremental MV as a joint incremental MV, and determining the second incremental MV by scaling the joint incremental MV according to at least one of the following: a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or a directional relationship between the first reference frame and the second reference frame relative to the current frame.

[0205] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; step 1840 may include, in response to a first absolute POC distance between the first reference frame and the current frame being less than a second absolute POC distance between the second reference frame and the current frame: determining the second incremental MV as a joint incremental MV, and determining the first incremental MV by scaling the joint incremental MV according to at least one of the following: a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or a directional relationship between the first reference frame and the second reference frame relative to the current frame.

[0206] In some embodiments, when the absolute POC distance between reference list 0 (or list 1) and the current frame is greater than the absolute POC distance between reference list 1 (or list 0) and the current frame, the delta MV in reference list 0 (or list 1) (which is the one with the larger absolute POC distance) may be set equal to joint_delta_mv. The delta MV in reference list 1 (or list 0) (which is the other with the smaller absolute POC distance) may be scaled from joint_delta_mv according to the POC distance from the reference frame to the current frame and / or the direction of the two reference frames.

[0207] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; and step 1840 may include, in response to the first absolute POC distance between the first reference frame and the current frame being equal to the second absolute POC distance between the second reference frame and the current frame: determining the first incremental MV as a joint incremental MV, in response to the first reference frame and the second reference frame having the same directional relationship relative to the current frame: determining the second incremental MV as a joint incremental MV, and in response to the first reference frame and the second reference frame having an opposite directional relationship relative to the current frame: determining the second incremental MV as the joint incremental MV multiplied by -1.

[0208] In some embodiments, when the absolute POC distance between reference list 1 and the current frame is the same as the absolute POC distance between reference list 0 and the current frame, the delta MV in reference list 0 can be set equal to joint_delta_mv. When the two reference frames have the same orientation, the delta MV in reference list 1 can also be set equal to joint_delta_mv. Otherwise, when the two reference frames have different orientations, the delta MV of reference list 1 is set to joint_delta_mv multiplied by -1.

[0209] In various embodiments / implementations of the present disclosure, the second incremental MV is obtained by scaling the joint incremental MV according to at least one of the following: a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or a directional relationship between the first reference frame and the second reference frame relative to the current frame, the scaling may include a linear scaling method, that is, the absolute value of the scaled incremental MV may be proportional to the ratio of the second POC distance divided by the first POC distance, and the sign of the scaled incremental MV may be determined according to the directional relationship. For one example, when the first POC distance is 4 and the second POC distance is 8 and the directional relationship between the first and second reference frames relative to the current frame is the same direction, the second incremental MV is obtained by scaling / multiplying the joint incremental MV by a factor of 2 (=8 / 4); and the second incremental MV has the same sign as the joint incremental MV because the directional relationship is the same direction. For another example, when the first POC distance is 3 and the second POC distance is -9 and the directional relationship of the first and second reference frames relative to the current frame is in opposite directions, the second incremental MV is obtained by scaling / multiplying the joint incremental MV by a factor of -3 (=-9 / 3); and the second incremental MV has an opposite sign to the joint incremental MV because the directional relationship is in opposite directions.

[0210] In some other embodiments, the inter prediction mode of the current block is NEW_NEARMV; and the second delta MV may be slightly adjusted by a predefined weight. Step 1840 may include determining the first delta MV as a joint delta MV, and determining the second delta MV by scaling the joint delta MV according to at least one of the following: a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, a directional relationship between the first reference frame and the second reference frame relative to the current frame, or a predefined weighting factor.

[0211] In some other embodiments, the inter prediction mode of the current block is NEAR_NEWMV; and the first delta MV may be slightly adjusted by a predefined weight. Step 1840 may include determining the second delta MV as a joint delta MV, and determining the first delta MV by scaling the joint delta MV according to at least one of the following: a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, a directional relationship between the first reference frame and the second reference frame relative to the current frame, or a predefined weighting factor.

[0212] In some other implementations, the predefined weighting factor is a fraction between -1 and 1.

[0213] In some other implementations, the predefined weighting factors are signaled in a high-level syntax that includes at least one of a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a tile header, a slice header, a frame header, a coding tree unit (CTU) header, or a super-block header.

[0214] In some embodiments, when the current block is in NEW_NEAR mode (or NEAR_NEW mode) and joint_mvd_flag indicates that the delta MVs of reference list 0 and reference list 1 are jointly signaled, the delta MV of reference list 0 (or list 1) can be set equal to joint_delta_mv, and the delta MV of list 1 (or list 0) can be scaled from joint_delta_mv based on one or more coded information, including but not limited to the POC distance from the reference frame to the current frame, the direction of the two reference frames, the difference between the MV predictors of the two MVs, and / or a predefined weighting factor w. In some embodiments, the predefined weighting factor can be a number between -1 and 1, such as 1 / 2. In some other embodiments, the predefined weighting factor can be signaled in a high-level syntax, including but not limited to SPS, VPS, PPS, picture header, tile header, slice header, frame header, CTU (or super block) header.

[0215] In various embodiments / implementations of the present disclosure, a weight-adjusted delta MV is obtained by scaling the joint delta MV according to at least one of the following: a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, a directional relationship between the first reference frame and the second reference frame relative to the current frame, or a predefined weighting factor. The scaling may include a linear scaling method, i.e., the absolute value of the scaled delta MV may be proportional to the ratio of the second POC distance divided by the first POC distance (and then multiplied by the predefined weighting factor). The sign of the scaled delta MV may be determined according to the directional relationship. For example, when the first POC distance is 4 and the second POC distance is 8 and the directional relationship between the first and second reference frames relative to the current frame is the same direction and the predefined weighting factor is 1 / 2, the weight-adjusted delta MV is obtained by scaling / multiplying the joint delta MV by a total factor of 1, the total factor 1 being calculated according to a factor of 2 (=8 / 4) multiplied by a weighting factor of 1 / 2; and the weight-adjusted delta MV has the same sign as the joint delta MV because the directional relationship is the same direction. For another example, when the first POC distance is 3 and the second POC distance is -9 and the directional relationship of the first and second reference frames relative to the current frame is in opposite directions and the predefined weighting factor is 1 / 2, the weight-adjusted delta MV is obtained by scaling / multiplying the joint delta MV by a total factor of -3 / 2, which is calculated based on the factor -3 (=-9 / 3) multiplied by the weighting factor 1 / 2. The weight-adjusted delta MV has an opposite sign to the joint delta MV because the directional relationship is in opposite directions.

[0216] The embodiments of the present disclosure may be used alone or in combination in any order. As needed, any steps and / or operations in any embodiment of the present disclosure may be combined or arranged in any number or order. Two or more steps and / or operations in any embodiment of the present disclosure may be performed in parallel. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium. The embodiments of the present disclosure may be applied to luminance blocks or chrominance blocks.

[0217] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 19 A computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0218] Computer software may be encoded using any suitable machine code or computer language that may be assembled, compiled, linked, or similar mechanisms to create code comprising instructions that may be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly or through interpretation, microcode execution, etc.

[0219] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, and the like.

[0220] Figure 19 The components shown for the computer system (2000) are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Neither should the configuration of components be interpreted as constituting any dependency or requirement on any one or combination of components illustrated in the exemplary embodiment of the computer system (2000).

[0221] The computer system (2000) may include certain human interface input devices. Such human interface input devices may be responsive to input from one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, taps), visual input (such as gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as voice, music, ambient sounds), images (such as scanned images, photographic images obtained from a still camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0222] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard (2001), mouse (2002), touchpad (2003), touch screen (2010), data gloves (not shown), joystick (2005), microphone (2006), scanner (2007), camera (2008).

[0223] The computer system (2000) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses through, for example, tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback from a touch screen (2010), a data glove (not shown), or a joystick (2005), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (2009), headphones (not depicted)), visual output devices (such as screens (2010), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities—some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through, for example, stereoscopic output; virtual reality glasses (not depicted), holographic displays, and smoke canisters (not depicted)), and printers (not depicted).

[0224] The computer system (2000) may also include human-accessible storage devices and their associated media, such as optical media including media (2021) such as CD / DVD ROM / RW (2020) with CD / DVD, thumb drives (2022), removable hard drives or solid-state drives (2023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), and the like.

[0225] Those skilled in the art will also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other volatile signals.

[0226] The computer system (2000) may also include an interface (2054) to one or more communication networks (2055). The network may be, for example, wireless, wired, or optical. The network may also be local, wide-area, metropolitan, vehicular and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CAN buses, and the like. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (2049) (such as, for example, a USB port of the computer system (2000)); other networks are typically integrated into the core of the computer system (2000) by attaching to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2000) can communicate with other entities. Such communications can be one-way receive-only (e.g., broadcast TV), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way, such as to other computer systems using a local area digital network or a wide area digital network. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.

[0227] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (2040) of the computer system (2000).

[0228] The core (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (2043), hardware accelerators (2044) for certain tasks, a graphics adapter (2050), and the like. These devices, along with read-only memory (ROM) (2045), random access memory (2046), and internal mass storage (2047) such as an internal non-user accessible hard drive, SSD, and the like, may be connected via a system bus (2048). In some computer systems, the system bus (2048) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices may be attached to the core's system bus (2048) directly or via a peripheral bus (2049). In one example, a screen (2010) may be connected to a graphics adapter (2050). Peripheral bus architectures include PCI, USB, and the like.

[0229] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can execute certain instructions, the combination of which can constitute the aforementioned computer code. The computer code can be stored in ROM (2045) or RAM (2046). Transient data can also be stored in RAM (2046), while permanent data can be stored in, for example, internal mass storage (2047). Fast storage and retrieval of any memory device can be enabled by using cache memory, which can be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc.

[0230] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The media and computer code may be those specially designed and constructed for the purposes of this disclosure, or they may be of a type well known and available to those skilled in the art of computer software.

[0231] As a non-limiting example, a computer system (2000) having an architecture, and in particular the core (2040), can provide functionality as a result of one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as certain memories of the core (2040) having non-volatile properties, such as core internal mass storage (2047) or ROM (2045). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2040). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the core (2040) and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of hard-wired logic or otherwise embodied in circuitry (e.g., accelerator (2044)) that may operate in place of or in conjunction with software to perform specific processes or specific portions of specific processes described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0232] Although a particular invention has been described with reference to illustrative embodiments, this description is not intended to be limiting. Based on this description, various modifications to the illustrative embodiments and additional embodiments of the present invention will be apparent to those of ordinary skill in the art. Those skilled in the art will readily recognize that these and various other modifications may be made to the exemplary embodiments illustrated and described herein without departing from the spirit and scope of the present invention. Therefore, it is contemplated that the appended claims will cover any such modifications and alternative embodiments. Certain proportions in the illustrations may be exaggerated, while other proportions may be minimized. Therefore, this disclosure and the accompanying drawings are to be considered illustrative and not restrictive.

[0233] The following is a list of abbreviations, some of which may appear in this disclosure:

[0234] JEM: Joint Exploration Model

[0235] VVC: universal video coding

[0236] BMS: Benchmark set

[0237] MV: Motion Vector

[0238] HEVC: High Efficiency Video Coding

[0239] SEI: Supplementary Enhancement Information

[0240] VUI: Video Usability Information

[0241] GOPs: Groups of Pictures

[0242] TUs: Transform Units

[0243] PUs: Prediction Units

[0244] CTUs: Coding Tree Units

[0245] CTBs: Coding Tree Blocks

[0246] PBs: Prediction Blocks

[0247] HRD: Hypothetical Reference Decoder

[0248] SNR: Signal-to-Noise Ratio

[0249] CPUs: Central Processing Units

[0250] GPUs: Graphics Processing Units

[0251] CRT: Cathode Ray Tube

[0252] LCD: Liquid-Crystal Display

[0253] OLED: Organic Light-Emitting Diode

[0254] CD: Compact Disc

[0255] DVD: Digital Video Disc

[0256] ROM: Read-Only Memory

[0257] RAM: Random Access Memory

[0258] ASIC: Application-Specific Integrated Circuit

[0259] PLD: Programmable Logic Device

[0260] LAN: Local Area Network

[0261] GSM: Global System for Mobile communications

[0262] LTE: Long-Term Evolution

[0263] CANBus: Controller Area Network Bus

[0264] USB: Universal Serial Bus

[0265] PCI: Peripheral Component Interconnect

[0266] FPGA: Field Programmable Gate Areas

[0267] SSD: solid-state drive

[0268] IC:Integrated Circuit

[0269] HDR: High Dynamic Range

[0270] SDR: Standard Dynamic Range

[0271] JVET: Joint Video Exploration Team

[0272] MPM: Most Probable Mode

[0273] WAIP: Wide-Angle Intra Prediction

[0274] CU: Coding Unit

[0275] PU: Prediction Unit

[0276] TU: Transform Unit

[0277] CTU: Coding Tree Unit

[0278] PDPC: Position dependent prediction combination

[0279] ISP: Intra Sub-Partitions

[0280] SPS: Sequence Parameter Setting

[0281] PPS: Picture Parameter Set

[0282] APS: Adaptation Parameter Set

[0283] VPS: Video Parameter Set

[0284] DPS: Decoding Parameter Set

[0285] ALF: Adaptive Loop Filter

[0286] SAO: Sample Adaptive Offset

[0287] CC-ALF: Cross-Component Adaptive Loop Filter

[0288] CDEF: Constrained Directional Enhancement Filter

[0289] CCSO: Cross-Component Sample Offset

[0290] LSO: Local Sample Offset

[0291] LR: Loop Restoration Filter

[0292] AV1: AOM Media Video 1

[0293] AV2: AOM Media Video 2

[0294] MVD: Motion Vector difference

[0295] CfL: Chroma from Luma

[0296] SDT: Semi Decoupled Tree

[0297] SDP: Semi Decoupled Partitioning

[0298] SST: Semi Separate Tree

[0299] SB: Super Block

[0300] IBC (or IntraBC): Intra Block Copy

[0301] CDF: Cumulative Density Function

[0302] SCC: Screen Content Coding

[0303] GBI: Generalized Bi-prediction

[0304] BCW: Bi-prediction with CU-level Weights

[0305] CIIP: Combined intra-inter prediction

[0306] POC: Picture Order Count

[0307] RPS: Reference Picture Set

[0308] DPB: Decoded Picture Buffer

[0309] MMVD: Merge Mode with Motion Vector Difference.

Claims

1. A video encoding method, characterized in that: The method comprises: For a current block in a current frame of video data, determining whether to apply an inter-frame prediction mode and an incremental motion vector; generating a first flag joint_mvd_flag indicating whether to jointly signal a first delta MV of a first reference frame in reference list 0 and a second delta MV of a second reference frame in reference list 1 for the current block; and The first flag joint_mvd_flag is encoded in the code stream of the video data.

2. The method according to claim 1, characterized in that The method further comprises: When the inter prediction mode of the current block is determined to be NEW_NEWMV, and the first delta MV and the second delta MV are jointly signaled: determining the first incremental MV as the joint incremental MV; The second delta MV is determined based on the first delta MV and at least one of: a first picture order count (POC) distance between a first reference frame and said current frame, a second POC distance between a second reference frame and the current frame, and The directional relationship between the first reference frame and the second reference frame relative to the current frame.

3. The method according to claim 1 or 2, characterized in that The method further comprises: signaling the joint increment MV in a code stream of the video data.

4. The method according to claim 1 or 2, characterized in that The method further comprises: When the inter prediction mode of the current block is determined to be NEW_NEWMV, and a first delta MV and a second delta MV are jointly signaled, and when a first absolute POC distance between a first reference frame and the current frame is greater than a second absolute POC distance between a second reference frame and the current frame: determining the first delta MV as the joint delta MV, and The second delta MV is determined by scaling the joint delta MV according to at least one of the following: a first POC distance between a first reference frame and the current frame, a second POC distance between a second reference frame and the current frame, and The directional relationship between the first reference frame and the second reference frame relative to the current frame.

5. The method according to claim 1 or 2, characterized in that When the inter prediction mode of the current block is determined to be NEW_NEWMV, and a first delta MV and a second delta MV are jointly signaled, and when a first absolute POC distance between a first reference frame and the current frame is less than a second absolute POC distance between a second reference frame and the current frame: determining the second delta MV as the joint delta MV, and The first delta MV is determined by scaling the joint delta MV according to at least one of the following: The first POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, and The directional relationship between the first reference frame and the second reference frame relative to the current frame.

6. The method according to claim 1 or 2, characterized in that When the inter prediction mode of the current block is determined to be NEW_NEWMV, and the first delta MV and the second delta MV are jointly signaled: determining the first delta MV as a joint delta MV when a first absolute POC distance between the first reference frame and the current frame is equal to a second absolute POC length between the second reference frame and the current frame; When the first reference frame and the second reference frame have the same direction relationship relative to the current frame, determining the second incremental MV as a joint incremental MV; as well as When the first reference frame and the second reference frame have opposite directional relationships relative to the current frame, the second delta MV is determined as the joint delta MV multiplied by -1.

7. The method according to claim 1 or 2, characterized in that When the inter prediction mode of the current block is determined to be NEW_NEARMV, and the first delta MV and the second delta MV are jointly signaled: determining the first delta MV as the joint delta MV, and The second delta MV is determined by scaling the joint delta MV according to at least one of: The first POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, The directional relationship between the first reference frame and the second reference frame relative to the current frame, and Predefined weighting factors.

8. The method according to claim 1 or 2, characterized in that When the inter prediction mode of the current block is determined to be NEAR_NEWMV, and a first delta MV and a second delta MV are jointly signaled: determining the second delta MV as the joint delta MV, and The first delta MV is determined by scaling the joint delta MV according to at least one of: The first POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, The directional relationship between the first reference frame and the second reference frame relative to the current frame, and Predefined weighting factors.

9. The method according to claim 8, characterized in that The predefined weighting factor is a fraction between -1 and 1.

10. The method according to claim 1 or 2, characterized in that The predefined weighting factors are signaled in a high-level syntax, the high-level syntax comprising at least one of a sequence parameter set (SPS), a video parameter set (VPS), a picture parameter set (PPS), a picture header, a tile header, a slice header, a frame header, a coding tree unit (CTU) header, or a super-block header.

11. The method according to claim 1 or 2, characterized in that When the inter prediction mode of the current block is determined to be NEAR_NEARMV, and the first delta MV and the second delta MV are jointly signaled, the first flag joint_mvd_flag is determined to be a default value.

12. The method according to claim 11, characterized in that A default value equal to 0 indicates that the first delta MV of the first reference frame and the second delta MV of the second reference frame are not jointly signaled.

13. A method for storing a video stream, characterized in that: Execute the video encoding method according to any one of claims 1 to 12 to generate a video stream, and store the video stream.

14. A method for transmitting a video stream, characterized in that: Execute the video encoding method according to any one of claims 1 to 12 to generate a video stream, and transmit the video stream.

15. A method for processing a video block, comprising converting the video block into a video stream, wherein the video stream comprises: A first indication for determining whether to apply an inter-prediction mode and an incremental motion vector to a video block; a first flag joint_mvd_flag, the first flag joint_mvd_flag indicating whether to jointly signal a first delta MV of a first reference frame in reference list 0 and a second delta MV of a second reference frame in reference list 1 for the current block; When the first flag joint_mvd_flag indicates that the first incremental MV and the second incremental MV are jointly signaled, the video code stream further includes a joint incremental MV for the first incremental MV and the second incremental MV.

16. An electronic device, characterized in that: include: at least one memory configured to store program code; as well as At least one processor is configured to read the program code and operate according to instructions of the program code to execute the method according to any one of claims 1 to 15.

17. A computer-readable storage medium storing a computer program / instruction and a video stream, wherein: When the computer program / instruction is executed by a processor, the steps of the video encoding method according to any one of claims 1 to 12 are implemented to generate the video code stream.