Methods, Apparatuses, and Storage Media for Video Encoding and Decoding
By extracting and processing the motion vector differences between video frames, video encoding and decoding methods improve the compression efficiency of video data, and solve the problem of high demand for video data storage and transmission in the prior art.
Patent Information
- Application Number
- CN202280008009.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-22
- Filing Date
- 2022-04-15
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-15
AI Technical Summary
When existing video encoding technologies deal with motion vector differences between video frames, they are inefficient and difficult to effectively compress video data, resulting in high bandwidth and storage space requirements.
A video encoding and decoding method is proposed. By extracting the inter prediction mode and joint incremental motion vector in the current frame, it is determined whether it is necessary to jointly signal the incremental motion vector, and then derive the incremental motion vector and decode the current frame.
Improve the efficiency of video encoding and decoding, reduce the redundant information of video data, and reduce bandwidth and storage space requirements.
Smart Images

Figure CN116584097B_ABST
Abstract
Description
[0001] INCORPORATION BY REFERENCE
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 245,655, filed on September 17, 2021, which is hereby incorporated by reference in its entirety. This application also claims the benefit of priority to U.S. Non-Provisional Application No. 17 / 700,745, filed on March 22, 2022, which is hereby incorporated by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to video encoding and / or decoding techniques, and more particularly to an improved design and signaling of joint motion vector differences for encoding and / or decoding. BACKGROUND OF THE DISCLOSURE
[0004] The background description provided herein is intended to present the background of the present application as a whole. The extent to which the work of the presently named inventors, described in the background section and in various aspects of this specification, was carried out does not indicate that it was prior art at the time of filing of the present application, and is never expressly or implicitly admitted to be prior art of the present application.
[0005] Inter-picture prediction with motion compensation can be used for video encoding and decoding. Uncompressed digital video may include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated full-sampled or subsampled chrominance samples. The series of pictures has a fixed or variable picture rate (or frame rate), such as 60 pictures per second or 60 frames per second. Uncompressed video has a specific bitrate requirement. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and a chrominance subsampling of 4:2:0, with 8 bits per pixel per color channel, requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GB of storage space.
[0006] One purpose of video encoding and decoding is to reduce the redundant information of the uncompressed input video signal through compression. Video compression can help reduce the requirements for the above-mentioned bandwidth and / or storage space, and in some cases, can reduce it by two or more orders of magnitude. Lossless compression, lossy compression, and combinations of both can be used. Lossless compression refers to a technique in which an exact copy of the original signal is reconstructed from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully preserved during encoding and cannot be fully recovered during decoding. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be usable for the intended application, despite some information loss. In the case of video, lossy compression is widely used in many applications. The allowable amount of distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable through a specific encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher allowable distortion generally allows an encoding algorithm with higher losses and higher compression ratios.
[0007] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.
[0008] Video codec technology can include known intra-frame encoding techniques. In intra-frame encoding, sample values are represented without referring to samples or other data of previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in the intra-frame mode, the picture can be called an intra-frame picture. Intra-frame pictures and their derivatives (such as independent decoder refresh pictures) can be used to reset the decoder state and can thus be used as the first picture in an encoded video bitstream and a video session, or as a still image. Then, the samples of the blocks after intra-frame prediction can be transformed into the frequency domain, and the transform coefficients thus generated can be quantized before entropy coding. Intra-frame prediction represents a technique that minimizes the sample values in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding for a given quantization step size.
[0009] As is known from, for example, MPEG-2 generation coding techniques, traditional intra-frame coding does not use intra-frame prediction. However, some newer video compression techniques include attempts to encode / decode blocks based on, for example, surrounding sample data and / or metadata that are obtained during spatially adjacent encoding and / or decoding and that precede in decoding order the data block being intra-frame encoded or decoded. Such techniques are hereafter referred to as "intra-frame prediction" techniques. Note that in at least some cases, intra-frame prediction uses only reference data from the current picture in reconstruction and not reference data from other reference pictures.
[0010] There can be many different forms of intra-frame prediction. When more than one such technique is available in a given video coding technique, the technique used can be referred to as an intra-frame prediction mode. One or more intra-frame prediction modes can be provided in a particular codec. In some cases, a mode can have sub-modes and / or can be associated with various parameters, and the mode / sub-mode information and intra-frame coding parameters for a video block can be included in a mode codeword and can be encoded separately or jointly. Which codeword is used for a given mode, sub-mode, and / or parameter combination can affect the coding efficiency gain through intra-frame prediction, and the same is true for the entropy coding technique used to convert the codeword into a bitstream.
[0011] A certain mode of intra-frame prediction was introduced with H.264, modified in H.265, and further modified in newer coding techniques such as Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Model Set (BMS). Generally, for intra-frame prediction, available adjacent sample values can be used to form a predictor block. For example, the available values of a particular set of adjacent samples along a particular direction and / or row can be copied into the predictor block. The reference to the direction used can be encoded in the bitstream or can itself be predicted.
[0012] Referring Figure 1A , depicted in the lower right is a subset of 9 predictor directions specified among the 33 possible intra-frame predictor directions in H.265 (corresponding to the 33 angular modes of the 35 intra-frame modes specified in H.265). The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the directions according to which the sample at 101 is predicted using adjacent samples. For example, arrow (102) indicates that sample (101) is predicted based on one or more adjacent samples in the upper right at a 45-degree angle to the horizontal direction. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more adjacent samples in the lower left of sample (101) at a 22.5-degree angle to the horizontal direction.
[0013] Still referring Figure 1A, a square block (104) including 4×4 samples is shown in the upper left corner (represented by a thick dashed line). The square block (104) consists of 16 samples, and each sample is labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y dimension and the X dimension within the block (104). Since the block is of 4×4 sample size, S44 is located in the lower right corner. Example reference samples following a similar numbering scheme are also shown. The reference samples are labeled with "R", its Y position (e.g., row index) relative to the block (104), and its X position (e.g., column index). In H.264 and H.265, neighboring prediction samples adjacent to the block in reconstruction are used.
[0014] Intra prediction within the picture of block 104 can start by copying the reference sample values from neighboring samples according to the signal - notified prediction direction. For example, assume that the encoded video bitstream includes signaling, and for this block 104, the signaling indicates the prediction direction of arrow (102) - that is, to predict samples based on one or more prediction samples in the upper - right direction at a 45 - degree angle to the horizontal direction. In such cases, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then sample S44 is predicted based on reference sample R08.
[0015] In some cases, for example, through interpolation, the values of multiple reference samples can be combined to calculate a reference sample, especially when the direction is not divisible by 45 degrees.
[0016] As video coding technology continues to develop, the number of possible directions increases. For example, in H.264 (in 2003), 9 different directions can be used for intra prediction. This increased to 33 in H.265 (in 2013), and JEM / VVC / BMS can support up to 65 directions at the time of this disclosure. Experimental studies have been conducted to help identify the most suitable intra - prediction directions, and certain techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, accepting a certain bit cost for the directions. Additionally, the direction itself can sometimes be predicted based on the neighboring directions used for intra prediction of already - decoded neighboring blocks.
[0017] Figure 1B A schematic diagram (180) depicting 65 intra - prediction directions according to JEM is shown to illustrate the increase in the number of prediction directions in various coding techniques over time.
[0018] The manner in which bits representing intra prediction directions are mapped to prediction directions in an encoded video bitstream can vary with different video coding techniques; and can range, for example, from a simple direct mapping from prediction directions to intra prediction modes, to codewords, to complex adaptive schemes involving most probable modes and the like. However, in all cases, there may be certain directions for intra prediction in the video content that are statistically less likely to occur than some other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, those less likely directions will be representable by a larger number of bits than the more likely directions.
[0019] Inter-picture prediction or inter-frame prediction can be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) can be used for prediction of a newly reconstructed picture or picture portion (e.g., block) after spatial shifting in the direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture can be the same as the picture in the current reconstruction. The MV can have two dimensions X and Y, or three dimensions, where the third dimension is an indication of the reference picture in use (approximate temporal dimension).
[0020] In some video compression techniques, the current MV applicable to a certain region of sample data can be predicted from other MVs, for example, from those other MVs that are related to other regions of sample data spatially adjacent to the region in the reconstruction and that are prior to the current MV in the decoding order. Doing so can significantly reduce the total amount of data required to encode the MVs by relying on removing redundancy in the related MVs, thereby increasing the compression efficiency. MV prediction can be performed effectively, for example, because when encoding an input video signal derived from a camera (referred to as natural video), there is a statistical likelihood that regions larger than the region to which a single MV applies move in a similar direction in the video sequence. Thus, in some cases, similar motion vectors derived from the MVs of adjacent regions can be used for prediction. This results in the actual MV of a given region being similar or identical to the MV predicted from the surrounding MVs. After entropy coding, such MVs can in turn be represented by a smaller number of bits than would be used if the MVs were directly encoded rather than predicted from one or more adjacent MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, the MV prediction itself may be lossy, for example, due to rounding errors when calculating the predicted value from several surrounding MVs.
[0021] Various MV prediction mechanisms are described in H.265 / HEVC (Recommendation ITU-T H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms specified in H.265, the technique described in this application is what is hereinafter referred to as "spatial merge".
[0022] Please refer to Figure 2 , the current block (201) includes samples that have been discovered by the encoder during the motion search process, and the samples can be predicted based on a previous block of the same size that has generated a spatial offset. Additionally, the MV can be derived from the metadata associated with one or more reference pictures instead of directly encoding the MV. For example, using the MV associated with any one of the five surrounding samples A0, A1 and B0, B1, B2 (corresponding to 202 to 206 respectively), the MV is derived from the metadata of the nearest reference picture (in decoding order). In H.265, the MV prediction can use the prediction value of the same reference picture that is also used by adjacent blocks. Summary of the Invention
[0023] This disclosure describes various embodiments of methods, apparatuses, and storage media for video encoding and decoding.
[0024] According to one aspect, embodiments of the present disclosure provide a method for video decoding. The method includes receiving an encoded video bitstream. The method further includes extracting, from the encoded video bitstream, an inter prediction mode and a joint delta motion vector MV of a current block in a current frame; extracting a first flag from the encoded video bitstream, the first flag indicating whether a first delta MV of a first reference frame in reference list 0 and a second delta MV of a second reference frame in reference list 1 are jointly signaled; in response to the first flag indicating that the first delta MV and the second delta MV are jointly signaled, deriving the first delta MV and the second delta MV based on the joint delta MV; and decoding the current block in the current frame based on the first delta MV and the second delta MV. The present application also discloses a method for video encoding, including: receiving an encoded video bitstream; signaling, in the encoded video bitstream, an inter prediction mode and a joint delta motion vector MV (joint_delta_mv) of a current block in a current frame; signaling, in the encoded video bitstream, a first flag (joint_mvd_flag), the first flag (joint_mvd_flag) indicating whether a first delta MV of a first reference frame in reference list 0 and a second delta MV of a second reference frame in reference list 1 are jointly signaled; wherein, when the first flag indicates that the first delta MV and the second delta MV are jointly signaled, the joint delta MV is used to derive the first delta MV and the second delta MV, and the first delta MV and the second delta MV are used to decode the current block in the current frame.
[0025] According to another aspect, embodiments of the present disclosure provide an apparatus for video encoding and / or decoding. The apparatus includes a memory storing instructions; and a processor communicatively coupled to the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above methods for video decoding and / or encoding.
[0026] Aspects of the present disclosure also provide a video encoding or decoding device or apparatus, including circuitry configured to perform any of the above methods.
[0027] In another aspect, embodiments of the present disclosure provide a non-volatile computer-readable medium storing instructions that, when executed by a computer for video encoding and decoding, cause the computer to perform the above methods for video encoding and decoding.
[0028] The above aspects and other aspects and their implementations are described in more detail in the drawings, the specification, and the claims. Description of the Drawings
[0029] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0030] Figure 1A A schematic diagram showing an example subset of intra prediction direction modes.
[0031] Figure 1B An illustration showing an exemplary intra prediction direction.
[0032] Figure 2 A schematic diagram showing spatial merge candidates for motion vector prediction for the current block and its surroundings in one example.
[0033] Figure 3 A schematic diagram showing a simplified block diagram of a communication system according to an example embodiment.
[0034] Figure 4 A schematic diagram showing a simplified block diagram of a communication system according to an example embodiment.
[0035] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment.
[0036] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment.
[0037] Figure 7 A block diagram of a video encoder according to another example embodiment.
[0038] Figure 8 A block diagram of a video decoder according to another example embodiment.
[0039] Figure 9 A scheme for coding block partitioning according to an example embodiment of the present disclosure.
[0040] Figure 10 Another scheme for coding block partitioning according to an example embodiment of the present disclosure.
[0041] Figure 11 Another scheme for coding block partitioning according to an example embodiment of the present disclosure.
[0042] Figure 12 An example of dividing a basic block into coding blocks according to an example partitioning scheme.
[0043] Figure 13 An example ternary partitioning scheme.
[0044] Figure 14 An example quadtree binary tree coding block partitioning scheme.
[0045] Figure 15 A scheme for partitioning an encoding block into a plurality of transform blocks and the encoding order of the transform blocks according to an exemplary embodiment of the present disclosure is shown.
[0046] Figure 16 Another scheme for partitioning an encoding block into a plurality of transform blocks and the encoding order of the transform blocks according to an exemplary embodiment of the present disclosure is shown.
[0047] Figure 17 Another scheme for partitioning an encoding block into a plurality of transform blocks according to an exemplary embodiment of the present disclosure is shown.
[0048] Figure 18 A flowchart of a method according to an exemplary embodiment of the present disclosure is shown.
[0049] Figure 19 A schematic diagram of a computer system according to an exemplary embodiment of the present disclosure is shown. Detailed Description of the Invention
[0050] The present invention will now be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, it should be noted that the present invention can be embodied in various different forms, and thus, the subject matter covered or claimed is intended to be construed as not limited to any of the embodiments set forth below. It should also be noted that the present invention can be embodied as a method, apparatus, component, or system. Therefore, the embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.
[0051] Throughout the specification and claims, terms may have subtle meanings suggested or implied in the context that go beyond the explicitly stated meanings. As used herein, the phrase "in one embodiment" or "in some embodiments" does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" does not necessarily refer to different embodiments. Similarly, as used herein, the phrase "in one implementation" or "in some implementations" does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" does not necessarily refer to different implementations. For example, the claimed subject matter is intended to include combinations of all or part of the exemplary embodiments / implementations.
[0052] In general, terms may be understood, at least in part, based on their use in context. For example, terms such as "and", "or", or "and / or" as used herein may have multiple meanings, which may depend, at least in part, on the context in which such terms are used. Generally, "or" when used in connection with a list, such as A, B, or C, is intended to mean A, B, and C, used in an inclusive sense herein, as well as A, B, or C, used in an exclusive sense herein. Additionally, the terms "one or more" or "at least one" as used herein, may, at least in part, depend on context, be used to describe any feature, structure, or characteristic in a singular sense, or may be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as "a", "an", or "the" may also be understood to convey either a singular usage or a plural usage, at least in part, depending on context. Further, the terms "based on" or "determined by" may be understood to not necessarily denote an exclusive set of factors, but may allow for the existence of additional factors that are not necessarily explicitly described again, at least in part, depending on context.
[0053] Figure 3 is a simplified block diagram of a communication system (300) according to an embodiment disclosed in the present application. The communication system (300) includes a plurality of terminal devices, and the terminal devices can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first terminal device (310) and a second terminal device (320) interconnected through a network (350). In Figure 3 an embodiment of, the first terminal device (310) and the second terminal device (320) perform unidirectional data transmission. For example, the first terminal device (310) may encode video data (e.g., a video picture stream captured by the first terminal device (310)) for transmission through the network (350) to the second terminal device (320). The encoded video data is transmitted in the form of one or more encoded video bitstreams. The second terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to recover the video data, and display video pictures based on the recovered video data. Unidirectional data transmission is relatively common in applications such as media services.
[0054] In another embodiment, a communication system (300) includes a third terminal device (330) and a fourth terminal device (340) that perform two-way transmission of encoded video data, which may be implemented, for example, during a video conference. For two-way data transmission, each of the third terminal device (330) and the fourth terminal device (340) may encode video data (such as a video picture stream captured by the terminal device) for transmission over a network (350) to the other of the third terminal device (330) and the fourth terminal device (340). Each of the third terminal device (330) and the fourth terminal device (340) may also receive the encoded video data transmitted by the other of the third terminal device (330) and the fourth terminal device (340), may decode the encoded video data to recover the video data, and may display video pictures on an accessible display device based on the recovered video data.
[0055] In Figure 3 the embodiment, the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340) may be servers, personal computers, and smart phones, but the scope of application of the basic principles disclosed in this application is not limited thereto. The embodiments disclosed in this application are applicable to desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, and the like. The network (350) represents any number or type of network that conveys encoded video data between the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched, packet-switched, and / or other types of channels. The network may include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless explicitly explained herein, the architecture and topology of the network (350) may be immaterial to the operations disclosed in this application.
[0056] As an example, Figure 4 shows the placement of a video encoder and a video decoder in a video streaming environment. The subject matter disclosed in this application is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, compressed video storage on digital media including CDs, DVDs, memory sticks, and the like.
[0057] A video streaming system may include an acquisition subsystem (413), and the acquisition subsystem may include a video source (401) such as a digital camera, etc., to create an uncompressed video picture or image stream (402). In an embodiment, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. Compared with the encoded video data (404) (or the encoded video bitstream), the uncompressed video picture stream (402) is depicted as a thick line to emphasize the high-data-volume video picture stream. The video picture stream (402) may be processed by an electronic device (420), and the electronic device (420) includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the uncompressed video picture stream (402), the encoded video data (404) (or the encoded video bitstream (404)) is depicted as a thin line to emphasize the lower-data-volume encoded video data (404) (or the encoded video bitstream (404)), which may be stored on a streaming server (405) for future use or directly used for downstream video devices (not shown). One or more streaming client subsystems, such as Figure 4 the client subsystem (406) and the client subsystem (408) in, may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an uncompressed output video picture stream (411) that can be presented on a display (412) (such as a display screen) or another presentation device (not depicted). The video decoder 410 may be configured to perform some or all of the various functions described in the present disclosure. In some streaming systems, the encoded video data (404), the video data (407), and the video data (409) (such as a video bitstream) may be encoded according to certain video coding / compression standards. Embodiments of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application may be used in the context of the VVC standard and other video coding standards.
[0058] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may further include a video encoder (not shown).
[0059] Figure 5 is a block diagram of a video decoder (510) according to an embodiment disclosed below in the present application. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 the video decoder (410) in the embodiment.
[0060] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same embodiment or another embodiment, one encoded video sequence is decoded at a time, wherein the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or images. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data and other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective processing circuits (not labeled). The receiver (531) may separate the encoded video sequence from other data. To prevent network jitter, a buffer memory (515) may be configured between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be provided outside the video decoder (510) and separated from the video decoder (510) (not labeled). In other applications, a buffer memory (not labeled) is provided outside the video decoder (510) to, for example, prevent network jitter, and another buffer memory (515) may exist inside the video decoder (510) to, for example, handle the playback timing. And when the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous synchronous network, it may also be possible not to configure the buffer memory (515), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and the buffer memory may be relatively large. Such a buffer memory may have an adaptive size and may be at least partially implemented in an operating system or a similar element (not labeled) outside the video decoder (510).
[0061] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a display device (such as a display screen), like the display device (512), which may or may not be part of an electronic device (530), but can be coupled to the electronic device (530), as shown in Figure 5 shown. The control information for the display device can be a Supplemental Enhancement Information (SEI message) or a parameter set segment (not labeled) of Video Usability Information (VUI). The parser (520) may perform parsing / entropy decoding on the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence can be carried out according to video coding techniques or standards and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (e.g., transform coefficients), quantizer parameter values, motion vectors, and so on.
[0062] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0063] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different processing or functional units. Which units are involved and the way they are involved can be controlled by the subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and the multiple processing or functional units below are not described.
[0064] In addition to the functional blocks already mentioned, the video decoder (510) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these functional units interact closely with each other and can be integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, a conceptual subdivision of the functional units is adopted in the following disclosure.
[0065] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives the quantized transform coefficients as symbols (521) and control information from the parser (520), including information indicating which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output blocks including sample values, and the sample values may be input into an aggregator (555).
[0066] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to intra-coded blocks; for example, blocks that do not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed surrounding block information stored in the current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) adds the prediction information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0067] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (553) may access the reference picture memory (557) to extract samples for inter-picture prediction. After motion-compensating the extracted samples according to the symbol (521), these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 is referred to as residual samples or a residual signal), thereby generating output sample information. The motion compensation prediction unit (553) obtaining the prediction samples from an address within the reference picture memory (557) may be controlled by a motion vector, and the motion vector is in the form of the symbol (521) for use by the motion compensation prediction unit (553), where the symbol (521) includes, for example, X, Y components (displacements) and a reference picture component (time). Motion compensation may also include interpolation of the sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, and motion compensation may also be associated with a motion vector prediction mechanism, etc.
[0068] The output samples of the aggregator (555) may be employed in the loop filter unit (554) by various loop filtering techniques. Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and the parameters may be used in the loop filter unit (556) as symbols (521) from the parser (520). However, in other embodiments, the video compression techniques may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. Some types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in further detail below.
[0069] The output of the loop filter unit (556) may be a sample stream that may be output to the display device (512) and stored in the reference picture memory (557) for subsequent inter-picture prediction.
[0070] Once fully reconstructed, certain encoded pictures may be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct subsequent encoded pictures.
[0071] The video decoder (510) may perform decoding operations according to a predetermined video compression technique employed, for example, in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, megasamples per second), the maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0072] In one embodiment, the receiver (531) may receive additional (redundant) data together with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (510) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0073] Figure 6 is a block diagram of a video encoder (603) according to an embodiment disclosed in the present application. The video encoder (603) is provided in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the embodiment.
[0074] The video encoder (603) may receive video samples from a video source (601) (not Figure 6 part of the electronic device (620) in the embodiment), and the video source may capture video images to be encoded by the video encoder (603). In another embodiment, the video source (601) may be implemented as part of the electronic device (620).
[0075] A video source (601) can provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (603). The digital video sample stream can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 YCrCb, RGB, XYZ, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures or images, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc. being used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0076] According to an embodiment, the video encoder (603) can encode and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below. For simplicity, the couplings are not labeled in the figure. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, the maximum allowed motion vector search range, etc. The controller (650) can be used for other suitable functions that relate to optimizing the video encoder (503) for a particular system design.
[0077] In some embodiments, the video encoder (603) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data, even though the embedded decoder 633 processes the encoded video stream through the source encoder 630 without entropy encoding (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is used to improve the encoding quality.
[0078] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder, for example, which has been described in detail above in connection with Figure 5 the video decoder (510). However, briefly referring additionally to Figure 5 , when the symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633) of the encoder.
[0079] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, this application sometimes focuses on the decoder operation, which is related to the decoding part of the encoder. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. Only some areas or aspects of the encoder will be described in more detail below.
[0080] During operation, in some embodiments, the source encoder (630) may perform motion-compensated predictive coding. Referring to one or more previously encoded pictures designated as "reference pictures" in a video sequence, the motion-compensated predictive coding performs predictive coding on an input picture. In this way, the encoding engine (632) encodes the difference (residual) between a pixel block in a color channel of the input picture and a pixel block of a reference picture, which may be selected as a prediction reference for the input picture. The term "residual" and its adjective form "residual" may be used interchangeably.
[0081] The local video decoder (633) may decode the encoded video data of a picture that may be designated as a reference picture, based on symbols created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 6 not shown), the reconstructed video sequence is typically a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on a reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture, which has the same content (no transmission errors) as the reconstructed reference picture that will be obtained by a remote video decoder.
[0082] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata, such as a reference picture motion vector, block shape, etc., that may serve as an appropriate prediction reference for the new picture. The predictor (635) may operate on a per-pixel basis for sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (635), it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0083] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0084] The outputs of all the above functional units may be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0085] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over the communication channel (660), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can merge the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0086] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture can typically be assigned to any of the following picture types:
[0087] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.
[0088] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0089] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0090] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined according to the encoding assignments of the corresponding pictures applied to the blocks. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures. For other purposes, source pictures or pictures in intermediate processing can be subdivided into other types of blocks. The partitioning of encoded blocks and other types of blocks may or may not follow the same manner, as described in further detail below.
[0091] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0092] In an embodiment, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, SEI messages, VUI parameter set fragments, etc.
[0093] The captured video can be multiple source pictures (video pictures) in a time series. Intra picture prediction (often simplified to intra prediction) utilizes the spatial correlation within a given picture, while inter picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0094] In some embodiments, bidirectional prediction techniques can be used for inter - picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past or future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be jointly predicted by a combination of the first reference block and the second reference block.
[0095] In addition, the merge mode technique can be used in inter - picture prediction to improve coding efficiency.
[0096] According to some embodiments disclosed in the present application, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis. For example, a picture in a video picture sequence is segmented into coding tree units (CTUs) for compression. The CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three parallel coding tree blocks (CTBs): one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) in a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs. Each of one or more 32×32 blocks can be further divided into 4 16×16 - pixel CUs. In one embodiment, each CU can be analyzed during encoding to determine the prediction type for the CU among various prediction types, such as inter - frame prediction type or intra - frame prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed on a prediction - block basis. The splitting of the CU into PUs (or PBs of different color channels) can be performed in various spatial patterns. For example, a luminance or chrominance PB can include a matrix of sample values (e.g., luminance values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, and 16x8 samples, etc.
[0097] Figure 7FIG. is a diagram of a video encoder (703) according to another illustrative embodiment disclosed in the present application. The video encoder (703) is configured to receive sample values within a current video picture in a sequence of video pictures in a processing block (e.g., a prediction block), and encode the processing block into an encoded picture that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used to replace Figure 4 the video encoder (303) in the embodiment.
[0098] In one embodiment, the video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. Then, the video encoder (703) uses, for example, rate-distortion optimization (RDO) to determine whether to use an intra mode, an inter mode, or a bi-predictive mode to encode the processing block. When it is determined to encode the processing block in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into the encoded picture; and when it is determined to encode the processing block in the inter mode or the bi-predictive mode, the video encoder (703) may use inter prediction or bi-predictive techniques respectively to encode the processing block into the encoded picture. In some illustrative embodiments, the merge mode may be a sub-mode of inter picture prediction, where a motion vector is derived from one or more motion vector prediction values without relying on encoded motion vector components external to the prediction values. In certain other illustrative embodiments, there may be motion vector components applicable to the subject block. Thus, the video encoder (703) includes components explicitly shown in Figure 7 such as a mode decision module (not shown) for determining the prediction mode of the processing block.
[0099] In Figure 7 the embodiment of, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in the exemplary arrangement in Figure 7 FIG..
[0100] The inter encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., redundancy information description, motion vector, merge mode information according to inter coding techniques), and calculate an inter prediction result (e.g., a prediction block) based on the inter prediction information using any suitable technique. In some embodiments, the reference picture is decoded using a decoding unit 633 (as shown in Figure 6 the example encoder 620 embedded inFigure 7 As shown by the residual decoder 728 (see details below), a decoded reference picture decoded based on the encoded video information.
[0101] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture in some cases, generate quantization coefficients after transformation, and also generate intra prediction information in some cases (e.g., based on intra prediction direction information of one or more intra coding techniques). The intra encoder (722) calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.
[0102] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines a prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra prediction information and add the intra prediction information to the bitstream; and when the prediction mode for the block is the inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter prediction information and add the inter prediction information to the bitstream.
[0103] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result of a block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is configured to convert the residual data from the time domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various illustrative embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra prediction information. The decoded blocks are appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.
[0104] An entropy encoder (725) is used to format a bitstream to produce an encoded block and perform entropy encoding. The entropy encoder (725) is configured to include various information in the bitstream. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. It should be noted that when encoding a block in the merge sub-mode of the inter mode or the bi-prediction mode, there is no residual information.
[0105] Figure 8 is a diagram of a video decoder (810) according to another embodiment disclosed in the present application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (810) is used to replace Figure 4 the video decoder (410) in the embodiment.
[0106] In Figure 8 an embodiment, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as schematically arranged in Figure 8 .
[0107] The entropy decoder (871) can be used to reconstruct certain symbols according to the encoded picture, and these symbols represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode used to encode the block (e.g., intra mode, inter mode, bi-prediction mode, merge sub-mode, or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata for prediction by the intra decoder (872) or the inter decoder (880), residual information in the form of, for example, quantized transform coefficients, and so on. In an embodiment, when the prediction mode is the inter or bi-prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is the intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).
[0108] The inter decoder (880) is used to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0109] The intra decoder (872) is used to receive the intra prediction information and generate a prediction result based on the intra prediction information.
[0110] The residual decoder (873) is used to perform inverse quantization to extract the dequantized transform coefficients, and processes the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also use certain control information (for obtaining the quantizer parameter QP), which may be provided by the entropy decoder (871) (the data path is not marked as this is only low-volume control information).
[0111] The reconstruction module (874) is used to combine the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, and the reconstructed block forms a part of the reconstructed picture, and the reconstructed picture may be a part of the reconstructed video. It should be noted that other appropriate operations such as deblocking operations may be performed to improve the visual quality.
[0112] It should be noted that any suitable technology may be used to implement the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810). In an embodiment, one or more integrated circuits may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810).
[0113] Turn to block partitioning for encoding and decoding. General partitioning can start from a basic block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After dividing or partitioning the basic block following any one or a combination of the example partitioning processes or other processes described below, a final set of partitions or encoding / decoding blocks can be obtained. Each of these partitions can be at one of various partition levels in a partition hierarchy and can have various shapes. Each of the partitions can be referred to as an encoding / decoding block (CB). For the various example partitioning implementations described further below, each resulting CB can have any of the allowed sizes and partition levels. Such partitions are called encoding / decoding blocks because they can form units for which some basic encoding / decoding decisions can be made and for which encoding / decoding parameters can be optimized and determined for concurrent signaling in an encoded video bitstream. The highest or deepest level in the final partition represents the depth of the encoding / decoding block partitioning structure of the tree. The encoding / decoding blocks can be luma encoding / decoding blocks or chroma encoding / decoding blocks. The CB tree structure for each color can be referred to as an encoding / decoding block tree (CBT).
[0114] The encoding / decoding blocks of all color channels can be collectively referred to as encoding / decoding units (CUs). The hierarchical structures of all color channels can be collectively referred to as coding tree units (CTUs). The partitioning patterns or structures of the various color channels in a CTU can be the same or can be different.
[0115] In some implementations, the partitioning tree scheme or structure for the luma and chroma channels may not have to be the same. In other words, the luma and chroma channels can have separate coding tree structures or patterns. Further, whether the luma and chroma channels use the same or different encoding / decoding partitioning tree structures and the actual encoding / decoding partitioning tree structures to be used can depend on whether the slice being encoded is a P, B, or I slice. For example, for an I slice, the chroma channel and the luma channel can have their respective encoding / decoding partitioning tree structures or encoding / decoding partitioning tree structure patterns, while for a P or B slice, the luma and chroma channels can share the same encoding / decoding partitioning tree scheme. When applying separate encoding / decoding partitioning tree structures or patterns, the luma channel can be partitioned into CBs by one encoding / decoding partitioning tree structure, and the chroma channel can be partitioned into chroma CBs by another encoding / decoding partitioning tree structure.
[0116] In some example implementations, a predefined partitioning pattern can be applied to the basic block. As Figure 9As shown, an exemplary 4-way partition tree can be employed. The 4-way partition tree can start from a first predefined level (e.g., 64×64 block level or other sizes, such as the basic block size), and the basic block can be hierarchically partitioned downward to a predefined lowest level (e.g., 4×4 level). For example, the basic block can be limited to four predefined partition options or patterns indicated by 902, 904, 906, and 908, where the partition specified as R allows recursive partitioning. Figure 9 The same partition options shown can be repeated at a lower scale until the lowest level (e.g., 4x4 level). In some embodiments, additional restrictions can be applied to Figure 9 the partition scheme. In Figure 9 the embodiments, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) can be allowed, but they may not be allowed to be recursive, while square partitions are allowed to be recursive. If needed, the final set of coded blocks is generated according to the Figure 9 recursive partitioning. The coding tree depth can be further defined to indicate the depth of splitting from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 64x64 block) can be set to 0, and after the root block is further split once according to Figure 9 , the coding tree depth is incremented by 1. For the above scheme, the maximum or deepest level from the 64x64 basic block to the 4x4 minimum partition is 4 (starting from level 0). Such a partition scheme can be applied to one or more color channels. Each color channel can be partitioned independently according to the Figure 9 scheme (e.g., the partition pattern or option in the predefined pattern can be independently determined for each color channel at each hierarchical level). Alternatively, two or more color channels can share the Figure 9 same hierarchical pattern tree (e.g., the same partitioning pattern or option in the predefined pattern can be selected for two or more color channels at each hierarchical level).
[0117] Figure 10 shows another exemplary predefined partition pattern that allows recursive partitioning to form a partition tree. As Figure 10 shown, an exemplary 10-way partition structure or pattern can be predefined. The root block can start at a predefined level (e.g., from the 128×128 level, or the basic block at the 64×64 level). Figure 10 The exemplary partition structure includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Figure 10 In the second row of , 1002, 1004, 1006, and 1008 indicate partition types with 3 sub-partitions, which can be referred to as "T-shaped" partitions. The "T-shaped" partitions 1002, 1004, 1006, and 1008 can be referred to as left T-shaped, top T-shaped, right T-shaped, and bottom T-shaped. In some exemplary embodiments, Figure 10Rectangular partitions are not allowed to be further subdivided. The coding tree depth can be further defined to indicate the depth of splitting from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128×128 block) can be set to 0, and after the root block is further split once according to Figure 10 the pattern, the coding tree depth is incremented by 1. In some embodiments, only all square partitions in 1010 may be allowed to be recursively partitioned to the next level of the partition tree according to Figure 10 the pattern. In other words, for square partitions in the T-shaped patterns 1002, 1004, 1006, and 1008, recursive partitioning may not be allowed. If needed, the final set of coded blocks is generated according to the recursive partitioning process of Figure 10 Such a scheme can be applied to one or more color channels. In some embodiments, more flexibility can be added when using partitions below the 8x8 level. For example, 2x2 chrominance inter prediction can be used in some cases.
[0118] In some other example embodiments for codec block partitioning, a quadtree structure can be used to split a basic block or an intermediate block into quadtree partitions. Such quadtree splitting can be applied hierarchically and recursively to any square partition. Whether the basic block or intermediate block / partition is further quadtree split can be adapted to various local characteristics of the basic block or intermediate block / partition. The quadtree partitioning at the picture boundary can be further adjusted. For example, an implicit quadtree split can be performed at the picture boundary so that the blocks will remain quadtree split until the size fits the picture boundary.
[0119] In some other example embodiments, hierarchical binary partitioning can be used starting from the basic block. For such a scheme, the basic block or an intermediate-level block can be partitioned into two partitions. The binary partitioning can be horizontal or vertical. For example, a horizontal binary partitioning can split the basic block or intermediate block into equal right and left partitions. Similarly, a vertical binary partitioning can split the basic block or intermediate block into equal upper and lower partitions. Such binary splitting can be hierarchical and recursive. A decision can be made at each of the basic blocks or intermediate blocks as to whether the binary partitioning scheme should continue, and if the scheme does continue further, a decision can be made as to whether horizontal or vertical binary partitioning should be used. In some embodiments, further partitioning is stopped at a predefined minimum partition size (in one or both dimensions). Optionally, once a predefined partition level or depth from the basic block is reached, further partitioning can be stopped. In some embodiments, the aspect ratio of the partitions can be restricted. For example, the aspect ratio of the partitions can be not less than 1:4 (or greater than 4:1). Thus, a vertical bar partition with a vertical-to-horizontal aspect ratio of 4:1 can only be further vertically binary partitioned into upper and lower partitions each having a vertical-to-horizontal aspect ratio of 2:1.
[0120] In still other examples, a ternary partitioning scheme can be used to partition a basic block or any intermediate block, such as Figure 13 shown. The ternary pattern can be implemented vertically as shown in 1302 of Figure 13 or horizontally as shown in 1304 of Figure 13 . Although the example vertical or horizontal split ratios in Figure 13 are shown as 1:2:1, other ratios can also be predefined. In some embodiments, two or more different ratios can be predefined. Such ternary partitioning schemes can be used to complement quadtree or binary partitioning structures because such ternary tree partitioning can capture an object located at the center of a block in a continuous partition, while quadtree and binary trees always split along the center of the block and thus split the object into separate partitions. In some embodiments, the width and height of the partitions of an example ternary tree are always powers of 2 to avoid additional transformations.
[0121] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree and binary partitioning schemes can be combined to partition a basic block into a quadtree - binary tree (QTBT) structure. In such a scheme, the basic block or intermediate block / partition can be either quadtree split or binary split, subject to a set of predefined conditions if specified. A specific example is shown in Figure 14 . In the example of Figure 14 , the basic block is first quadtree split into four partitions as shown in 1402, 1404, 1406, and 1408. Thereafter, each of the resulting partitions is either quadtree partitioned into four further partitions (such as 1408), or binary split into two further partitions (horizontally or vertically, such as 1402 or 1406, both of which are symmetric) at the next level, or not split (such as 1404). For square partitions, binary or quadtree partitioning can be allowed recursively, as shown in the overall example partitioning pattern of 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree splits and dashed lines represent binary splits. Flags can be used for each binary split node (non - leaf binary partition) to indicate whether the binary split is horizontal or vertical. For example, as shown in 1420 and consistent with the partitioning structure of 1410, the flag "0" can represent a horizontal binary split, and the flag "1" can represent a vertical binary split. For quadtree - split partitions, there is no need to indicate the split type because quadtree splitting always splits the block or partition horizontally and vertically to produce 4 sub - blocks / partitions of equal size. In some embodiments, the flag "1" can represent a horizontal binary split, and the flag "0" can represent a vertical binary split.
[0122] In some example embodiments of QTBT, the quadtree and binary tree partitioning rules can be represented by the following predefined parameters and corresponding functions associated therewith:
[0123] — CTU size: The size of the root node of the quadtree (the size of the basic block)
[0124] — MinQTSize: The minimum allowable size of the quadtree leaf node
[0125] — MaxBTSize: The maximum allowable size of the binary tree root node
[0126] — MaxBTDepth: The maximum allowable depth of the binary tree
[0127] — MinBTSize: The minimum allowable size of the binary tree leaf node
[0128] In some example embodiments of the QTBT partitioning structure, the CTU size can be set to 128×128 luma samples having two corresponding 64×64 chroma sample blocks (when considering and using example chroma subsampling), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. The quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have a size ranging from their minimum allowable size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If the node is 128×128, it will not be first split by the binary tree because the size exceeds MaxBTSize (i.e., 64×64). Otherwise, the nodes that do not exceed MaxBTSize can be partitioned by the binary tree. In Figure 14 the example, the basic block is 128×128. According to the predefined ruleset, the basic block can only be split by the quadtree. The partitioning depth of the basic block is 0. Each of the resulting four partitions is 64×64, which does not exceed MaxBTSize and can be further split by the quadtree or the binary tree at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be considered not to be performed. When the width of the binary tree node is equal to MinBTSize (i.e., 4), further horizontal splitting can be considered not to be performed. Similarly, when the height of the binary tree node is equal to MinBTSize, further vertical splitting can be considered not to be performed.
[0129] In some example embodiments, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure or respective QTBT structures for luminance and chrominance. For example, for P and B slices, the luminance and chrominance CTBs in a CTU can share the same QTBT structure. However, for I slices, the luminance CTB can be partitioned into CUs by a QTBT structure, and the chrominance CTB can be partitioned into chrominance CUs by another QTBT structure. This means that CUs can be used to represent different color channels in an I slice. For example, an I slice can consist of coding blocks of the luminance component or coding blocks of two chrominance components, and a CU in a P or B slice can consist of coding blocks of all three color components.
[0130] In some other embodiments, the QTBT scheme can be supplemented with the above ternary scheme. Such embodiments can be referred to as multi-type tree (MTT) structures. For example, in addition to the binary splitting of nodes, one of the ternary partitioning patterns can be selected. Figure 13 In some embodiments, only square nodes can be ternary split. An additional flag can be used to indicate whether the ternary partitioning is horizontal or vertical.
[0131] The design of two-level or multi-level trees, such as the QTBT embodiment and the QTBT embodiment supplemented by ternary splitting, can be mainly for reducing complexity. Theoretically, the complexity of traversing the tree is T D , where T represents the number of splitting types, and D is the depth of the tree. A trade-off can be made by using multiple types (T) to reduce the depth (D) simultaneously.
[0132] In some embodiments, a CU can be further partitioned. For example, for the purpose of intra-frame or inter-frame prediction during the encoding and decoding processes, a CU can be further partitioned into multiple prediction blocks (PBs). In other words, a CU can be further divided into different sub-partitions where individual prediction decisions / configurations can be made. In parallel, for the purpose of depicting the level of performing the transform or inverse transform of video data, a CU can be further partitioned into multiple transform blocks (TBs). The partitioning schemes of a CU into PBs and TBs can be the same or can be different. For example, its own process can be used to perform each partitioning scheme based on various characteristics of the video data, such as. In some example embodiments, the PB and TB partitioning schemes can be independent. In some other example embodiments, the PB and TB partitioning schemes and boundaries can be related. In some embodiments, for example, TBs can be partitioned after PB partitioning. In particular, each PB is further partitioned into one or more TBs after the partitioning of the coding block is determined. For example, in some embodiments, a PB can be split into one, two, four, or other numbers of TBs.
[0133] In some embodiments, to partition a basic block into coding blocks and further into prediction blocks and / or transform blocks, the luminance channel and the chrominance channel may be processed differently. For example, in some embodiments, coding blocks of the luminance channel may be allowed to be partitioned into prediction blocks and / or transform blocks, while coding blocks of the chrominance channel may not be allowed to be partitioned into prediction blocks and / or transform blocks. In such embodiments, the transformation and / or prediction of luminance blocks may thus be performed only at the coding block level. For another example, the minimum transform block size of the luminance channel and the chrominance channel may be different. For example, coding blocks of the luminance channel may be allowed to be partitioned into smaller transform blocks and / or prediction blocks than those of the chrominance channel. For yet another example, the maximum depth of partitioning a coding block into transform blocks and / or prediction blocks may be different between the luminance channel and the chrominance channel. For example, coding blocks of the luminance channel may be allowed to be partitioned into deeper transform blocks and / or prediction blocks than those of the chrominance channel. For a specific example, luminance coding blocks may be partitioned into transform blocks of multiple sizes, which may be represented by recursive partitioning down to up to 2 levels, and transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4 and transform block sizes from 4×4 to 64×64 may be allowed. However, for chrominance blocks, only the maximum possible transform blocks specified for luminance blocks may be allowed.
[0134] In some example embodiments for partitioning coding blocks into PBs, the depth, shape, and / or other characteristics of partitioning the PBs may depend on whether the PB is intra-coded or inter-coded.
[0135] In various example scenarios, partitioning a coding block (or prediction block) into transform blocks may be implemented, which includes but is not limited to quadtree splitting and predefined pattern splitting, either recursively or non-recursively, and additional consideration of transform blocks at the boundaries of the coding block or prediction block. Generally, the resulting transform blocks may be at different splitting levels, may not have the same size, and may not need to be square-shaped (e.g., they may be rectangles with some allowed sizes and aspect ratios). Further examples will be described in more detail below in conjunction with Figure 15 、 16 and 17.
[0136] However, in some other embodiments, the CBs obtained via any of the above partitioning schemes can be used as the basic or minimum decoding blocks for prediction and / or transformation. In other words, no further splitting is performed for the purpose of performing inter-frame prediction / intra-frame prediction and / or for the purpose of transformation. For example, the CBs obtained from the above QTBT scheme can be directly used as the units for performing prediction. Specifically, such QTBT structures remove the concept of multiple partitioning types, i.e., it removes the separation of CUs, PUs, and TUs, and supports more flexibility in the CU / CB partitioning shapes as described above. In such QTBT block structures, the CU / CB can have a square or rectangular shape. The leaf nodes of such QTBTs are used as the units for prediction and transformation processing without any further partitioning. This means that the CUs, PUs, and TUs have the same block size in such example QTBT codec block structures.
[0137] The above various CB partitioning schemes and the further partitioning of CBs into PBs and / or TBs (including no PB / TB partitioning) can be combined in any way. The following specific embodiments are provided as non-limiting examples.
[0138] Specific example embodiments of the coding block and transform block partitioning are described below. In such example embodiments, the recursive quadtree splitting or predefined splitting patterns (such as Figure 9 and 10 those in
[0139] can be used to split the basic block into coding blocks. At each level, whether the further quadtree splitting of a particular partition should continue can be determined by the local video data characteristics. The resulting CBs can be at various quadtree splitting levels of various sizes. A decision can be made at the CB level (or CU, for all three color channels) as to whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode the picture region. Each CB can be further split into one, two, four, or other numbers of PBs according to a predefined PB splitting type. Within a PB, the same prediction process can be applied, and the relevant information can be sent to the decoder on the basis of the PB. After obtaining the residual block by applying the prediction process based on the PB splitting type, the CB can be partitioned into TBs according to another quadtree structure similar to the coding tree of the CB. In this specific embodiment, the CB or TB can be, but is not limited to, square. In addition, in this specific example, the PB can be square or rectangular for inter-frame prediction and can be only square for intra-frame prediction. The coding block can be split into, for example, four square-shaped TBs. Each TB can be further recursively split (using quadtree splitting) into smaller TBs, called Residual Quadtree (RQT).
[0139] Another example embodiment for partitioning a basic block into CBs, PBs, and / or TBs is described below. For example, a quadtree of nested multi-type trees can be used, which has a binary and ternary split segmentation structure (e.g., QTBT as described above or QTBT with ternary split), rather than using multiple partition unit types such as Figure 9 or those shown in 10. The separation of the CB, PB, and TB concepts (i.e., partitioning the CB into PBs and / or TBs, and partitioning the PB into TBs) can be waived, unless a CB of a size too large for the maximum transform length is needed, in which case such a CB may need to be further segmented. This example partitioning scheme can be designed to support more flexibility in the CB partitioning shape, such that both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, a CU can have a square or rectangular shape. Specifically, a coding tree block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a nested multi-type tree structure. Figure 11 An example of a nested multi-type tree structure using binary or ternary split is shown in Figure 11 . Specifically, Figure 11 the example multi-type tree structure includes four split types, which are referred to as vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical ternary split (SPLIT_TT_VER) (1106), and horizontal ternary split (SPLIT_TT_HOR) (1108). Then, a CB corresponds to a leaf of the multi-type tree. In this example embodiment, this segmentation is used for prediction and transformation processing without any further partitioning, unless the CB is too large for the maximum transform length. This means that, in most cases, CBs, PBs, and TBs have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CB. In some embodiments, in addition to binary or ternary split, a nested pattern such as
[0140] Figure 12 An example of a quadtree with a nested multi-type tree coding block structure for block partitioning (including quadtree, binary, and ternary split options) for one basic block is shown in Figure 12 . More specifically, Figure 11 shows that the basic block 1200 is split into four square partitions 1202, 1204, 1206, and 1208 by a quadtree. It is decided to further use Figure 12In the example, partition 1204 is not further divided. Partitions 1202 and 1208 each undergo another quadtree division. For partition 1202, the upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree division respectively undergo the third-level division of the quadtree, Figure 11 the horizontal binary division 1104, no division, and Figure 11 the horizontal ternary division 1108. Partition 1208 undergoes another quadtree division, and the upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree division respectively undergo Figure 11 the third-level division of the vertical ternary division 1106, no division, no division, and Figure 11 the horizontal binary division 1104. Two of the sub-partitions in the third-level upper-left partition of 1208 are further divided respectively according to Figure 11 the horizontal binary division 1104 and the horizontal ternary division 1108. Partition 1206 is divided into two partitions according to the second-level division pattern of the vertical binary division 1102 in accordance with Figure 11 , and the two partitions are further divided at the third level according to Figure 11 the horizontal ternary division 1108 and the vertical binary division 1102. According to Figure 11 the horizontal binary division 1104, the fourth-level division is further applied to one of them.
[0141] For the above specific example, the maximum luminance transform size can be 64×64, and the maximum supported chrominance transform size can be different from the luminance at, for example, 32×32. Although the example CB above Figure 12 is generally not further divided into smaller PBs and / or TBs, when the width or height of a luminance coding block or a chrominance coding block is greater than the maximum transform width or height, the luminance coding block or the chrominance coding block can be automatically divided along the horizontal and / or vertical directions to meet the transform size limit in that direction.
[0142] In the specific example for partitioning a basic block into the above CBs, as described above, the coding tree scheme can support the ability for luminance and chrominance to have separate block tree structures. For example, for P and B slices, the luminance and chrominance CTBs in a CTU can share the same coding tree structure. For example, for I slices, luminance and chrominance can have separate coded block tree structures. When applying separate block tree structures, the luminance CTB is partitioned into CBs through one coding tree structure, and the chrominance CTB is partitioned into chrominance CBs through another coding tree structure. This means that a CU in an I slice can consist of coded blocks of the luminance component or coded blocks of two chrominance components, and a CU in a P or B slice always consists of coded blocks of all three color components, unless the video is monochrome.
[0143] When a coding block is further partitioned into multiple transform blocks, the transform blocks therein are in order in the bitstream, following various orders or scanning patterns. Example implementation schemes for partitioning a coding block or a prediction block into transform blocks and the coding order of the transform blocks are further described in detail below. In some example implementation schemes, as described above, the transform partitioning can support multiple shapes. For example, transform blocks of 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, where the transform block size ranges from, for example, 4×4 to 64×64. In some implementation schemes, if the coding block is less than or equal to 64×64, the transform block partitioning can be applied only to the luminance component, such that for the chrominance blocks, the transform block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, then the luminance and chrominance coding blocks can be implicitly divided into multiple times of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.
[0144] In some example implementation schemes of transform block partitioning, for both intra-coded blocks and inter-coded blocks, the coding block can be further partitioned into multiple transform blocks with a partition depth of up to a predetermined number of levels (e.g., 2 levels). The transform block partition depth and size can be related. For some example embodiments, an example mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below.
[0145] Table 1: Transform Partition Size Settings
[0146]
[0147] Based on the example mapping in Table 1, for a 1:1 square block, the next-level transform split can create four 1:1 square sub-transform blocks. The transform partitioning can stop at 4×4, for example. In this way, the transform size at the current depth of 4×4 corresponds to the same size of 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next-level transform split will create two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next-level transform split will create two 1:2 / 2:1 sub-transform blocks.
[0148] In some example implementation schemes, for the luminance component of an intra-coded block, additional restrictions can be applied to the transform block partitioning. For example, for each level of transform partitioning, all sub-transform blocks can be restricted to have the same size. For example, for a 32×16 coding block, the level 1 transform split creates two 16×16 sub-transform blocks, and the level 2 transform split creates eight 8×8 sub-transform blocks. In other words, the second-level split must be applied to all first-level sub-blocks to keep the transform unit sizes equal. In Figure 15An example of transform block partitioning for intra-coded square blocks shown in Table 1 is illustrated, along with the coding order indicated by the arrow diagram. Specifically, 1502 shows the square coding block. In 1504, a first-level partitioning into 4 equally sized transform blocks is shown according to Table 1, where the coding order is indicated by the arrow. In 1506, a second-level partitioning of all the first-level equally sized blocks into 16 equally sized transform blocks is shown according to Table 1, where the coding order is indicated by the arrow.
[0149] In some example embodiments, for the luminance component of an inter-coded block, the above restrictions on intra-coding may not be applied. For example, after the first-level transform partitioning, any one of the sub-transform blocks may be further independently partitioned one more level. Thus, the resulting transform blocks may or may not have the same size. In Figure 16 an example of the transform blocks into which an inter-coded block is partitioned is shown, with the transform blocks having a coding order. In Figure 16 the example of, according to Table 1, the inter-coded block 1602 is partitioned into transform blocks at the second level. At the first level, the inter-coded block is partitioned into four equally sized transform blocks. Then, as shown in 1604, only one (not all) of the four transform blocks is further partitioned into four sub-transform blocks, resulting in a total of 7 transform blocks having two different sizes. The example coding order of these 7 transform blocks is indicated by the Figure 16 arrows in 1604 of
[0150] In some example embodiments, for one or more chrominance components, some additional restrictions on the transform blocks may be applied. For example, for one or more chrominance components, the transform block size may be as large as the coding block size, but not less than a predefined size, such as 8×8.
[0151] In some other example embodiments, for coding blocks with a width (W) or height (H) greater than 64, both the luminance and chrominance coding blocks may be implicitly partitioned into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively. Here, in the present disclosure, "min(a, b)" may return the smaller value between a and b.
[0152] Figure 17 A further alternative example scheme for partitioning a coding block or a prediction block into transform blocks is illustrated. As Figure 17 shown, instead of using recursive transform partitioning, a predefined set of partitioning types may be applied to the coding block according to the transform type of the coding block. In the specific example shown in Figure 17 one of 6 example partitioning types may be applied to partition the coding block into various numbers of transform blocks. Such a scheme for generating transform block partitioning may be applied to a coding block or a prediction block.
[0153] More specifically, Figure 17 the partitioning scheme of provides up to six example partitioning types for any given transform type (where the transform type refers to, for example, the type of preliminary transform, such as ADST and others). In this scheme, for example, a transform partitioning type can be assigned to each coding block or prediction block based on the rate-distortion cost. In an example, the transform partitioning type assigned to a coding block or prediction block can be determined based on the transform type of the coding block or prediction block. A particular transform partitioning type can correspond to a transform block split size and pattern, as shown by the six transform partitioning types illustrated in Figure 17 The correspondence between various transform types and various transform partitioning types can be predefined. The following shows a corresponding example, where the capitalized labels indicate the transform partitioning types that can be assigned to a coding block or prediction block based on the rate-distortion cost:
[0154] · PARTITION_NONE: Assigns a transform size equal to the block size.
[0155] · PARTITION_SPLIT: The assigned transform size is 1 / 2 of the width of the block size and 1 / 2 of the height of the block size.
[0156] · PARTITION_HORZ: Assigns a transform size having the same width as the block size and 1 / 2 of the height of the block size.
[0157] · PARTITION_VERT: Assigns a transform size having 1 / 2 of the width of the block size and the same height as the block size.
[0158] · PARTITION_HORZ4: Assigns a transform size having the same width as the block size and 1 / 4 of the height of the block size.
[0159] · PARTITION_VERT4: Assigns a transform size having 1 / 4 of the width of the block size and the same height as the block size.
[0160] In the above example, all the transform partitioning types shown as Figure 17 include a unified transform size for the partitioned transform block. This is merely an example and not a limitation. In some other embodiments, a mixed transform block size can be used for the partitioned transform block of a specific partitioning type (or mode).
[0161] Then, a PB (or CB, which is also referred to as a PB when not further partitioned into prediction blocks) obtained from any of the above partitioning schemes can become the respective blocks for encoding and decoding through intra-frame or inter-frame prediction. For inter-frame prediction of the current PB, the residual between the current block and the prediction block can be generated, encoded, and included in the encoded bitstream.
[0162] Inter-frame prediction can be implemented, for example, in a single-reference mode or a composite-reference mode. In some embodiments, a skip flag may first be included in the bitstream of a current block (or at a higher level) to indicate whether the current block is inter-frame coded and will not be skipped. If the current block is inter-frame coded, another flag may be further included in the bitstream to signal whether the single-reference mode or the composite-reference mode is used for the current block. For the single-reference mode, one reference block may be used to generate a predicted block of the current block. For the composite-reference mode, two or more reference blocks may be used to generate a predicted block, for example, by weighted averaging. The composite-reference mode may be referred to as a multiple-reference mode, a two-reference mode, or a multi-reference mode. One or more reference frame indices may be used, and additionally, one or more corresponding motion vectors may be used to identify one or more reference blocks, where the one or more corresponding motion vectors indicate one or more shifts in position (e.g., in horizontal and vertical pixels) between the one or more reference blocks and the current block. For example, an inter-frame predicted block of the current block may be generated from a single reference block identified by a motion vector in a reference frame as a predicted block in the single-reference mode, while for the composite-reference mode, the predicted block may be generated by weighted averaging of two reference blocks in two reference frames indicated by two motion vectors. The one or more motion vectors may be encoded and included in the bitstream in various ways.
[0163] In some embodiments, an encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be held in the DPB waiting for display (in a decoding system), and some images / pictures in the DPB may be used as reference frames to achieve inter-frame prediction. In some embodiments, the reference frames in the DPB may be marked as short-term references or long-term references for the current image being encoded or decoded. For example, a short-term reference frame may include a frame used for inter-frame prediction of blocks in the current frame, or a frame including blocks for inter-frame prediction of a predefined number (e.g., 2) of subsequent video frames closest to the current frame in the decoding order. A long-term reference frame may include frames in the DPB that may be used to predict image blocks in frames that are more than a predefined number of frames from the current frame in the decoding order. Information about such tags for short-term and long-term reference frames may be referred to as a reference picture set (RPS), and may be added to the header of each frame in the encoded bitstream. Each frame in the encoded video bitstream may be identified by a picture order counter (POC), and each frame is numbered absolutely according to the playback sequence, or numbered relative to a group of pictures starting from, for example, an I-frame.
[0164] In some example embodiments, one or more reference picture lists may be formed based on information in the RPS, the list containing the identities of short-term and long-term reference frames for inter-frame prediction. For example, a single picture reference list may be formed for uni-directional inter-frame prediction, denoted as the L0 reference (or reference list 0), while two picture reference lists may be formed for bi-directional inter-frame prediction, denoted as L0 (or reference list 0) and L1 (or reference list 1), for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be sorted in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. When multiple references used to generate a prediction block by weighted averaging in a combined prediction mode are on the same side of the block to be predicted, uni-directional inter-frame prediction may be in a single-reference mode or a combined-reference mode. Bi-directional inter-frame prediction can only be in a combined mode because bi-directional inter-frame prediction involves at least two reference blocks.
[0165] In some embodiments, a merge mode (MM) for inter-frame prediction may be implemented. Generally, for the merge mode, one or more of the motion vectors in the single-reference prediction of the current PB or the motion vectors in the combined-reference prediction may be derived from one or more other motion vectors, rather than being independently calculated and signaled. For example, in an encoding system, one or more current motion vectors of the current PB may be reduced to one or more differences between one or more current motion vectors and one or more previously encoded motion vectors (referred to as reference motion vectors). One or more differences of one or more motion vectors (rather than all of the one or more current motion vectors) may be encoded and included in the bitstream, and may be linked to one or more reference motion vectors. Correspondingly, in a decoding system, one or more motion vectors corresponding to the current PB may be derived based on the decoded motion vector differences and one or more decoded reference motion vectors linked thereto. As a specific form of general merge mode (MM) inter-frame prediction, such inter-frame prediction based on motion vector differences may be referred to as merge mode with motion vector difference (MMVD). Thus, either general MM or specific MMVD may be implemented to exploit the correlation between motion vectors associated with different PBs to improve encoding and decoding efficiency. For example, adjacent PBs may have similar motion vectors. For another example, for blocks similarly positioned / placed in space, the motion vectors may be temporally correlated (between frames).
[0166] In some example embodiments, an MM flag may be included in a bitstream during an encoding process to indicate whether a current PB is in merge mode. Additionally, or alternatively, an MMVD flag may be included in and signaled in the bitstream during the encoding process to indicate whether the current PB is in MMVD mode. The MM and / or MMVD flag or indicator may be provided at levels such as PB level, CB level, CU level, CTB level, CTU level, slice level, picture level, etc. For a particular example, for a current CU, both an MM flag and an MMVD flag may be included, and the MMVD flag may be signaled immediately after the skip flag and the MM flag to specify whether the MMVD mode is used for the current CU.
[0167] In some example embodiments of MMVD, a merge candidate list for motion vector prediction may be formed for a block being predicted. The merge candidate list may contain a predetermined number (e.g., 2) of MV predictor candidate blocks, and the motion vectors of the MV predictor candidate blocks may be used to predict the current motion vector. The MVD candidate blocks may include blocks selected from adjacent blocks and / or temporal blocks in the same frame (e.g., blocks at the same location in a previous or subsequent frame of the current frame). These options represent blocks that may have a motion vector similar to or the same as the current block at a spatial or temporal position relative to the current block. The size of the MV predictor candidate list may be predetermined. For example, the list may contain two candidates. In order to be on the merge candidate list, a candidate block, for example, needs to have the same one (or more) reference frames as the current block, must exist (e.g., when the current block is close to the edge of the frame, a boundary check needs to be performed), and must have been encoded during the encoding process and / or decoded during the decoding process. In some embodiments, if spatial adjacent blocks are available and meet the above conditions, the merge candidate list may first be filled with spatial adjacent blocks (scanned in a specific predefined order), and then if there is still space in the list, temporal blocks may be filled. For example, adjacent candidate blocks may be selected from the left block and the top block of the current block. The merged MV predictor candidate list may be signaled in the bitstream.
[0168] In some embodiments, the actual merge candidate being used as a reference motion vector for predicting the motion vector of the current block may be signaled. In the case where the merge candidate list contains two candidates, a one-bit flag called the merge candidate flag may be used to indicate the selection of the reference merge candidate. For a current block being predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor may be associated with a reference motion vector from the merge candidate list.
[0169] In some example embodiments of MMVD, after selecting a merge candidate and using it as a base motion vector predictor for a motion vector to be predicted, a motion vector difference (MVD or delta MV, representing the difference between the motion vector to be predicted and a reference candidate motion vector) can be calculated in an encoding system. Such MVDs can include information representing the magnitude of the MV difference and the direction of the MV difference, which can be signaled in a bitstream. The motion difference magnitude and the motion difference direction can be signaled in various ways.
[0170] In some example embodiments of MMVD, a distance index can be used to specify the magnitude information of the motion vector difference and to indicate one of a set of predefined offsets that represent a predefined motion vector difference from a starting point (the reference motion vector). Then, the MV offset according to the signaled index can be added to either the horizontal or vertical component of the starting (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset is determined by example direction information of the MVD. An example predefined relationship between the distance index and the predefined offsets is specified in Table 2.
[0171] Table 2 - Example relationship between distance index and predefined MV offsets
[0172]
[0173] In some example embodiments of MMVD, the direction index can be further signaled and used to represent the direction of the MVD relative to a reference motion vector. In some embodiments, the direction can be restricted to one of a horizontal direction and a vertical direction. An exemplary 2-bit direction index is shown in Table 3. In the example of Table 3, the interpretation of the MVD can vary according to the information of the start / reference MV. For example, when the start / reference MV corresponds to a uni-directional prediction block or corresponds to a bi-directional prediction block (where the two reference frame lists point to the same side of the current picture) (i.e., the POCs of both reference pictures are greater than the POC of the current picture, or both are less than the POC of the current picture), the signs in Table 3 can specify the sign (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a bi-directional prediction block (where the two reference pictures are at different sides of the current picture) (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the signs in Table 3 can specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 can have an opposite value (opposite sign of the offset). Conversely, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the signs in Table 3 can specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset of the reference MV associated with picture reference list 0 has an opposite value.
[0174] Table 3 - Example embodiments of the signs of MV offsets specified by the direction index
[0175] Direction IDX 00 01 10 11 x-axis (horizontal) + - Not applicable Not applicable y-axis (vertical) Not applicable Not applicable + -
[0176] In some example embodiments, the MVD can be scaled according to the difference in POC in each direction. If the differences in POC in the two lists are the same, no scaling is required. Otherwise, if the difference in POC in reference list 0 is greater than the difference in POC in reference list 1, scale the MVD of reference list 1. If the POC difference in reference list 1 is greater than that in list 0, the MVD of list 0 can be scaled in the same way. If the start MV is uni-directionally predicted, add the MVD to the available or reference MV.
[0177] In some example embodiments of MVD encoding / decoding and signaling for bidirectional composite prediction, in addition to or as an alternative to separately encoding and signaling two MVDs, symmetric MVD encoding / decoding may be implemented such that only one MVD needs to be signaled and the other MVD may be derived from the signaled MVD. In such embodiments, motion information is signaled that includes reference picture indices for both list-0 and list-1. However, only the MVD associated with, for example, reference list-0 is signaled, while the MVD associated with reference list-1 is not signaled but derived. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" may be included in the bitstream to indicate whether reference list-1 is not signaled in the bitstream. If this flag is 1, indicating that reference list-1 is equal to zero (and thus not signaled), the bidirectional prediction flag (referred to as "BiDirPredFlag") may be set to 0, meaning there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, if the closest reference picture in list-0 and the closest reference picture in list-1 form a forward-backward pair of reference pictures or a backward-forward pair of reference pictures, the BiDirPredFlag may be set to 1 and both the list-0 and list-1 reference pictures are short-term reference pictures. Otherwise, the BiDirPredFlag is set to 0. A BiDirPredFlag of 1 indicates that a symmetric mode flag is additionally signaled in the bitstream. When the BiDirPredFlag is 1, the decoder may extract the symmetric mode flag from the bitstream. For example, the symmetric mode flag may be signaled (if needed) at the CU level, and it indicates whether the symmetric MVD encoding / decoding mode is being used for the corresponding CU. When the symmetric mode flag is 1, it indicates that the symmetric MVD encoding / decoding mode is used, and the reference picture indices for both list-0 and list-1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") and the MVD associated with list-0 (referred to as "MVD0") are signaled, and the other motion vector difference "MVD1" will be derived rather than signaled. For example, MVD1 may be derived as -MVD0. Thus, only one MVD is signaled in the example symmetric MVD mode. In some other example embodiments of MV prediction, a coordinated scheme may be used to implement the general merge mode, MMVD, and some other types of MV prediction for both single-reference and composite-reference mode MV prediction. Various syntax elements may be used to signal the way to predict the MV of the current block.
[0178] For example, for the single-reference mode, the following MV prediction modes may be signaled:
[0179] NEARMV - directly uses one of the motion vector predictors (MVPs) in the list without using any MVD, and this MVP is indicated by the DRL (Dynamic Reference List) index.
[0180] NEWMV - uses one of the motion vector predictors (MVPs) in the list as a reference and applies an increment to the MVP (e.g., uses MVD), and this MVP is signaled by the DRL index.
[0181] GLOBALMV - uses a motion vector based on frame - level global motion parameters.
[0182] Similarly, for the composite reference inter - prediction mode that uses two reference frames corresponding to two MV to be predicted, the following MV prediction modes can be signaled:
[0183] NEAR_NEARMV - for each of the two MV to be predicted, uses one of the motion vector predictors (MVPs) in the list without using MVD, and this MVP is signaled by the DRL index.
[0184] NEAR_NEWMV - to predict the first of the two motion vectors, uses one of the motion vector predictors (MVPs) in the list as a reference MV without MVD, and this MVP is signaled by the DRL index; to predict the second of the two motion vectors, uses one of the motion vector predictors (MVPs) in the list as a reference MV and combines it with an additionally signaled incremental MV (MVD), and this MVP is signaled by the DRL index.
[0185] NEW_NEARMV - to predict the second of the two motion vectors, uses one of the motion vector predictors (MVPs) in the list as a reference MV without MVD, and this MVP is signaled by the DRL index; to predict the first of the two motion vectors, uses one of the motion vector predictors (MVPs) in the list as a reference MV and combines it with an additionally signaled incremental MV (MVD), and this MVP is signaled by the DRL index.
[0186] NEW_NEWMV - uses one of the motion vector predictors (MVPs) in the list as a reference MV and combines it with an additionally signaled incremental MV to predict each of the two MV, where this MVP is signaled by the DRL index.
[0187] GLOBAL_GLOBALMV - uses the MV from each reference based on frame - level global motion parameters.
[0188] Accordingly, the above term "NEAR" refers to MV prediction using a reference MV (without MVD) as a general merge mode, while the term "NEW" refers to MV prediction involving using a reference MV and offsetting it with a signaled MVD as in the MMVD mode. For composite inter prediction, both the above reference base motion vector and motion vector difference can generally be different or independent between two references (even if they can be correlated and such correlation can be exploited to reduce the amount of information needed to signal the two motion vector differences). In such cases, joint signaling of the two MVDs can be implemented and indicated in the bitstream.
[0189] The above dynamic reference list (DRL) can be used to hold a set of indexed motion vectors that are dynamically maintained and considered as candidate motion vector predictors.
[0190] In some example embodiments, for composite prediction, an optical flow-based approach can be used to correct the motion vector (MV) by sub-block. Specifically, the optical flow equation can be applied to formulate a least squares problem, from which the fine motion can be derived from the gradients of the composite inter prediction samples. Using these fine motions, the MV of each sub-block can be corrected within the prediction block, which can enhance the inter prediction quality. Some codec features can be an extension of the concept of bidirectional optical flow (BDOF) as it supports MV correction when the two reference blocks have an arbitrary temporal distance to the current block.
[0191] In some embodiments, the following four additional composite inter modes can be added: NEAR_NEARMV_OPTFLOW, NEAR_NEWMV_OPTFLOW, NEW_NEARMV_OPTFLOW, and / or NEW_NEWMV_OPTFLOW.
[0192] These modes can be referred to as optical flow modes, and the reference MV type can be defined as a conventional composite mode (e.g., NEAR_NEWMV_OPTFLOW has the same reference MV type as NEAR_NEWMV). Composite prediction can be performed based on the sub-block corrected MV rather than the original MV.
[0193] The various embodiments and / or implementations described in this disclosure can be used alone or in any order combination. Further, a part, all, or any part or all combination of these embodiments and / or implementations can be embodied as part of an encoder and / or decoder and can be implemented in hardware and / or software. For example, they can be hard-coded in a dedicated processing circuit (e.g., one or more integrated circuits). In another example, they can be implemented by one or more processors executing a program stored in a non-volatile computer-readable medium.
[0194] There may be some concerns / issues associated with some implementations of the signaling method for motion vector differences. For example, how to signal one or more delta MVs in the NEW_NEARMV mode, NEAR_NEWMV mode, or NEW_NEWMV mode. One concern / issue is that the correlation of the motion vector differences in the two reference lists is not utilized, thus reducing the encoding / decoding efficiency and performance.
[0195] The present disclosure describes various embodiments for signaling the motion vector difference (MVD or delta MV) for inter-frame prediction mode encoding and / or decoding, which address at least one of the above concerns / issues and implement an effective software / hardware implementation for improved inter-frame prediction mode encoding / decoding.
[0196] In various embodiments, referring to Figure 18 , a method 1800 for video decoding is provided, which is executed in a device including a memory storing instructions and a processor communicating with the memory. Method 1800 may include some or all of the following steps: step 1810, receiving an encoded video bitstream; step 1820, extracting the inter-frame prediction mode and the joint delta motion vector MV of a current block in a current frame from the encoded video bitstream; step 1830, extracting a first flag from the encoded video bitstream indicating whether a first delta MV of a first reference frame and a second delta MV of a second reference frame are jointly signaled; step 1840, in response to the first flag indicating that the first delta MV and the second delta MV are jointly signaled, deriving the first delta MV and the second delta MV based on the joint delta MV; and / or step 1850, decoding the current block in the current frame based on the first delta MV and the second delta MV.
[0197] In some implementations, the joint delta MV may be represented as joint_delta_mv, which may be an element indicating the joint delta MV. In some implementations, the first flag indicating whether a first delta MV of a first reference frame and a second delta MV of a second reference frame are jointly signaled may be represented as joint_mvd_flag. In some implementations, the first reference frame may be a frame in a reference list (reference list 0); and / or the second reference frame may be a frame in another reference list (reference list 1).
[0198] In some implementations, step 1830 may include extracting a first flag (joint_mvd_flag) from the encoded video bitstream, and the first flag (joint_mvd_flag) indicates whether a first delta MV of a first reference frame in reference list 0 and a second delta MV of a second reference frame in reference list 1 are jointly signaled.
[0199] In various embodiments of the present disclosure, the size of a block (such as, but not limited to, a coding block, a prediction block, or a transform block) may refer to the width or height of the block. The width or height of the block may be an integer in pixels. In various embodiments of the present disclosure, the size of the block may refer to the area size of the block. The area size of the block may be an integer calculated by multiplying the width of the block by the height of the block in pixels. In some different embodiments of the present disclosure, the size of the block may refer to the maximum value of the width or height of the block, the minimum value of the width and height of the block, or the aspect ratio of the block. The aspect ratio of the block may be calculated as the width of the block divided by the height of the block, or may be calculated as the height divided by the width of the block.
[0200] Here, in some embodiments of the present disclosure, the "first" reference frame may refer not only to "one" reference frame, but also to the "first" reference frame among multiple reference frames (e.g., having the smallest index, or appearing earliest in the sequence), and the "second" reference frame may refer not only to "another" reference frame, but also to the "second" reference frame among multiple reference frames (e.g., having the second smallest index, or appearing second earliest in the sequence).
[0201] Here, in various embodiments of the present disclosure, "signaling XYZ" may refer to encoding XYZ into the encoded bitstream during the encoding process; and / or, after sending the encoded bitstream from one device to another device, "signaling XYZ" may refer to decoding / extracting XYZ from the encoded bitstream during the decoding process.
[0202] Here, in various embodiments of the present disclosure, the direction of a reference frame may be determined by whether the reference frame is before or after the current frame in the display order. In some implementations of the composite reference mode, when the picture order count (POC) of both reference frames of a motion vector pair is greater than or less than the POC of the current frame, the directions of the two reference frames are the same. Otherwise, when the POC of one reference frame is greater than the POC of the current frame while the POC of the other reference frame is less than the POC of the current frame, the directions of the two reference frames are different.
[0203] Here, in various embodiments of the present disclosure, a "block" may refer to a prediction block, an encoding / decoding block, a transform block, or a coding unit (CU).
[0204] Referring to step 1810, the device may be Figure 5 the electronic device (530) in Figure 8 or the video decoder (810) in Figure 6 In some implementations, the device may be the decoder (633) in the encoder (620) in Figure 5a part of the electronic device (530) in, Figure 8 a part of the video decoder (810) in, or Figure 6 a part of the decoder (633) in the encoder (620) in. The encoded video bitstream may be Figure 8 an encoded video sequence in, or Figure 6 or Figure 7 intermediate encoded data in.
[0205] Referring to step 1820, extract the inter - frame prediction mode and the joint delta motion vector (MV) of the current block in the current frame from the encoded video bitstream. The current block may be in a composite reference mode. The inter - frame prediction mode may include one of the following: NEAR_NEAR mode, NEW_NEARMV mode, NEAR_NEWMV mode, or NEW_NEWMV mode. The joint delta MV may be referred to as the MV difference (MVD).
[0206] Referring to step 1830, extract a first flag from the encoded video bitstream that indicates whether the first delta MV of the first reference frame and the second delta MV of the second reference frame are jointly signaled.
[0207] In some embodiments of some one or more inter - frame prediction modes, the first flag is encoded in the encoded video bitstream and can be extracted from the encoded video bitstream.
[0208] In some embodiments of some one or more inter - frame prediction modes, the first flag is not encoded in the encoded video bitstream and can be derived based on a default value according to one or more inter - frame prediction models. For example, the inter - frame prediction mode of the current block is NEAR_NEARMV; and step 1830 may include determining the first flag as the default value. In some embodiments, the default value is 0, indicating that the first delta MV of the first reference frame and the second delta MV of the second reference frame are not jointly signaled.
[0209] In some embodiments of the composite reference mode, the first flag (which may be named joint_mvd_flag) may be sent to the device to indicate whether the delta MVs of the first reference list (reference list 0) and the second reference list (reference list 1) are jointly signaled.
[0210] In some other embodiments, in response to the value of the first flag (joint_mvd_flag) indicating that the incremental MVs of reference list 0 and reference list 1 are signaled jointly, only one joint incremental MV (which may be named joint_delta_mv) is signaled and sent to the decoder, and the incremental MVs of reference list 0 and reference list 1 can be derived from the joint incremental MV (joint_delta_mv). In response to the value of the first flag (joint_mvd_flag) indicating that the incremental MVs of reference list 0 and reference list 1 are not signaled jointly, zero or one or two incremental MVs of reference list 0 and / or reference list 1 can be signaled separately based on the inter-frame prediction mode.
[0211] In some embodiments, a value of 0 for the first flag may indicate that the incremental MVs of reference list 0 and reference list 1 are signaled jointly, and only one joint incremental MV is signaled and sent; and a value of 1 for the first flag may indicate that the incremental MVs of reference list 0 and reference list 1 are not signaled jointly, and zero (or one or two) joint incremental MVs can be signaled and sent. Vice versa, in some other embodiments, a value of 1 for the first flag may indicate that the incremental MVs of reference list 0 and reference list 1 are signaled jointly, and only one joint incremental MV is signaled and sent; and a value of 0 for the first flag may indicate that the incremental MVs of reference list 0 and reference list 1 are not signaled jointly, and zero (or one or two) joint incremental MVs can be signaled and sent.
[0212] Referring to step 1840, in response to the first flag indicating that the first incremental MV and the second incremental MV are signaled jointly, the first incremental MV and the second incremental MV can be derived based on the joint incremental MV.
[0213] In some embodiments, when the inter-frame prediction mode of the current block is NEW_NEWMV and the first flag (e.g., joint_mvd_flag) indicates that the incremental MVs of reference list 0 and reference list 1 are signaled jointly, the incremental MVs of reference list 0 and / or reference list 1 can be derived from joint_delta_mv based on the POC distances from the first reference frame and the second reference frame to the current frame and the directions of the two reference frames.
[0214] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; step 1840 may include determining the first delta MV as a joint delta MV, and determining the second delta MV by scaling the joint delta MV according to at least one of the following: the first picture order count (POC) distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame. Here, the directional relationship is, for example, the same direction or the opposite direction.
[0215] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; step 1840 may include determining the second delta MV as a joint delta MV, and determining the first delta MV by scaling the joint delta MV according to at least one of the following: the first picture order count (POC) distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame.
[0216] In some embodiments, the delta MV in reference list 0 (or list 1) may always be set to be equal to joint_delta_mv, and the delta MV in reference list 1 (or list 0) may be scaled from joint_delta_mv according to the POC distance from the reference frame to the current frame and / or the direction of the two reference frames.
[0217] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; step 1840 may include, in response to the first absolute POC distance between the first reference frame and the current frame being greater than the second absolute POC distance between the second reference frame and the current frame: determining the first delta MV as a joint delta MV, and determining the second delta MV by scaling the joint delta MV according to at least one of the following: the first POC distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame.
[0218] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; step 1840 may include, in response to the first absolute POC distance between the first reference frame and the current frame being less than the second absolute POC distance between the second reference frame and the current frame: determining the second delta MV as a joint delta MV, and determining the first delta MV by scaling the joint delta MV according to at least one of the following: the first POC distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, or the directional relationship between the first reference frame and the second reference frame with respect to the current frame.
[0219] In some embodiments, when the absolute POC distance between reference list 0 (or list 1) and the current frame is greater than the absolute POC distance between reference list 1 (or list 0) and the current frame, the delta MV in reference list 0 (or list 1), which is the one with the larger absolute POC distance, may be set equal to joint_delta_mv. The delta MV in reference list 1 (or list 0), which is the other with the smaller absolute POC distance, may be scaled from joint_delta_mv according to the POC distance from the reference frame to the current frame and / or the directions of the two reference frames.
[0220] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEWMV; and step 1840 may include, in response to the first absolute POC distance between the first reference frame and the current frame being equal to the second absolute POC distance between the second reference frame and the current frame: determining the first delta MV as the joint delta MV, in response to the direction relationships of the first reference frame and the second reference frame with respect to the current frame being the same: determining the second delta MV as the joint delta MV, and in response to the direction relationships of the first reference frame and the second reference frame with respect to the current frame being opposite: determining the second delta MV as the joint delta MV multiplied by -1.
[0221] In some embodiments, when the absolute POC distance between reference list 1 and the current frame is the same as the absolute POC distance between reference list 0 and the current frame, the delta MV in reference list 0 may be set equal to joint_delta_mv. When the directions of the two reference frames are the same, the delta MV in reference list 1 may also be set equal to joint_delta_mv. Otherwise, when the directions of the two reference frames are different, the delta MV of reference list 1 is set to joint_delta_mv multiplied by -1.
[0222] In various embodiments of the present disclosure, a second incremental MV is obtained by scaling a combined incremental MV according to at least one of the following: a first POC distance between a first reference frame and the current frame, a second POC distance between a second reference frame and the current frame, or a directional relationship of the first reference frame and the second reference frame relative to the current frame. The scaling may include a linear scaling method, i.e., the absolute value of the scaled incremental MV may be proportional to the ratio of the second POC distance divided by the first POC distance, and the sign of the scaled incremental MV may be determined according to the directional relationship. For one example, when the first POC distance is 4 and the second POC distance is 8 and the directional relationship of the first and second reference frames relative to the current frame is the same direction, the second incremental MV is obtained by scaling / multiplying the combined incremental MV by a factor of 2 (= 8 / 4); and the second incremental MV has the same sign as the combined incremental MV because the directional relationship is the same direction. For another example, when the first POC distance is 3 and the second POC distance is -9 and the directional relationship of the first and second reference frames relative to the current frame is the opposite direction, the second incremental MV is obtained by scaling / multiplying the combined incremental MV by a factor of -3 (= -9 / 3); and the second incremental MV has the opposite sign as the combined incremental MV because the directional relationship is the opposite direction.
[0223] In some other embodiments, the inter-frame prediction mode of the current block is NEW_NEARMV; and the second incremental MV may be slightly adjusted by a predefined weight. Step 1840 may include: determining the first incremental MV as the combined incremental MV, and determining the second incremental MV by scaling the combined incremental MV according to at least one of the following: a first POC distance between a first reference frame and the current frame, a second POC distance between a second reference frame and the current frame, a directional relationship of the first reference frame and the second reference frame relative to the current frame, or a predefined weighting factor.
[0224] In some other embodiments, the inter-frame prediction mode of the current block is NEAR_NEWMV; and the first incremental MV may be slightly adjusted by a predefined weight. Step 1840 may include: determining the second incremental MV as the combined incremental MV, and determining the first incremental MV by scaling the combined incremental MV according to at least one of the following: a first POC distance between a first reference frame and the current frame, a second POC distance between a second reference frame and the current frame, a directional relationship of the first reference frame and the second reference frame relative to the current frame, or a predefined weighting factor.
[0225] In some other embodiments, the predefined weighting factor is a fraction between -1 and 1.
[0226] In some other embodiments, a predefined weighting factor is signaled in a high-level syntax that includes at least one of the following: Sequence Parameter Set (SPS), Video Parameter Set (VPS), Picture Parameter Set (PPS), picture header, tile header, slice header, frame header, Coding Tree Unit (CTU) header, or superblock header.
[0227] In some embodiments, when the current block is in the NEW_NEAR mode (or NEAR_NEW mode) and joint_mvd_flag indicates that the delta MVs for reference list 0 and reference list 1 are jointly signaled, the delta MV for reference list 0 (or list 1) can be set to be equal to joint_delta_mv, and the delta MV for list 1 (or list 0) is scaled based on one or more coded information from joint_delta_mv, the one or more coded information including but not limited to the POC distance from the reference frame to the current frame, the directions of two reference frames, the difference between the MV predictors of two MVs, and / or a predefined weighting factor w. In some embodiments, the predefined weighting factor can be a number between -1 and 1, such as 1 / 2. In some other embodiments, the predefined weighting factor can be signaled in a high-level syntax that includes but is not limited to SPS, VPS, PPS, picture header, tile header, slice header, frame header, CTU (or superblock) header.
[0228] In various embodiments of the present disclosure, a weighted-adjusted incremental MV is obtained by scaling a combined incremental MV according to at least one of the following: a first POC distance between a first reference frame and a current frame, a second POC distance between a second reference frame and the current frame, a directional relationship of the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor. The scaling may include a linear scaling method, i.e., the absolute value of the scaled incremental MV may be proportional to the ratio of the second POC distance divided by the first POC distance (and then multiplied by the predefined weighting factor). The sign of the scaled incremental MV may be determined according to the directional relationship. For one example, when the first POC distance is 4 and the second POC distance is 8 and the directional relationship of the first and second reference frames with respect to the current frame is the same direction and the predefined weighting factor is 1 / 2, the weighted-adjusted incremental MV is obtained by scaling / multiplying the combined incremental MV by a total factor 1, which is calculated according to the factor 2 (=8 / 4) multiplied by the weighting factor 1 / 2; and the weighted-adjusted incremental MV has the same sign as the combined incremental MV because the directional relationship is the same direction. For another example, when the first POC distance is 3 and the second POC distance is -9 and the directional relationship of the first and second reference frames with respect to the current frame is the opposite direction and the predefined weighting factor is 1 / 2, the weighted-adjusted incremental MV is obtained by scaling / multiplying the combined incremental MV by a total factor -3 / 2, which is calculated according to the factor -3 (= -9 / 3) multiplied by the weighting factor 1 / 2. The weighted-adjusted incremental MV has the opposite sign as the combined incremental MV because the directional relationship is the opposite direction.
[0229] The embodiments in the present disclosure may be used alone or in any combination in any order. As needed, any steps and / or operations in any embodiment in the present disclosure may be combined or arranged in any number or order. Two or more steps and / or operations in any embodiment of the present disclosure may be performed in parallel. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium. The embodiments of the present disclosure may be applied to a luminance block or a chrominance block.
[0230] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 19 FIG. shows a computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter.
[0231] Computer software can be encoded using any suitable machine code or computer language, which can be created through assembly, compilation, linking, or similar mechanisms to create code including instructions that can be executed directly or through interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0232] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0233] Figure 19 The components shown for the computer system (2000) are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependence on or requirement for any one component or combination thereof illustrated in the exemplary embodiments of the computer system (2000).
[0234] The computer system (2000) may include certain human-machine interface input devices. Such human-machine interface input devices can respond to input from one or more human users through, for example, tactile input (such as: keystrokes, swipes, data glove movements), audio input (such as: voice, taps), visual input (such as: gestures), olfactory input (not shown). The human-machine interface device can also be used to capture certain media not necessarily directly related to a human's conscious input, such as audio (such as: voice, music, ambient sound), images (such as: scanned images, photographic images obtained from a still image camera), video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0235] The input human-machine interface devices can include one or more of the following (only one of each is depicted): keyboard (2001), mouse (2002), touchpad (2003), touch screen (2010), data glove (not shown), joystick (2005), microphone (2006), scanner (2007), camera (2008).
[0236] The computer system (2000) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as the tactile feedback of a touch screen (2010), a data glove (not shown), or a joystick (2005), but there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (2009), headphones (not depicted)), visual output devices (such as a screen (2010), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted), holographic displays, and fog machines (not depicted)), and printers (not depicted).
[0237] The computer system (2000) may also include human-accessible storage devices and their associated media, such as optical media including media (2021) such as CD / DVD ROM / RW (2020) with CD / DVD, thumb drives (2022), removable hard disk drives, or solid-state drives (2023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.
[0238] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other volatile signals.
[0239] The computer system (2000) may also include an interface (2054) to one or more communication networks (2055). The network can be, for example, wireless, wired, optical. The network can also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CAN bus, etc. Some networks typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (2049) (such as, for example, the USB port of the computer system (2000)); other networks are typically integrated into the core of the computer system (2000) by attaching to the system bus as described below (for example, an Ethernet interface into a PC computer system or a cellular network interface into a smart phone computer system). Using any of these networks, the computer system (2000) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way, such as to other computer systems using local digital networks or wide area digital networks. Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.
[0240] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core (2040) of the computer system (2000).
[0241] The core (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), special-purpose programmable processing units in the form of field-programmable gate arrays (FPGAs) (2043), hardware accelerators (2044) for certain tasks, graphics adapters (2050), etc. These devices, together with read-only memory (ROM) (2045), random access memory (2046), internal mass storage such as internal non-user-accessible hard disk drives, SSDs, etc. (2047), can be connected via a system bus (2048). In some computer systems, the system bus (2048) can be accessed in the form of one or more physical plugs to enable expansion by attaching additional CPUs, GPUs, etc. Peripheral devices can be attached directly or via a peripheral bus (2049) to the system bus (2048) of the core. In one example, a screen (2010) can be connected to the graphics adapter (2050). The architecture of the peripheral bus includes PCI, USB, etc.
[0242] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can execute certain instructions, and the combination of these instructions can constitute the above-mentioned computer code. This computer code can be stored in the ROM (2045) or RAM (2046). Transitional data can also be stored in the RAM (2046), while permanent data can be stored in, for example, the internal mass storage (2047). Fast storage and retrieval of any memory device can be enabled by using a cache memory, which can be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc.
[0243] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be those specially designed and constructed for the purposes of this disclosure, or they can be of the type well-known and available to those skilled in the computer software art.
[0244] As a non-limiting example, a computer system (2000) having an architecture, and in particular a core (2040), can provide functionality as a result of software embodied in one or more tangible computer-readable media being executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be media associated with the user-accessible mass storage introduced above, as well as certain memories of the core (2040) having a non-volatile nature, such as the core internal mass storage (2047) or ROM (2045). The software implementing various embodiments of this disclosure can be stored in such devices and executed by the core (2040). Depending on specific needs, the computer-readable medium can include one or more memory devices or chips. The software can cause the core (2040) and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processes described herein or specific parts of specific processes, including defining data structures stored in the RAM (2046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of being logically hardwired or otherwise embodied in circuitry (e.g., accelerator (2044)), which can operate in place of or in conjunction with the software to perform specific processes described herein or specific parts of specific processes. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to a computer-readable medium can include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. This disclosure encompasses any suitable combination of hardware and software.
[0245] While a particular invention has been described with reference to illustrative embodiments, such description is not meant to be limiting. Various modifications of the illustrative embodiments and additional embodiments of the present invention will be apparent to those of ordinary skill in the art from this description. Those skilled in the art will readily recognize that these and various other modifications can be made to the exemplary embodiments illustrated and described herein without departing from the spirit and scope of the present invention. Accordingly, it is contemplated that the appended claims will cover any such modifications and alternative embodiments. Certain proportions in the figures may be exaggerated while others may be minimized. Accordingly, the present disclosure and the figures are considered illustrative rather than restrictive.
[0246] The following is a list of abbreviations, some of which may appear in this disclosure:
[0247] JEM: joint exploration model
[0248] VVC: versatile video coding
[0249] BMS: benchmark set
[0250] MV: Motion Vector
[0251] HEVC: High Efficiency Video Coding
[0252] SEI: Supplementary Enhancement Information
[0253] VUI: Video Usability Information
[0254] GOPs: Groups of Pictures
[0255] TUs: Transform Units
[0256] PUs: Prediction Units
[0257] CTUs: Coding Tree Units
[0258] CTBs: Coding Tree Blocks
[0259] PBs: Prediction Blocks
[0260] HRD: Hypothetical Reference Decoder
[0261] SNR: Signal Noise Ratio
[0262] CPUs: Central Processing Units
[0263] GPUs: Graphics Processing Units
[0264] CRT: Cathode Ray Tube
[0265] LCD: Liquid-Crystal Display
[0266] OLED: Organic Light-Emitting Diode
[0267] CD: Compact Disc
[0268] DVD: Digital Video Disc
[0269] ROM: Read-Only Memory
[0270] RAM: Random Access Memory
[0271] ASIC: Application-Specific Integrated Circuit
[0272] PLD: Programmable Logic Device
[0273] LAN: Local Area Network
[0274] GSM: Global System for Mobile communications
[0275] LTE: Long-Term Evolution
[0276] CANBus: Controller Area Network Bus
[0277] USB: Universal Serial Bus
[0278] PCI: Peripheral Component Interconnect
[0279] FPGA: Field Programmable Gate Areas
[0280] SSD: solid-state drive
[0281] IC: Integrated Circuit
[0282] HDR: High Dynamic Range
[0283] SDR: Standard Dynamic Range
[0284] JVET: Joint Video Exploration Team
[0285] MPM: Most Probable Mode
[0286] WAIP: Wide-Angle Intra Prediction
[0287] CU: Coding Unit
[0288] PU: Prediction Unit
[0289] TU: Transform Unit
[0290] CTU: Coding Tree Unit
[0291] PDPC: position dependent prediction combination
[0292] ISP: Intra Sub-Partitions
[0293] SPS: Sequence Parameter Setting
[0294] PPS: Picture Parameter Set
[0295] APS: Adaptation Parameter Set
[0296] VPS: Video Parameter Set
[0297] DPS: Decoding Parameter Set
[0298] ALF: Adaptive Loop Filter
[0299] SAO: Sample Adaptive Offset
[0300] CC-ALF: Cross-Component Adaptive Loop Filter
[0301] CDEF: Constrained Directional Enhancement Filter
[0302] CCSO: Cross-Component Sample Offset
[0303] LSO: Local Sample Offset
[0304] LR: Loop Restoration Filter
[0305] AV1: AOMedia Video 1
[0306] AV2: AOMedia Video 2
[0307] MVD: Motion Vector difference
[0308] CfL: Chroma from Luma
[0309] SDT: Semi Decoupled Tree
[0310] SDP: Semi Decoupled Partitioning
[0311] SST: Semi Separate Tree
[0312] SB: Super Block
[0313] IBC(or IntraBC): Intra Block Copy
[0314] CDF: Cumulative Density Function
[0315] SCC: Screen Content Coding
[0316] GBI: Generalized Bi-prediction
[0317] BCW: Bi-prediction with CU-level Weights
[0318] CIIP: Combined intra-inter prediction
[0319] POC: Picture Order Count
[0320] RPS: Reference Picture Set
[0321] DPB: Decoded Picture Buffer
[0322] MMVD: Merge Mode with Motion Vector Difference
Claims
1. A method for video decoding, characterized in that, The method includes: Receiving an encoded video bitstream; Extracting, from the encoded video bitstream, an inter prediction mode and a joint delta motion vector MV of a current block in a current frame, denoted as joint_delta_mv; Extracting a first flag joint_mvd_flag from the encoded video bitstream, the first flag joint_mvd_flag indicating whether a first delta MV of a first reference frame in reference list 0 and a second delta MV of a second reference frame in reference list 1 are jointly signaled; In response to the first flag joint_mvd_flag indicating that the first delta MV and the second delta MV are jointly signaled, deriving the first delta MV and the second delta MV based on the joint delta MV; and Decoding the current block in the current frame based on the first delta MV and the second delta MV; wherein, when the inter prediction mode of the current block is NEW_NEARMV, where NEW_NEARMV indicates: using a motion vector predictor MVP indicated by a dynamic reference list DRL index as a first reference MV, and signaling a delta MV for the first reference MV, and using a motion vector predictor MVP indicated by another DRL index as a second reference MV, The deriving the first delta MV and the second delta MV based on the joint delta MV includes: Determining the first delta MV as the joint delta MV, and determining the second delta MV by scaling the joint delta MV according to at least one of: a first picture order count POC distance between the first reference frame and the current frame, a second picture order count POC distance between the second reference frame and the current frame, a directional relationship between the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor, wherein the predefined weighting factor is used to determine a scaling factor for the joint delta MV according to the following: multiplying a ratio of dividing the second POC distance by the first POC distance by the predefined weighting factor; or Determining the second delta MV as the joint delta MV, and determining the first delta MV by scaling the joint delta MV according to at least one of: the first POC distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, the directional relationship between the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor, wherein the predefined weighting factor is used to determine a scaling factor for the joint delta MV according to the following: multiplying a ratio of dividing the second POC distance by the first POC distance by the predefined weighting factor; wherein the predefined weighting factor is a fraction between -1 and 1.
2. The method according to claim 1, wherein: When the inter - frame prediction mode of the current block is NEW_NEWMV, where NEW_NEWMV indicates that: the motion vector predictor MVP indicated by one dynamic reference list DRL index is used as the first reference MV, and the motion vector predictor MVP indicated by another dynamic reference list DRL index is used as the second reference MV, The deriving of the first incremental MV and the second incremental MV based on the joint incremental MV includes: Determining the first incremental MV as the joint incremental MV, and Determining the second incremental MV by scaling the joint incremental MV according to at least one of the following: The first picture order count POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, and The directional relationship between the first reference frame and the second reference frame with respect to the current frame.
3. The method according to claim 2, wherein: The absolute value of the scaled joint incremental MV is proportional to the ratio of the second POC distance divided by the first POC distance, and the sign of the scaled joint incremental MV is determined according to the directional relationship.
4. The method according to claim 1, wherein: When the inter - frame prediction mode of the current block is NEW_NEWMV, The deriving of the first incremental MV and the second incremental MV based on the joint incremental MV includes: Determining the second incremental MV as the joint incremental MV, and Determining the first incremental MV by scaling the joint incremental MV according to at least one of the following: The first picture order count POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or The directional relationship between the first reference frame and the second reference frame with respect to the current frame.
5. The method according to claim 1, wherein: When the inter - frame prediction mode of the current block is NEW_NEWMV, The deriving of the first incremental MV and the second incremental MV based on the joint incremental MV includes: In response to the first absolute POC distance between the first reference frame and the current frame being greater than the second absolute POC distance between the second reference frame and the current frame: Determining the first incremental MV as the joint incremental MV, and Determining the second incremental MV by scaling the joint incremental MV according to at least one of the following: The first POC distance between the first reference frame and the current frame, The second POC distance between the second reference frame and the current frame, or The directional relationship between the first reference frame and the second reference frame with respect to the current frame.
6. The method according to claim 1, wherein: When the inter - frame prediction mode of the current block is NEW_NEWMV, The deriving of the first incremental MV and the second incremental MV based on the joint incremental MV includes: In response to a first absolute POC distance between the first reference frame and the current frame being less than a second absolute POC distance between the second reference frame and the current frame: determine the second delta MV as the joint delta MV, and determine the first delta MV by scaling the joint delta MV according to at least one of the following: a first POC distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, or a directional relationship of the first reference frame and the second reference frame with respect to the current frame.
7. The method according to claim 1, wherein: in the case where the inter prediction mode of the current block is NEW_NEWMV, the deriving the first delta MV and the second delta MV based on the joint delta MV includes: in response to a first absolute POC distance between the first reference frame and the current frame being equal to a second absolute POC distance between the second reference frame and the current frame: determine the first delta MV as the joint delta MV, in response to a directional relationship of the first reference frame and the second reference frame with respect to the current frame being the same: determine the second delta MV as the joint delta MV, and in response to the directional relationship of the first reference frame and the second reference frame with respect to the current frame being opposite: determine the second delta MV as the joint delta MV multiplied by -1.
8. The method according to claim 1, wherein: signaling the predefined weighting factor in a high-level syntax, the high-level syntax including at least one of the following: sequence parameter set SPS, video parameter set VPS, picture parameter set PPS, picture header, tile header, slice header, frame header, coding tree unit CTU header, or super block header.
9. The method according to claim 1, wherein: the inter prediction mode of the current block is NEAR_NEARMV; and the extracting the first flag joint_mvd_flag, the first flag joint_mvd_flag indicating whether the first delta MV of the first reference frame and the second delta MV of the second reference frame are signaled jointly, includes: determine the first flag joint_mvd_flag as a default value.
10. The method according to claim 9, wherein: the default value is 0, indicating that the first delta MV of the first reference frame and the second delta MV of the second reference frame are not signaled jointly.
11. A method for video encoding, characterized in that, The method includes: acquiring a source video sequence; encoding the source video sequence to obtain an encoded video bitstream; in the encoded video bitstream, signaling an inter prediction mode and a joint delta motion vector MV, denoted as joint_delta_mv, of a current block in a current frame; In an encoded video bitstream, a first flag, joint_mvd_flag, is signaled, where the first flag, joint_mvd_flag, indicates whether a first delta MV of a first reference frame in reference list 0 and a second delta MV of a second reference frame in reference list 1 are jointly signaled; wherein, when the first flag, joint_mvd_flag, indicates that the first delta MV and the second delta MV are jointly signaled, the joint delta MV is used to derive the first delta MV and the second delta MV, and the first delta MV and the second delta MV are used to decode the current block in the current frame; wherein, when the inter prediction mode of the current block is NEW_NEARMV, where NEW_NEARMV indicates: using a motion vector predictor, MVP, indicated by one dynamic reference list (DRL) index as a first reference MV, and signaling a delta MV for the first reference MV, and using a motion vector predictor, MVP, indicated by another DRL index as a second reference MV, the joint delta MV is used to derive the first delta MV and the second delta MV according to the following: determining the first delta MV as the joint delta MV, and determining the second delta MV by scaling the joint delta MV according to at least one of: a first picture order count (POC) distance between the first reference frame and the current frame, a second POC distance between the second reference frame and the current frame, a directional relationship between the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor, where the predefined weighting factor is used to determine a scaling factor for the joint delta MV according to: multiplying a ratio of dividing the second POC distance by the first POC distance by the predefined weighting factor; or determining the second delta MV as the joint delta MV, and determining the first delta MV by scaling the joint delta MV according to at least one of: the first POC distance between the first reference frame and the current frame, the second POC distance between the second reference frame and the current frame, the directional relationship between the first reference frame and the second reference frame with respect to the current frame, or a predefined weighting factor, where the predefined weighting factor is used to determine a scaling factor for the joint delta MV according to: multiplying a ratio of dividing the second POC distance by the first POC distance by the predefined weighting factor; where the predefined weighting factor is a fraction between -1 and 1.
12. A video decoding device, characterized in that, The apparatus comprises: a memory storing instructions; and a processor in communication with the memory, where when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 1 to 10.
13. A video encoding device, characterized in that, The apparatus comprises: a memory storing instructions; and A processor, communicating with the memory, wherein when the processor executes the instructions, the processor is configured to cause the device to perform the method according to claim 11.
14. A non-volatile computer-readable storage medium stores instructions, characterized in that, When the instructions are executed by the processor, the instructions are configured to cause the processor to perform the method according to any one of claims 1 to 11.
15. A method for processing a video bitstream, characterized in that The video bitstream is generated according to the encoding method of claim 11, or decoded according to the method of any one of claims 1-10.
Citation Information
Patent Citations
Joint signaling method for motion vector difference
CN117063471A
Method and apparatus for processing video signals on basis of inter prediction
US20220038707A1